An AI employee pilot should answer one purchase or deployment question: Can this platform perform one recurring role well enough, safely enough, and economically enough to justify the next level of commitment?
That requires a baseline, representative work, controlled authority, explicit metrics, negative tests, and decision thresholds written before results arrive. A collection of impressive demos is not a pilot. A 30-day trial with no comparison, no acceptance definition, and no stop condition is only extended product exploration.
Use the AI employee platform evaluation framework to choose candidates. Then run the same bounded experiment for each qualified option.
On this page · 16 sectionsOpen
- What Is an AI Employee Pilot?
- What Decision Should the Pilot Produce?
- How Do You Choose the Right Pilot Role?
- How Do You Establish a Baseline Before the Pilot?
- How Do You Build a Representative Pilot Sample?
- What Pilot Modes Should You Use?
- Which Controls Must Exist Before the Pilot?
- How Long Should an AI Employee Pilot Run?
- Which Metrics and Thresholds Should Decide the Result?
- How Do You Run the Pilot Week by Week?
- What Evidence Should the Final Pilot Report Contain?
- Who Should Own and Review the Pilot?
- How Do You Control Bias and Uncertainty in the Result?
- How Does CellCog Fit a Pilot?
- Which Pilot Mistakes Invalidate the Decision?
- Define the Decision Before Starting the Pilot
- Define the decision before creating the account: go, revise and retest, switch candidate, or stop.
- Pilot one role with one accountable owner, one recurring outcome, representative normal and hard cases, and a measured human or current-system baseline.
- Start in test or shadow mode, then advance to draft and approved-action modes only after evidence gates pass. Do not grant independent external authority because the calendar reached week 3.
- Measure accepted outcomes, first-pass quality, correction time, escalation, reliability, latency, usage, total cost, and worst-error severity.
- Write pass thresholds, knockout failures, budget limits, and confidence requirements before the first scored run.
- End with a decision record and evidence export. Do not let an inconclusive pilot roll quietly into production.
§ 01What Is an AI Employee Pilot?
An AI employee pilot is a bounded, time-limited evaluation of a defined standing role under conditions that approximate real operation without exposing the organization to uncontrolled production risk.
| Pilot property | Required definition | Failure if missing |
|---|---|---|
| Decision | Go, revise, switch, or stop | Trial continues without a conclusion |
| Role | Outcome, owner, boundaries, sources, actions | Product demo replaces job evaluation |
| Baseline | Current quality, time, cost, risk, and volume | Improvement cannot be calculated |
| Sample | Normal, hard, boundary, failure, and prohibited cases | Happy-path performance is overstated |
| Modes | Test, shadow, draft, approved action | Authority expands by convenience |
| Metrics | Formulas, sources, thresholds, owner | Dashboard activity replaces outcomes |
| Controls | Permissions, approvals, logs, stop path | Experiment creates production exposure |
| Duration | Evidence gates and maximum calendar | Pilot drifts indefinitely |
| Exit | Export, cleanup, handover, decision record | Learning and data are lost |
The pilot is narrower than onboarding. The AI employee onboarding guide defines the full graduated operating path. The pilot is the experiment that decides whether the role and platform should advance.
A pilot tests a claim
Write one falsifiable claim:
For the weekly acquisition-review role, the candidate will produce at least 16 accepted reports across representative cases, meet the defined quality and risk gates, reduce median review time versus baseline, and remain inside the agreed cost ceiling.
The numbers are an illustrative contract. Your volume and thresholds must come from the actual workload and consequence.
A pilot is not a showcase
Exclude:
- vendor-selected prompts only;
- one polished artifact;
- public benchmark results alone;
- employee activity without accepted outcomes;
- a synthetic task unrelated to the real role;
- unlimited human repair hidden from the result;
- broad credentials used for setup convenience; and
- qualitative enthusiasm as the decision threshold.
These may support exploration. They do not support a purchase decision.
A pilot is not production by another name
Production use has real recipients, records, money, customers, and consequences. A pilot should minimize those exposures while testing enough of the operating path to reveal decisive uncertainty.
Use shadow results, drafts, sandbox systems, test accounts, synthetic or minimized data, explicit approvals, reversible actions, and narrow recipient lists until the role passes the relevant gates.
§ 02What Decision Should the Pilot Produce?
Choose the terminal states before selecting tasks.
Go
Advance when:
- all knockout controls pass;
- the sample is sufficiently representative;
- minimum quality and outcome thresholds pass;
- review and correction remain within ceiling;
- cost fits the approved case;
- no unresolved severe failure remains;
- owner and operating process exist; and
- the next authority level is explicitly bounded.
“Go” should authorize a specific next state, not unlimited deployment.
Revise and retest
Use this outcome when the role remains viable and one bounded cause is repairable:
- unclear instruction;
- missing source;
- tool schema defect;
- overly broad task;
- weak evaluation rubric;
- incorrect permission;
- known connector issue; or
- insufficient but obtainable sample.
Set one revision cycle, owner, changed component, regression set, budget, and deadline. Do not restart the whole pilot without preserving prior cases.
Switch candidate
Switch when the role is valid but the platform cannot meet a mandatory capability, control, cost, or support requirement.
Carry the same role card, cases, rubric, and evidence format to the next candidate. Changing the experiment would destroy comparability.
Stop
Stop when:
- the role is not worth delegating;
- useful work cannot be defined or evaluated;
- a mandatory risk boundary cannot be enforced;
- failures remain severe after the agreed repair;
- human review erases the value;
- total cost exceeds the high case;
- no accountable owner exists; or
- the current process is already better.
Stopping is a successful pilot outcome when it prevents an unsupported deployment.
§ 03How Do You Choose the Right Pilot Role?
Pick the smallest recurring responsibility that exposes the decision.
Use role-fit criteria
| Criterion | Strong pilot signal | Weak pilot signal |
|---|---|---|
| Recurrence | Weekly or higher frequency | Rare annual event |
| Digital inputs | Available through approved files, messages, or systems | Mostly physical or undocumented context |
| Acceptance | Output can be checked | “We will know good work when we see it” |
| Reversibility | Draft, recommend, or reversible write | Irreversible consequential action |
| Evidence | Sources, tool results, and state are visible | Success depends on private intuition |
| Volume | Enough cases inside the pilot window | Too few cases for any pattern |
| Owner | One person can accept and escalate | Committee without decision rights |
| Value | Time, quality, capacity, or risk benefit is plausible | Novelty is the main benefit |
Use the outcome-first AI employee hiring process if the role has not yet been defined. The pilot should verify a role contract, not invent it while the vendor is being evaluated.
Start with one recurring outcome
Good pilot outcomes include:
- accepted weekly market report;
- reconciled operating dashboard;
- approved content draft with verified claims;
- correctly triaged support queue;
- accepted account-research brief;
- prepared CRM update after approval;
- complete project status pack; or
- prioritized inbox with drafted responses.
Avoid “help the marketing team.” It has no stable unit of work.
Avoid the easiest and hardest role
The easiest task may prove only that the model can summarize. The hardest task may fail because the organization selected an unsafe or undefined process.
Choose a role with meaningful ambiguity, bounded consequence, enough cases, accessible sources, and a realistic path to recurring use.
§ 04How Do You Establish a Baseline Before the Pilot?
Measure the current process using the same outcome definition.
Capture the current workflow
For at least one representative period, record:
- work received;
- work started;
- work completed;
- work accepted;
- first-pass acceptance;
- reviewer minutes;
- correction minutes;
- elapsed time;
- exceptions;
- escalations;
- missed work;
- cost;
- severe errors; and
- sources or system state used.
Do not compare the AI’s active runtime with a person’s full workday. Compare the complete workflow cost and accepted output.
Use a metric contract
| Metric | Formula | Source | Why it matters |
|---|---|---|---|
| Accepted-outcome rate | Accepted / eligible outcomes | Task ledger + reviewer decision | Core usefulness |
| First-pass acceptance | Accepted without correction / submitted | Review record | Hidden rework |
| Review minutes/outcome | Reviewer minutes / accepted outcomes | Time log | Human bottleneck |
| Correction minutes/outcome | Correction minutes / accepted outcomes | Time log | Quality cost |
| Escalation recall | Required escalations raised / required escalations | Labeled cases | Safety judgment |
| False completion | Incomplete work marked complete / submitted | State audit | Reliability |
| Cost/accepted outcome | Total cost / accepted outcomes | Usage + labor + failures | Economic comparison |
| Worst-error severity | Maximum consequence class observed | Incident rubric | Tail risk |
The AI employee KPI framework gives the full formulas and anti-gaming rules. The pilot selects the smallest set needed for the decision.
Measure the baseline fairly
The human process may include expertise acquired over years, ad hoc help, hidden overtime, and unlogged correction. The AI process may include setup, prompt tuning, vendor support, and enthusiastic reviewer attention.
Record both sides completely. Do not make the human baseline artificially expensive or the AI pilot artificially free.
Freeze the comparison period
Avoid comparing a quiet pilot month with a peak baseline quarter. Segment by:
- workload type;
- difficulty;
- data availability;
- source version;
- recipient;
- urgency;
- risk tier; and
- known seasonal condition.
If the periods differ, state the limitation and use matched cases where possible.
§ 05How Do You Build a Representative Pilot Sample?
The sample should reflect the deployment distribution and deliberately include consequential edges.
Create five case groups
| Case group | Purpose | Illustrative share |
|---|---|---|
| Normal | Estimate ordinary performance | 50-70% |
| Hard but valid | Test complexity inside role | 15-25% |
| Boundary | Test when scope or evidence is unclear | 5-15% |
| Tool/source failure | Test recovery and incomplete state | 5-10% |
| Prohibited/adversarial | Test denial and escalation | At least several explicit cases |
The percentages are design examples, not statistical standards. Use the real production distribution, then oversample low-frequency high-consequence cases for control testing.
Use enough repeated work
One success proves little because agent outputs vary and operating conditions change.
For a small knowledge-work pilot, 20-50 representative cases can expose patterns if the cases are diverse and each is deeply reviewed. A high-volume routing role may need hundreds or thousands of cases to estimate rates with useful precision.
Choose volume from:
- baseline case frequency;
- consequence of error;
- expected performance gap;
- task diversity;
- output variability;
- reviewer capacity; and
- confidence required for the decision.
Do not claim statistical significance without a proper design.
Preserve a holdout
Keep some cases away from the configuration process. If every failed case becomes an example in the prompt and is rerun until it passes, the pilot measures memorization of the test pack.
Use:
- development cases for iteration;
- holdout cases for the purchase decision;
- regression cases from every discovered failure; and
- live shadow cases to test current variation.
Anthropic’s current agent-evaluation guide distinguishes capability tests from regression suites and recommends examining tasks, repeated trials, graders, and complete execution records. The pilot should retain enough evidence to explain why a score changed.
Test negative behavior
At minimum, include cases where the agent should:
- refuse a prohibited action;
- request missing information;
- identify conflicting sources;
- escalate uncertainty;
- stop after a tool failure;
- avoid duplicate execution;
- ignore an instruction embedded in untrusted content;
- avoid writing sensitive data to memory;
- honor an expired or denied approval; and
- leave the task open rather than claim false completion.
A role that succeeds only when everything works is not ready.
§ 06What Pilot Modes Should You Use?
Advance authority through evidence gates.
Mode 0: configuration and sandbox
Use test data or a sandbox to verify role, sources, tools, output schema, logs, permissions, approval, denial, and stop behavior.
Exit only when:
- the expected tool path works;
- prohibited tools fail;
- evidence is reconstructable;
- credentials can be revoked;
- cost and retry limits work; and
- no live external effect occurs.
Mode 1: retrospective replay
Run historical cases with known outcomes. This supports fast iteration and matched comparison.
Protect against leakage: if the expected answer appears in the source pack or prompt, the result is not a genuine test.
Mode 2: live shadow
The AI receives current work in parallel with the existing process but does not control the official output or action.
Shadow mode reveals:
- current source access;
- real timing;
- queue behavior;
- changing instructions;
- exception frequency;
- handoff quality; and
- reviewer burden.
It cannot prove live-action reliability because no live action occurs.
Mode 3: draft and recommendation
The AI creates a draft, recommendation, or proposed system change. A person accepts, edits, rejects, or executes.
Record the delta between proposal and accepted outcome. A final artifact that looks good after 45 minutes of hidden repair is not a first-pass success.
Mode 4: approved action
The AI prepares an exact action and waits for parameter-bound human approval. Examples include sending one message, updating specified CRM fields, or publishing an approved artifact.
Test:
- action preview;
- approver identity;
- exact parameters;
- expiration;
- modification after approval;
- duplicate submission;
- denial;
- execution result;
- rollback; and
- evidence.
OpenAI’s practical guide to building agents identifies high-risk action and repeated failure as important triggers for human intervention. Approval should therefore follow consequence and observed performance, not a generic “human in the loop” label.
Mode 5: bounded independent action
Allow only after prior modes pass and only for low-risk, reversible, observable actions inside explicit limits.
Independent does not mean unsupervised. Monitoring, alerts, periodic review, stop conditions, and accountable human ownership remain.
§ 07Which Controls Must Exist Before the Pilot?
Controls should match the action and data risk.
Permission and approval contract
The action-specific permissions guide maps identity, environment, system, object, data, action, condition, limit, approval, evidence, and recovery.
For the pilot, document:
- what the agent can read;
- what it can prepare;
- what it can change only after approval;
- what it can change inside a limit;
- what is always blocked;
- which identity performs each operation;
- who can approve;
- how approval expires;
- how access is revoked; and
- what evidence proves enforcement.
Data and memory boundary
Specify:
- approved sources;
- prohibited data;
- inference path;
- storage location;
- allowed memory writes;
- retention;
- deletion;
- export;
- test-data handling; and
- whether vendor personnel may access pilot data.
Do not use production secrets or sensitive records merely because setup is temporary.
Monitoring and stop path
The owner needs visibility into:
- trigger received;
- task created;
- current state;
- source accessed;
- tool requested and executed;
- permission and approval result;
- output and artifact;
- retries;
- usage and cost;
- error;
- escalation; and
- final acceptance.
Define who can pause schedules, stop a run, revoke credentials, cancel queued work, preserve evidence, and return ownership to a person.
The NIST AI RMF Core calls for evaluation under conditions similar to deployment, production monitoring, override, incident response, recovery, change management, and a decision about whether deployment should proceed. Those lifecycle outcomes fit the pilot’s evidence gates.
Security abuse cases
The current OWASP AI Agent Security Cheat Sheet recommends least privilege, approval for high-impact actions, memory isolation, structured logs, and explicit limits. It also recommends repeatable tests for prompt override, tool misuse, privilege escalation, memory poisoning, data exfiltration, recursive abuse, approval bypass, and multi-agent chaining.
Select the abuse cases relevant to the role and retain the tested version and result.
§ 08How Long Should an AI Employee Pilot Run?
Duration follows evidence, not a universal 30-day calendar.
Set a minimum and maximum
Use:
- minimum number of representative cases;
- minimum number of live recurring cycles;
- required negative tests;
- required approved actions, if relevant;
- maximum calendar duration;
- maximum budget; and
- maximum revision cycles.
A weekly reporting role may need 4-8 real cycles. A high-volume triage role may generate enough cases in 5 business days but still need time to observe operational drift and escalation.
Use evidence gates by phase
| Phase | Minimum evidence | Advance when |
|---|---|---|
| Configure | Role, data, tools, permissions, rubric | Controls and ordinary task work |
| Replay | Historical normal and edge cases | Minimum offline thresholds pass |
| Shadow | Current workload and timing | Quality, escalation, and reliability hold |
| Draft | Human decision on every output | Review and correction fit ceiling |
| Approved action | Exact action evidence | Approval, execution, and recovery pass |
| Decision | Full ledger and cost | Confidence and all knockout gates pass |
Do not advance after a date alone.
Stop early when evidence is decisive
Stop for:
- prohibited action succeeds;
- sensitive data crosses a forbidden boundary;
- access cannot be revoked;
- severe repeated hallucination affects the role’s purpose;
- logs cannot reconstruct consequential action;
- budget or retry cap fails;
- vendor cannot resolve a mandatory gap; or
- the role owner withdraws accountability.
An early stop protects both budget and risk.
§ 09Which Metrics and Thresholds Should Decide the Result?
Use a balanced scorecard with non-compensating risk gates.
Outcome and quality
Measure:
- eligible outcomes;
- accepted outcomes;
- first-pass acceptance;
- rubric score;
- material omission;
- grounded-claim rate where applicable;
- calculation reconciliation;
- false completion; and
- reopen rate.
Volume never compensates for low acceptance.
Human effort
Measure:
- setup hours;
- reviewer minutes;
- correction minutes;
- approval wait time;
- escalation handling;
- incident handling; and
- weekly management time.
If a role saves production time and creates equal supervision time, the business case may fail.
Reliability and speed
Measure:
- trigger success;
- completion rate;
- end-to-end latency;
- timeout;
- duplicate action;
- retry;
- stuck task age;
- tool error;
- recovery time; and
- schedule adherence.
Report distributions, not averages alone. A median can hide a severe tail.
Risk
Measure:
- prohibited-action attempts and successes;
- missed required escalation;
- unnecessary escalation;
- sensitive-data exposure;
- approval bypass;
- memory-policy breach;
- unsupported claim severity;
- security test result; and
- worst observed consequence.
Keep worst-error severity next to the average quality score.
Economics
Pilot cost per accepted outcome = (platform + usage + setup amortization + integration + review + correction + monitoring + expected failure) ÷ accepted outcomes
Use current, measured usage and include all 7 cost layers rather than the plan price alone.
Illustrative threshold sheet
| Decision metric | Illustrative pass | Review band | Knockout |
|---|---|---|---|
| Accepted-outcome rate | 90% or higher | 80-89% | - |
| First-pass acceptance | 75% or higher | 60-74% | - |
| Review minutes/outcome | 10 or fewer | 11-20 | - |
| Required-escalation recall | 100% on severe cases | Below 100% on non-severe | Missed severe escalation |
| False completion | 0 severe | Low non-severe rate | Severe false completion |
| Prohibited action | 0 successes | - | Any success |
| Cost/accepted outcome | At/below approved base | Between base and high | Above hard budget |
| Evidence completeness | 100% for consequential actions | Repairable gaps | Unreconstructable consequential action |
These thresholds are illustrative only. Set them from the baseline, risk tolerance, and business case before the pilot.
§ 10How Do You Run the Pilot Week by Week?
A 30-day structure works for many recurring knowledge roles when enough cases arrive.
Before day 1
- approve the role card;
- approve the data and action boundary;
- measure baseline;
- freeze task pack and holdout;
- define metrics and thresholds;
- assign reviewers;
- configure logging and cost limits;
- test revocation and stop;
- record product, model, prompt, tool, and policy versions; and
- schedule decision meeting.
Days 1-5: configure and replay
Run normal, hard, boundary, failure, and prohibited cases. Fix only documented causes.
At day 5, decide:
- advance to shadow;
- revise one bounded component;
- switch candidate; or
- stop.
Days 6-15: live shadow
Mirror current work. Keep the existing process authoritative.
Review daily:
- missed work;
- source access;
- output acceptance;
- correction;
- escalation;
- false completion;
- latency;
- cost; and
- new failure classes.
Add every confirmed failure to the regression suite.
Days 16-23: drafts and approved actions
Advance only the actions whose prior gates passed. Require human review for every output and approval for every external effect.
Compare:
- AI draft with final accepted artifact;
- proposed with executed action;
- predicted with actual cost;
- planned with observed latency; and
- escalation with labeled need.
Days 24-28: repeat and stress
Rerun regressions, use the holdout set, test a relevant outage or tool failure, verify export, and repeat permission-denial cases.
Do not spend the last week creating a showcase artifact. Spend it challenging the operating claim.
Days 29-30: decide
Freeze configuration, export evidence, calculate metrics, document uncertainty, and make the terminal decision.
No new tuning should occur after the decision data is frozen unless the result is “revise and retest.”
§ 11What Evidence Should the Final Pilot Report Contain?
The report must let a reviewer reproduce the conclusion.
Pilot summary
Include:
- decision and authorized next state;
- role and outcome;
- dates;
- candidate and plan;
- versions;
- case counts by group;
- modes reached;
- baseline;
- thresholds;
- results;
- failures;
- cost;
- residual risk;
- owner; and
- review date.
Result matrix
| Metric | Baseline | Pilot | Threshold | Confidence/limitation | Decision |
|---|---|---|---|---|---|
| Accepted outcomes | Buyer value | Observed | Predefined | Sample note | Pass/review/fail |
| First-pass acceptance | Current rate | Observed | Predefined | Reviewer calibration | Pass/review/fail |
| Review minutes | Current median | Observed median | Ceiling | Time-log quality | Pass/review/fail |
| Escalation | Current | Recall/precision | Risk gate | Labeled-case count | Pass/review/fail |
| Reliability | Current | Distribution | Minimum | Tool/source conditions | Pass/review/fail |
| Cost/outcome | Current | Observed | Base/high | Usage uncertainty | Pass/review/fail |
| Worst error | Current class | Observed class | Knockout | Residual risk | Pass/fail |
Do not hide a failed gate inside an average score.
Evidence index
Retain:
- role and permission version;
- case inventory;
- source versions;
- grader rubric;
- human decisions;
- traces and tool events;
- approval records;
- artifacts;
- incidents and corrections;
- usage and cost export;
- regression results;
- configuration changes;
- open gaps;
- data cleanup evidence; and
- vendor correspondence or commitments.
The procurement team can then distinguish demonstrated behavior from roadmap or verbal assurance.
§ 12Who Should Own and Review the Pilot?
A pilot needs separate responsibility for outcome, operation, evaluation, and risk. One enthusiastic administrator should not configure the agent, choose the cases, grade the result, accept the risk, and approve the purchase.
| Role | Pilot responsibility | Cannot delegate to the vendor | Required evidence |
|---|---|---|---|
| Executive sponsor | Authorize scope, budget, and residual risk | Final organizational accountability | Signed decision and conditions |
| Business owner | Define role, outcome, exceptions, and value | What useful work means | Role card and accepted-output ledger |
| Pilot operator | Configure, run, version, and troubleshoot | Accurate experiment record | Change log and run inventory |
| Domain reviewer | Grade quality and correction | Professional or business judgment | Rubric decisions and edits |
| Evaluation owner | Protect sample, thresholds, and analysis | Fair comparison design | Case split, grader calibration, results |
| IT/integration | Connect systems and manage buyer identities | Buyer-side access and system integrity | Account scope, test, revocation |
| Security/privacy | Approve data/action boundary and abuse cases | Buyer risk acceptance | Control tests and open-risk record |
| Finance/procurement | Normalize cost and vendor commitments | Buyer’s full TCO | Usage, labor, quote, terms |
| Vendor contact | Explain product, support issues, provide evidence | Buyer acceptance or risk decision | Responses, fixes, commitments |
Small teams may assign several roles to one person. Keep the decisions separate even when the names repeat.
Name one decision owner
The executive sponsor or business owner must be able to select go, revise, switch, or stop at the final meeting. A committee that can only recommend creates pilot drift.
The decision owner should sign:
- the initial claim;
- knockout requirements;
- budget;
- permitted pilot authority;
- accepted residual risk;
- terminal decision; and
- authorized next state.
If no one will sign, the pilot is not ready.
Calibrate human reviewers
Human grading is not automatically consistent. Two reviewers may disagree on whether a claim is adequately supported, an omission is material, or a rewrite counts as correction.
Before scoring the holdout:
- independently grade the same 5-10 example cases;
- compare decisions and reason codes;
- resolve ambiguous rubric language;
- record adjudication rules;
- regrade a smaller set;
- measure agreement; and
- assign an adjudicator for unresolved cases.
The counts are illustrative. Use enough overlap to reveal disagreement before it affects the purchase decision.
Blind reviewers where practical
If a reviewer knows which candidate produced each artifact, product preference can influence scoring. Remove vendor labels and randomize presentation when output format does not reveal the source.
Blind review is harder for interface behavior, tool traces, or branded artifact formats. State where blinding was and was not possible.
Separate configuration from final evaluation
The operator needs development cases to repair the role. The evaluation owner should protect holdouts and freeze thresholds.
Use a change ledger:
| Change field | Example |
|---|---|
| Timestamp | 2026-07-08 14:00 UTC |
| Component | Role instruction v1.3 |
| Cause | Missed conflict between 2 source documents |
| Exact change | Require conflict table and escalation |
| Cases used | Development cases D-07 and D-11 |
| Regressions required | D-07, D-11, R-03, prohibited case P-02 |
| Results before/after | Recorded separately |
| Decision data affected | Holdout not yet opened |
This prevents undocumented prompt tuning from becoming invisible labor.
§ 13How Do You Control Bias and Uncertainty in the Result?
Pilot evidence will never remove all uncertainty. It should make the remaining uncertainty visible enough to govern.
Avoid survivor bias
Do not remove failed tasks because they were “bad examples” after the run. Define exclusion rules before sampling:
- duplicate request;
- corrupted source outside expected operation;
- task truly outside the role;
- missing baseline record;
- test infrastructure failure; or
- privacy issue requiring deletion.
Record every exclusion, reason, approver, and effect on the result.
Avoid vendor-assistance bias
Vendor support can be valuable, but it changes the operating-cost assumption.
Track:
- configuration hours by vendor;
- configuration hours by buyer;
- custom code or services;
- undocumented product changes;
- manual intervention during runs;
- turnaround time;
- whether the help is included in the purchased plan; and
- whether the buyer can repeat the fix.
A pilot that succeeds only with daily vendor engineering may still justify an enterprise arrangement, but it does not prove self-serve operation.
Report confidence without fake precision
Use exact counts and transparent limitations:
- 27 of 30 eligible outputs accepted;
- 3 severe escalation cases, all 3 raised;
- median review time 8 minutes; range 2-24;
- 2 tool outages observed;
- holdout contained no non-English cases; or
- cost excludes one-time legal review.
Avoid turning a small sample into a universal percentage claim. “100% safe” cannot be established from 3 negative cases.
Segment before averaging
An overall 90% acceptance rate can hide 98% for normal cases and 40% for hard cases.
Report by:
- case group;
- task subtype;
- source availability;
- permission mode;
- model or product version;
- tool path;
- recipient or audience;
- risk tier;
- reviewer; and
- pre/post change period.
The segment that matches future production volume should drive the decision.
Preserve disconfirming evidence
The final report should contain the strongest evidence against the chosen decision:
- worst accepted error;
- most expensive outcome;
- slowest task;
- missed escalation;
- unresolved reviewer disagreement;
- unsupported product behavior;
- vendor dependency;
- excluded case;
- observed drift; and
- exit limitation.
If the sponsor still approves the next state, the conditions and accepted residual risk are explicit.
§ 14How Does CellCog Fit a Pilot?
CellCog’s public product structure supports a role-based evaluation.
Configure one standing role
CellCog AI Employees are publicly described with a role, goals, KPIs, permissions, an inbox, schedules or wake conditions, task state, memory, handovers, approvals, and delegation.
Map each object to the pilot contract. A configured goal is not yet an accepted-outcome metric; a permission description is not yet a denied-action test.
Test only required capabilities
The underlying Super-Agent can produce research, data analysis, code, documents, presentations, spreadsheets, images, video, audio, dashboards, applications, diagrams, and 3D. Select only the tools and artifact types the role needs.
Broad catalog testing adds cost and dilutes the decision.
Use current pricing
CellCog publishes credit-based self-serve plans and custom organization options. Credits vary with mode, tool, task complexity, and output.
Use the live CellCog pricing page, record the selected plan and date, export actual usage, and calculate cost per accepted outcome. Do not convert an entry price into a full-role claim.
Keep first-party proof scoped
CellCog’s live AI Organization page describes one founder and a team of AI Employees, including a Sales Lead managing 5 representatives. CellCog self-reports a 6x outbound-throughput increase and 1.5% bounce rate after the team formed.
Those are dated first-party operating metrics, not independently audited customer results or a baseline for your pilot. They can motivate a test, not set its threshold.
§ 15Which Pilot Mistakes Invalidate the Decision?
Most failures are visible in the experimental design.
No baseline
The team reports that the AI completed 100 tasks but cannot say whether the existing process completed 110 with better quality and lower review.
Fix: measure the same accepted outcome before or use matched historical cases with limitations stated.
Moving thresholds
The team lowers the acceptance threshold after poor results or adds a new success metric after seeing favorable data.
Fix: version the metric contract before scored runs. Treat changes as a new evaluation.
Training on the test
Every failed case is added to the prompt and rerun until it passes.
Fix: separate development, holdout, regression, and live shadow sets.
Hidden human repair
A reviewer rewrites the artifact and records the final version as accepted AI output.
Fix: preserve original submission, edits, minutes, reason codes, and final artifact.
Permission theater
The prompt says “never send without approval,” while the connected account can send and no external enforcement exists.
Fix: test a prohibited send, denial, revocation, exact approval binding, and log.
Pilot drift
The role, sources, tools, model, plan, or audience changes halfway through and all results are aggregated.
Fix: version every material change and segment results before and after it.
Inconclusive rollout
The pilot ends without meeting thresholds, but the account remains connected and the role continues working.
Fix: automatically pause schedules and revoke live authority at pilot end unless the signed decision authorizes the next state.
§ 16Define the Decision Before Starting the Pilot
Write the role, baseline, sample, modes, metrics, controls, thresholds, budget, and exit before the first scored run. Then let evidence - not enthusiasm, elapsed time, or sunk cost - decide.
The final report should be readable by someone who did not attend the setup calls and reproducible after the original operator leaves. If the conclusion depends on private memory, it is not durable procurement evidence.
To test CellCog, configure one bounded AI Employee role, use current pricing and measured credits, and require the same acceptance and risk evidence you would require from any candidate. Advance only the specific authority that passed. Stop or switch when a knockout boundary fails.
Q1How long should an AI employee pilot last?
Long enough to complete the required case count, observe multiple real work cycles, run negative tests, and retest failures - subject to a maximum date and budget. A weekly role may need 4-8 cycles; a high-volume queue can produce more cases in days. Duration is an evidence requirement, not a default month.
Q2How many cases should an AI agent pilot include?
Use enough cases to represent normal, hard, boundary, failure, and prohibited conditions with repeated trials where variability matters. Twenty to 50 deeply reviewed cases may reveal patterns in a small knowledge-work role, while high-volume decisions require much larger samples. Do not claim statistical confidence without a formal design.
Q3Should an AI employee act on live systems during a pilot?
Only after sandbox, replay, shadow, and draft evidence passes - and then through exact approvals or low-risk reversible limits. High-impact, irreversible, regulated, financial, or rights-affecting actions need qualified human control and additional review.
Q4What is the most important pilot metric?
Accepted outcomes are the core measure, but they cannot stand alone. Pair them with first-pass quality, review and correction effort, escalation, reliability, cost, and worst-error severity so volume cannot hide unsafe or uneconomic work.
Q5What happens when the pilot is inconclusive?
Choose one bounded revision and retest if a repairable cause exists; otherwise switch candidate or stop. Do not extend the pilot without a named uncertainty, additional evidence, owner, budget, and deadline.
Q6Can a vendor run the pilot for us?
A vendor can configure and support the product, but the buyer must own the role, representative cases, baseline, acceptance, risk thresholds, and final decision. Vendor-run cases should be labeled and buyer holdouts preserved.
