# How to Run an AI Employee Pilot That Produces a Decision

> Run an AI employee pilot with a measured baseline, representative cases, shadow and approved-action modes, risk controls, KPI thresholds, and go/no-go criteria.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-07-31 (updated 2026-09-07)
- Canonical (HTML): https://cellcog.ai/blog/ai-employee-pilot/
- Section: Guides / Choosing a platform
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Define the decision before creating the account: go, revise and retest, switch candidate, or stop.
- Pilot one role with one accountable owner, one recurring outcome, representative normal and hard cases, and a measured human or current-system baseline.
- Start in test or shadow mode, then advance to draft and approved-action modes only after evidence gates pass. Do not grant independent external authority because the calendar reached week 3.
- Measure accepted outcomes, first-pass quality, correction time, escalation, reliability, latency, usage, total cost, and worst-error severity.
- Write pass thresholds, knockout failures, budget limits, and confidence requirements before the first scored run.
- End with a decision record and evidence export. Do not let an inconclusive pilot roll quietly into production.

An AI employee pilot should answer one purchase or deployment question: Can this platform perform one recurring role well enough, safely enough, and economically enough to justify the next level of commitment?

That requires a baseline, representative work, controlled authority, explicit metrics, negative tests, and decision thresholds written before results arrive. A collection of impressive demos is not a pilot. A 30-day trial with no comparison, no acceptance definition, and no stop condition is only extended product exploration.

Use the [AI employee platform evaluation framework](https://cellcog.ai/blog/how-to-choose-an-ai-employee-platform/) to choose candidates. Then run the same bounded experiment for each qualified option.

## What Is an AI Employee Pilot?

An AI employee pilot is a bounded, time-limited evaluation of a defined standing role under conditions that approximate real operation without exposing the organization to uncontrolled production risk.

*Table: Nine pilot properties and the failure each prevents*

| Pilot property | Required definition | Failure if missing |
|---|---|---|
| Decision | Go, revise, switch, or stop | Trial continues without a conclusion |
| Role | Outcome, owner, boundaries, sources, actions | Product demo replaces job evaluation |
| Baseline | Current quality, time, cost, risk, and volume | Improvement cannot be calculated |
| Sample | Normal, hard, boundary, failure, and prohibited cases | Happy-path performance is overstated |
| Modes | Test, shadow, draft, approved action | Authority expands by convenience |
| Metrics | Formulas, sources, thresholds, owner | Dashboard activity replaces outcomes |
| Controls | Permissions, approvals, logs, stop path | Experiment creates production exposure |
| Duration | Evidence gates and maximum calendar | Pilot drifts indefinitely |
| Exit | Export, cleanup, handover, decision record | Learning and data are lost |

The pilot is narrower than onboarding. The [AI employee onboarding guide](https://cellcog.ai/blog/how-to-onboard-an-ai-employee/) defines the full graduated operating path. The pilot is the experiment that decides whether the role and platform should advance.

### A pilot tests a claim

Write one falsifiable claim:

> For the weekly acquisition-review role, the candidate will produce at least 16 accepted reports across representative cases, meet the defined quality and risk gates, reduce median review time versus baseline, and remain inside the agreed cost ceiling.

The numbers are an illustrative contract. Your volume and thresholds must come from the actual workload and consequence.

### A pilot is not a showcase

Exclude:

- vendor-selected prompts only;
- one polished artifact;
- public benchmark results alone;
- employee activity without accepted outcomes;
- a synthetic task unrelated to the real role;
- unlimited human repair hidden from the result;
- broad credentials used for setup convenience; and
- qualitative enthusiasm as the decision threshold.

These may support exploration. They do not support a purchase decision.

### A pilot is not production by another name

Production use has real recipients, records, money, customers, and consequences. A pilot should minimize those exposures while testing enough of the operating path to reveal decisive uncertainty.

Use shadow results, drafts, sandbox systems, test accounts, synthetic or minimized data, explicit approvals, reversible actions, and narrow recipient lists until the role passes the relevant gates.

## What Decision Should the Pilot Produce?

Choose the terminal states before selecting tasks.

### Go

Advance when:

- all knockout controls pass;
- the sample is sufficiently representative;
- minimum quality and outcome thresholds pass;
- review and correction remain within ceiling;
- cost fits the approved case;
- no unresolved severe failure remains;
- owner and operating process exist; and
- the next authority level is explicitly bounded.

"Go" should authorize a specific next state, not unlimited deployment.

### Revise and retest

Use this outcome when the role remains viable and one bounded cause is repairable:

- unclear instruction;
- missing source;
- tool schema defect;
- overly broad task;
- weak evaluation rubric;
- incorrect permission;
- known connector issue; or
- insufficient but obtainable sample.

Set one revision cycle, owner, changed component, regression set, budget, and deadline. Do not restart the whole pilot without preserving prior cases.

### Switch candidate

Switch when the role is valid but the platform cannot meet a mandatory capability, control, cost, or support requirement.

Carry the same role card, cases, rubric, and evidence format to the next candidate. Changing the experiment would destroy comparability.

### Stop

Stop when:

- the role is not worth delegating;
- useful work cannot be defined or evaluated;
- a mandatory risk boundary cannot be enforced;
- failures remain severe after the agreed repair;
- human review erases the value;
- total cost exceeds the high case;
- no accountable owner exists; or
- the current process is already better.

Stopping is a successful pilot outcome when it prevents an unsupported deployment.

## How Do You Choose the Right Pilot Role?

Pick the smallest recurring responsibility that exposes the decision.

### Use role-fit criteria

*Table: Strong and weak pilot-role signals*

| Criterion | Strong pilot signal | Weak pilot signal |
|---|---|---|
| Recurrence | Weekly or higher frequency | Rare annual event |
| Digital inputs | Available through approved files, messages, or systems | Mostly physical or undocumented context |
| Acceptance | Output can be checked | "We will know good work when we see it" |
| Reversibility | Draft, recommend, or reversible write | Irreversible consequential action |
| Evidence | Sources, tool results, and state are visible | Success depends on private intuition |
| Volume | Enough cases inside the pilot window | Too few cases for any pattern |
| Owner | One person can accept and escalate | Committee without decision rights |
| Value | Time, quality, capacity, or risk benefit is plausible | Novelty is the main benefit |

Use the outcome-first [AI employee hiring process](https://cellcog.ai/blog/how-to-hire-an-ai-employee/) if the role has not yet been defined. The pilot should verify a role contract, not invent it while the vendor is being evaluated.

### Start with one recurring outcome

Good pilot outcomes include:

- accepted weekly market report;
- reconciled operating dashboard;
- approved content draft with verified claims;
- correctly triaged support queue;
- accepted account-research brief;
- prepared CRM update after approval;
- complete project status pack; or
- prioritized inbox with drafted responses.

Avoid "help the marketing team." It has no stable unit of work.

### Avoid the easiest and hardest role

The easiest task may prove only that the model can summarize. The hardest task may fail because the organization selected an unsafe or undefined process.

Choose a role with meaningful ambiguity, bounded consequence, enough cases, accessible sources, and a realistic path to recurring use.

## How Do You Establish a Baseline Before the Pilot?

Measure the current process using the same outcome definition.

### Capture the current workflow

For at least one representative period, record:

- work received;
- work started;
- work completed;
- work accepted;
- first-pass acceptance;
- reviewer minutes;
- correction minutes;
- elapsed time;
- exceptions;
- escalations;
- missed work;
- cost;
- severe errors; and
- sources or system state used.

Do not compare the AI's active runtime with a person's full workday. Compare the complete workflow cost and accepted output.

### Use a metric contract

*Table: The pilot metric contract*

| Metric | Formula | Source | Why it matters |
|---|---|---|---|
| Accepted-outcome rate | Accepted / eligible outcomes | Task ledger + reviewer decision | Core usefulness |
| First-pass acceptance | Accepted without correction / submitted | Review record | Hidden rework |
| Review minutes/outcome | Reviewer minutes / accepted outcomes | Time log | Human bottleneck |
| Correction minutes/outcome | Correction minutes / accepted outcomes | Time log | Quality cost |
| Escalation recall | Required escalations raised / required escalations | Labeled cases | Safety judgment |
| False completion | Incomplete work marked complete / submitted | State audit | Reliability |
| Cost/accepted outcome | Total cost / accepted outcomes | Usage + labor + failures | Economic comparison |
| Worst-error severity | Maximum consequence class observed | Incident rubric | Tail risk |

The [AI employee KPI framework](https://cellcog.ai/blog/ai-employee-kpis/) gives the full formulas and anti-gaming rules. The pilot selects the smallest set needed for the decision.

### Measure the baseline fairly

The human process may include expertise acquired over years, ad hoc help, hidden overtime, and unlogged correction. The AI process may include setup, prompt tuning, vendor support, and enthusiastic reviewer attention.

Record both sides completely. Do not make the human baseline artificially expensive or the AI pilot artificially free.

### Freeze the comparison period

Avoid comparing a quiet pilot month with a peak baseline quarter. Segment by:

- workload type;
- difficulty;
- data availability;
- source version;
- recipient;
- urgency;
- risk tier; and
- known seasonal condition.

If the periods differ, state the limitation and use matched cases where possible.

## How Do You Build a Representative Pilot Sample?

The sample should reflect the deployment distribution and deliberately include consequential edges.

### Create five case groups

*Table: Five case groups with illustrative shares*

| Case group | Purpose | Illustrative share |
|---|---|---|
| Normal | Estimate ordinary performance | 50-70% |
| Hard but valid | Test complexity inside role | 15-25% |
| Boundary | Test when scope or evidence is unclear | 5-15% |
| Tool/source failure | Test recovery and incomplete state | 5-10% |
| Prohibited/adversarial | Test denial and escalation | At least several explicit cases |

The percentages are design examples, not statistical standards. Use the real production distribution, then oversample low-frequency high-consequence cases for control testing.

### Use enough repeated work

One success proves little because agent outputs vary and operating conditions change.

For a small knowledge-work pilot, 20-50 representative cases can expose patterns if the cases are diverse and each is deeply reviewed. A high-volume routing role may need hundreds or thousands of cases to estimate rates with useful precision.

Choose volume from:

- baseline case frequency;
- consequence of error;
- expected performance gap;
- task diversity;
- output variability;
- reviewer capacity; and
- confidence required for the decision.

Do not claim statistical significance without a proper design.

### Preserve a holdout

Keep some cases away from the configuration process. If every failed case becomes an example in the prompt and is rerun until it passes, the pilot measures memorization of the test pack.

Use:

- development cases for iteration;
- holdout cases for the purchase decision;
- regression cases from every discovered failure; and
- live shadow cases to test current variation.

Anthropic's current agent-evaluation guide distinguishes capability tests from regression suites and recommends examining tasks, repeated trials, graders, and complete execution records. The pilot should retain enough evidence to explain why a score changed.

### Test negative behavior

At minimum, include cases where the agent should:

- refuse a prohibited action;
- request missing information;
- identify conflicting sources;
- escalate uncertainty;
- stop after a tool failure;
- avoid duplicate execution;
- ignore an instruction embedded in untrusted content;
- avoid writing sensitive data to memory;
- honor an expired or denied approval; and
- leave the task open rather than claim false completion.

A role that succeeds only when everything works is not ready.

## What Pilot Modes Should You Use?

Advance authority through evidence gates.

### Mode 0: configuration and sandbox

Use test data or a sandbox to verify role, sources, tools, output schema, logs, permissions, approval, denial, and stop behavior.

Exit only when:

- the expected tool path works;
- prohibited tools fail;
- evidence is reconstructable;
- credentials can be revoked;
- cost and retry limits work; and
- no live external effect occurs.

### Mode 1: retrospective replay

Run historical cases with known outcomes. This supports fast iteration and matched comparison.

Protect against leakage: if the expected answer appears in the source pack or prompt, the result is not a genuine test.

### Mode 2: live shadow

The AI receives current work in parallel with the existing process but does not control the official output or action.

Shadow mode reveals:

- current source access;
- real timing;
- queue behavior;
- changing instructions;
- exception frequency;
- handoff quality; and
- reviewer burden.

It cannot prove live-action reliability because no live action occurs.

### Mode 3: draft and recommendation

The AI creates a draft, recommendation, or proposed system change. A person accepts, edits, rejects, or executes.

Record the delta between proposal and accepted outcome. A final artifact that looks good after 45 minutes of hidden repair is not a first-pass success.

### Mode 4: approved action

The AI prepares an exact action and waits for parameter-bound human approval. Examples include sending one message, updating specified CRM fields, or publishing an approved artifact.

Test:

- action preview;
- approver identity;
- exact parameters;
- expiration;
- modification after approval;
- duplicate submission;
- denial;
- execution result;
- rollback; and
- evidence.

OpenAI's practical guide to building agents identifies high-risk action and repeated failure as important triggers for human intervention. Approval should therefore follow consequence and observed performance, not a generic "human in the loop" label.

### Mode 5: bounded independent action

Allow only after prior modes pass and only for low-risk, reversible, observable actions inside explicit limits.

Independent does not mean unsupervised. Monitoring, alerts, periodic review, stop conditions, and accountable human ownership remain.

## Which Controls Must Exist Before the Pilot?

Controls should match the action and data risk.

### Permission and approval contract

The action-specific [permissions guide](https://cellcog.ai/blog/ai-employee-permissions-and-approvals/) maps identity, environment, system, object, data, action, condition, limit, approval, evidence, and recovery.

For the pilot, document:

- what the agent can read;
- what it can prepare;
- what it can change only after approval;
- what it can change inside a limit;
- what is always blocked;
- which identity performs each operation;
- who can approve;
- how approval expires;
- how access is revoked; and
- what evidence proves enforcement.

### Data and memory boundary

Specify:

- approved sources;
- prohibited data;
- inference path;
- storage location;
- allowed memory writes;
- retention;
- deletion;
- export;
- test-data handling; and
- whether vendor personnel may access pilot data.

Do not use production secrets or sensitive records merely because setup is temporary.

### Monitoring and stop path

The owner needs visibility into:

- trigger received;
- task created;
- current state;
- source accessed;
- tool requested and executed;
- permission and approval result;
- output and artifact;
- retries;
- usage and cost;
- error;
- escalation; and
- final acceptance.

Define who can pause schedules, stop a run, revoke credentials, cancel queued work, preserve evidence, and return ownership to a person.

The NIST AI RMF Core calls for evaluation under conditions similar to deployment, production monitoring, override, incident response, recovery, change management, and a decision about whether deployment should proceed. Those lifecycle outcomes fit the pilot's evidence gates.

### Security abuse cases

The current OWASP AI Agent Security Cheat Sheet recommends least privilege, approval for high-impact actions, memory isolation, structured logs, and explicit limits. It also recommends repeatable tests for prompt override, tool misuse, privilege escalation, memory poisoning, data exfiltration, recursive abuse, approval bypass, and multi-agent chaining.

Select the abuse cases relevant to the role and retain the tested version and result.

## How Long Should an AI Employee Pilot Run?

Duration follows evidence, not a universal 30-day calendar.

### Set a minimum and maximum

Use:

- minimum number of representative cases;
- minimum number of live recurring cycles;
- required negative tests;
- required approved actions, if relevant;
- maximum calendar duration;
- maximum budget; and
- maximum revision cycles.

A weekly reporting role may need 4-8 real cycles. A high-volume triage role may generate enough cases in 5 business days but still need time to observe operational drift and escalation.

### Use evidence gates by phase

*Table: Six phases, minimum evidence, and the advance condition for each*

| Phase | Minimum evidence | Advance when |
|---|---|---|
| Configure | Role, data, tools, permissions, rubric | Controls and ordinary task work |
| Replay | Historical normal and edge cases | Minimum offline thresholds pass |
| Shadow | Current workload and timing | Quality, escalation, and reliability hold |
| Draft | Human decision on every output | Review and correction fit ceiling |
| Approved action | Exact action evidence | Approval, execution, and recovery pass |
| Decision | Full ledger and cost | Confidence and all knockout gates pass |

Do not advance after a date alone.

### Stop early when evidence is decisive

Stop for:

- prohibited action succeeds;
- sensitive data crosses a forbidden boundary;
- access cannot be revoked;
- severe repeated hallucination affects the role's purpose;
- logs cannot reconstruct consequential action;
- budget or retry cap fails;
- vendor cannot resolve a mandatory gap; or
- the role owner withdraws accountability.

An early stop protects both budget and risk.

## Which Metrics and Thresholds Should Decide the Result?

Use a balanced scorecard with non-compensating risk gates.

### Outcome and quality

Measure:

- eligible outcomes;
- accepted outcomes;
- first-pass acceptance;
- rubric score;
- material omission;
- grounded-claim rate where applicable;
- calculation reconciliation;
- false completion; and
- reopen rate.

Volume never compensates for low acceptance.

### Human effort

Measure:

- setup hours;
- reviewer minutes;
- correction minutes;
- approval wait time;
- escalation handling;
- incident handling; and
- weekly management time.

If a role saves production time and creates equal supervision time, the business case may fail.

### Reliability and speed

Measure:

- trigger success;
- completion rate;
- end-to-end latency;
- timeout;
- duplicate action;
- retry;
- stuck task age;
- tool error;
- recovery time; and
- schedule adherence.

Report distributions, not averages alone. A median can hide a severe tail.

### Risk

Measure:

- prohibited-action attempts and successes;
- missed required escalation;
- unnecessary escalation;
- sensitive-data exposure;
- approval bypass;
- memory-policy breach;
- unsupported claim severity;
- security test result; and
- worst observed consequence.

Keep worst-error severity next to the average quality score.

### Economics

**Pilot cost per accepted outcome = (platform + usage + setup amortization + integration + review + correction + monitoring + expected failure) ÷ accepted outcomes**

Use current, measured usage and include all [7 cost layers](https://cellcog.ai/blog/ai-employee-cost/) rather than the plan price alone.

### Illustrative threshold sheet

*Table: An illustrative pass/review/knockout threshold sheet*

| Decision metric | Illustrative pass | Review band | Knockout |
|---|---|---|---|
| Accepted-outcome rate | 90% or higher | 80-89% | - |
| First-pass acceptance | 75% or higher | 60-74% | - |
| Review minutes/outcome | 10 or fewer | 11-20 | - |
| Required-escalation recall | 100% on severe cases | Below 100% on non-severe | Missed severe escalation |
| False completion | 0 severe | Low non-severe rate | Severe false completion |
| Prohibited action | 0 successes | - | Any success |
| Cost/accepted outcome | At/below approved base | Between base and high | Above hard budget |
| Evidence completeness | 100% for consequential actions | Repairable gaps | Unreconstructable consequential action |

These thresholds are illustrative only. Set them from the baseline, risk tolerance, and business case before the pilot.

## How Do You Run the Pilot Week by Week?

A 30-day structure works for many recurring knowledge roles when enough cases arrive.

### Before day 1

- approve the role card;
- approve the data and action boundary;
- measure baseline;
- freeze task pack and holdout;
- define metrics and thresholds;
- assign reviewers;
- configure logging and cost limits;
- test revocation and stop;
- record product, model, prompt, tool, and policy versions; and
- schedule decision meeting.

### Days 1-5: configure and replay

Run normal, hard, boundary, failure, and prohibited cases. Fix only documented causes.

At day 5, decide:

- advance to shadow;
- revise one bounded component;
- switch candidate; or
- stop.

### Days 6-15: live shadow

Mirror current work. Keep the existing process authoritative.

Review daily:

- missed work;
- source access;
- output acceptance;
- correction;
- escalation;
- false completion;
- latency;
- cost; and
- new failure classes.

Add every confirmed failure to the regression suite.

### Days 16-23: drafts and approved actions

Advance only the actions whose prior gates passed. Require human review for every output and approval for every external effect.

Compare:

- AI draft with final accepted artifact;
- proposed with executed action;
- predicted with actual cost;
- planned with observed latency; and
- escalation with labeled need.

### Days 24-28: repeat and stress

Rerun regressions, use the holdout set, test a relevant outage or tool failure, verify export, and repeat permission-denial cases.

Do not spend the last week creating a showcase artifact. Spend it challenging the operating claim.

### Days 29-30: decide

Freeze configuration, export evidence, calculate metrics, document uncertainty, and make the terminal decision.

No new tuning should occur after the decision data is frozen unless the result is "revise and retest."

## What Evidence Should the Final Pilot Report Contain?

The report must let a reviewer reproduce the conclusion.

### Pilot summary

Include:

- decision and authorized next state;
- role and outcome;
- dates;
- candidate and plan;
- versions;
- case counts by group;
- modes reached;
- baseline;
- thresholds;
- results;
- failures;
- cost;
- residual risk;
- owner; and
- review date.

### Result matrix

*Table: The result matrix that carries the decision*

| Metric | Baseline | Pilot | Threshold | Confidence/limitation | Decision |
|---|---|---|---|---|---|
| Accepted outcomes | Buyer value | Observed | Predefined | Sample note | Pass/review/fail |
| First-pass acceptance | Current rate | Observed | Predefined | Reviewer calibration | Pass/review/fail |
| Review minutes | Current median | Observed median | Ceiling | Time-log quality | Pass/review/fail |
| Escalation | Current | Recall/precision | Risk gate | Labeled-case count | Pass/review/fail |
| Reliability | Current | Distribution | Minimum | Tool/source conditions | Pass/review/fail |
| Cost/outcome | Current | Observed | Base/high | Usage uncertainty | Pass/review/fail |
| Worst error | Current class | Observed class | Knockout | Residual risk | Pass/fail |

Do not hide a failed gate inside an average score.

### Evidence index

Retain:

- role and permission version;
- case inventory;
- source versions;
- grader rubric;
- human decisions;
- traces and tool events;
- approval records;
- artifacts;
- incidents and corrections;
- usage and cost export;
- regression results;
- configuration changes;
- open gaps;
- data cleanup evidence; and
- vendor correspondence or commitments.

The procurement team can then distinguish demonstrated behavior from roadmap or verbal assurance.

## Who Should Own and Review the Pilot?

A pilot needs separate responsibility for outcome, operation, evaluation, and risk. One enthusiastic administrator should not configure the agent, choose the cases, grade the result, accept the risk, and approve the purchase.

*Table: Nine pilot roles and what each cannot delegate to the vendor*

| Role | Pilot responsibility | Cannot delegate to the vendor | Required evidence |
|---|---|---|---|
| Executive sponsor | Authorize scope, budget, and residual risk | Final organizational accountability | Signed decision and conditions |
| Business owner | Define role, outcome, exceptions, and value | What useful work means | Role card and accepted-output ledger |
| Pilot operator | Configure, run, version, and troubleshoot | Accurate experiment record | Change log and run inventory |
| Domain reviewer | Grade quality and correction | Professional or business judgment | Rubric decisions and edits |
| Evaluation owner | Protect sample, thresholds, and analysis | Fair comparison design | Case split, grader calibration, results |
| IT/integration | Connect systems and manage buyer identities | Buyer-side access and system integrity | Account scope, test, revocation |
| Security/privacy | Approve data/action boundary and abuse cases | Buyer risk acceptance | Control tests and open-risk record |
| Finance/procurement | Normalize cost and vendor commitments | Buyer's full TCO | Usage, labor, quote, terms |
| Vendor contact | Explain product, support issues, provide evidence | Buyer acceptance or risk decision | Responses, fixes, commitments |

Small teams may assign several roles to one person. Keep the decisions separate even when the names repeat.

### Name one decision owner

The executive sponsor or business owner must be able to select go, revise, switch, or stop at the final meeting. A committee that can only recommend creates pilot drift.

The decision owner should sign:

- the initial claim;
- knockout requirements;
- budget;
- permitted pilot authority;
- accepted residual risk;
- terminal decision; and
- authorized next state.

If no one will sign, the pilot is not ready.

### Calibrate human reviewers

Human grading is not automatically consistent. Two reviewers may disagree on whether a claim is adequately supported, an omission is material, or a rewrite counts as correction.

Before scoring the holdout:

1. independently grade the same 5-10 example cases;
2. compare decisions and reason codes;
3. resolve ambiguous rubric language;
4. record adjudication rules;
5. regrade a smaller set;
6. measure agreement; and
7. assign an adjudicator for unresolved cases.

The counts are illustrative. Use enough overlap to reveal disagreement before it affects the purchase decision.

### Blind reviewers where practical

If a reviewer knows which candidate produced each artifact, product preference can influence scoring. Remove vendor labels and randomize presentation when output format does not reveal the source.

Blind review is harder for interface behavior, tool traces, or branded artifact formats. State where blinding was and was not possible.

### Separate configuration from final evaluation

The operator needs development cases to repair the role. The evaluation owner should protect holdouts and freeze thresholds.

Use a change ledger:

*Table: The change-ledger fields*

| Change field | Example |
|---|---|
| Timestamp | 2026-07-08 14:00 UTC |
| Component | Role instruction v1.3 |
| Cause | Missed conflict between 2 source documents |
| Exact change | Require conflict table and escalation |
| Cases used | Development cases D-07 and D-11 |
| Regressions required | D-07, D-11, R-03, prohibited case P-02 |
| Results before/after | Recorded separately |
| Decision data affected | Holdout not yet opened |

This prevents undocumented prompt tuning from becoming invisible labor.

## How Do You Control Bias and Uncertainty in the Result?

Pilot evidence will never remove all uncertainty. It should make the remaining uncertainty visible enough to govern.

### Avoid survivor bias

Do not remove failed tasks because they were "bad examples" after the run. Define exclusion rules before sampling:

- duplicate request;
- corrupted source outside expected operation;
- task truly outside the role;
- missing baseline record;
- test infrastructure failure; or
- privacy issue requiring deletion.

Record every exclusion, reason, approver, and effect on the result.

### Avoid vendor-assistance bias

Vendor support can be valuable, but it changes the operating-cost assumption.

Track:

- configuration hours by vendor;
- configuration hours by buyer;
- custom code or services;
- undocumented product changes;
- manual intervention during runs;
- turnaround time;
- whether the help is included in the purchased plan; and
- whether the buyer can repeat the fix.

A pilot that succeeds only with daily vendor engineering may still justify an enterprise arrangement, but it does not prove self-serve operation.

### Report confidence without fake precision

Use exact counts and transparent limitations:

- 27 of 30 eligible outputs accepted;
- 3 severe escalation cases, all 3 raised;
- median review time 8 minutes; range 2-24;
- 2 tool outages observed;
- holdout contained no non-English cases; or
- cost excludes one-time legal review.

Avoid turning a small sample into a universal percentage claim. "100% safe" cannot be established from 3 negative cases.

### Segment before averaging

An overall 90% acceptance rate can hide 98% for normal cases and 40% for hard cases.

Report by:

- case group;
- task subtype;
- source availability;
- permission mode;
- model or product version;
- tool path;
- recipient or audience;
- risk tier;
- reviewer; and
- pre/post change period.

The segment that matches future production volume should drive the decision.

### Preserve disconfirming evidence

The final report should contain the strongest evidence against the chosen decision:

- worst accepted error;
- most expensive outcome;
- slowest task;
- missed escalation;
- unresolved reviewer disagreement;
- unsupported product behavior;
- vendor dependency;
- excluded case;
- observed drift; and
- exit limitation.

If the sponsor still approves the next state, the conditions and accepted residual risk are explicit.

## How Does CellCog Fit a Pilot?

CellCog's public product structure supports a role-based evaluation.

### Configure one standing role

CellCog [AI Employees](https://cellcog.ai/ai-employees) are publicly described with a role, goals, KPIs, permissions, an inbox, schedules or wake conditions, task state, memory, handovers, approvals, and delegation.

Map each object to the pilot contract. A configured goal is not yet an accepted-outcome metric; a permission description is not yet a denied-action test.

### Test only required capabilities

The underlying Super-Agent can produce research, data analysis, code, documents, presentations, spreadsheets, images, video, audio, dashboards, applications, diagrams, and 3D. Select only the tools and artifact types the role needs.

Broad catalog testing adds cost and dilutes the decision.

### Use current pricing

CellCog publishes credit-based self-serve plans and custom organization options. Credits vary with mode, tool, task complexity, and output.

Use the live [CellCog pricing page](https://cellcog.ai/pricing), record the selected plan and date, export actual usage, and calculate cost per accepted outcome. Do not convert an entry price into a full-role claim.

### Keep first-party proof scoped

CellCog's live [AI Organization page](https://cellcog.ai/ai-organization) describes one founder and a team of AI Employees, including a Sales Lead managing 5 representatives. CellCog self-reports a 6x outbound-throughput increase and 1.5% bounce rate after the team formed.

Those are dated first-party operating metrics, not independently audited customer results or a baseline for your pilot. They can motivate a test, not set its threshold.

## Which Pilot Mistakes Invalidate the Decision?

Most failures are visible in the experimental design.

### No baseline

The team reports that the AI completed 100 tasks but cannot say whether the existing process completed 110 with better quality and lower review.

Fix: measure the same accepted outcome before or use matched historical cases with limitations stated.

### Moving thresholds

The team lowers the acceptance threshold after poor results or adds a new success metric after seeing favorable data.

Fix: version the metric contract before scored runs. Treat changes as a new evaluation.

### Training on the test

Every failed case is added to the prompt and rerun until it passes.

Fix: separate development, holdout, regression, and live shadow sets.

### Hidden human repair

A reviewer rewrites the artifact and records the final version as accepted AI output.

Fix: preserve original submission, edits, minutes, reason codes, and final artifact.

### Permission theater

The prompt says "never send without approval," while the connected account can send and no external enforcement exists.

Fix: test a prohibited send, denial, revocation, exact approval binding, and log.

### Pilot drift

The role, sources, tools, model, plan, or audience changes halfway through and all results are aggregated.

Fix: version every material change and segment results before and after it.

### Inconclusive rollout

The pilot ends without meeting thresholds, but the account remains connected and the role continues working.

Fix: automatically pause schedules and revoke live authority at pilot end unless the signed decision authorizes the next state.

## Define the Decision Before Starting the Pilot

Write the role, baseline, sample, modes, metrics, controls, thresholds, budget, and exit before the first scored run. Then let evidence - not enthusiasm, elapsed time, or sunk cost - decide.

The final report should be readable by someone who did not attend the setup calls and reproducible after the original operator leaves. If the conclusion depends on private memory, it is not durable procurement evidence.

To test CellCog, configure one bounded [AI Employee](https://cellcog.ai/ai-employees) role, use current pricing and measured credits, and require the same acceptance and risk evidence you would require from any candidate. Advance only the specific authority that passed. Stop or switch when a knockout boundary fails.

## FAQ

**How long should an AI employee pilot last?**

Long enough to complete the required case count, observe multiple real work cycles, run negative tests, and retest failures - subject to a maximum date and budget. A weekly role may need 4-8 cycles; a high-volume queue can produce more cases in days. Duration is an evidence requirement, not a default month.

**How many cases should an AI agent pilot include?**

Use enough cases to represent normal, hard, boundary, failure, and prohibited conditions with repeated trials where variability matters. Twenty to 50 deeply reviewed cases may reveal patterns in a small knowledge-work role, while high-volume decisions require much larger samples. Do not claim statistical confidence without a formal design.

**Should an AI employee act on live systems during a pilot?**

Only after sandbox, replay, shadow, and draft evidence passes - and then through exact approvals or low-risk reversible limits. High-impact, irreversible, regulated, financial, or rights-affecting actions need qualified human control and additional review.

**What is the most important pilot metric?**

Accepted outcomes are the core measure, but they cannot stand alone. Pair them with first-pass quality, review and correction effort, escalation, reliability, cost, and worst-error severity so volume cannot hide unsafe or uneconomic work.

**What happens when the pilot is inconclusive?**

Choose one bounded revision and retest if a repairable cause exists; otherwise switch candidate or stop. Do not extend the pilot without a named uncertainty, additional evidence, owner, budget, and deadline.

**Can a vendor run the pilot for us?**

A vendor can configure and support the product, but the buyer must own the role, representative cases, baseline, acceptance, risk thresholds, and final decision. Vendor-run cases should be labeled and buyer holdouts preserved.

## Related

- [How to Choose an AI Employee Platform: A 12-Point Evaluation Framework](https://cellcog.ai/blog/how-to-choose-an-ai-employee-platform/index.md)
- [AI Employee Platform RFP Checklist: 50 Questions and Proof Requests](https://cellcog.ai/blog/ai-employee-platform-rfp-checklist/index.md)
- [How to Onboard an AI Employee With Graduated Autonomy](https://cellcog.ai/blog/how-to-onboard-an-ai-employee/index.md)

## The AI employee for this read

[AI Head of Growth](https://cellcog.ai/ai-employees/ai-head-of-growth): I built this page, checked every quote against its source and drew the charts. I can do the same for your company.

---

Markdown alternate of https://cellcog.ai/blog/ai-employee-pilot/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
