Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentPricingBlogStoryContact

How to Run an AI Employee Pilot That Produces a Decision

Napkin-style sketch of a laboratory flask on a pedestal feeding a four-way decision signpost labeled go, revise, switch, stop, with an amber highlight on the go arrow
Fig 0Define the decision before creating the account: go, revise and retest, switch candidate, or stop.

An AI employee pilot should answer one purchase or deployment question: Can this platform perform one recurring role well enough, safely enough, and economically enough to justify the next level of commitment?

That requires a baseline, representative work, controlled authority, explicit metrics, negative tests, and decision thresholds written before results arrive. A collection of impressive demos is not a pilot. A 30-day trial with no comparison, no acceptance definition, and no stop condition is only extended product exploration.

Use the AI employee platform evaluation framework to choose candidates. Then run the same bounded experiment for each qualified option.

On this page · 16 sectionsOpen
  1. What Is an AI Employee Pilot?
  2. What Decision Should the Pilot Produce?
  3. How Do You Choose the Right Pilot Role?
  4. How Do You Establish a Baseline Before the Pilot?
  5. How Do You Build a Representative Pilot Sample?
  6. What Pilot Modes Should You Use?
  7. Which Controls Must Exist Before the Pilot?
  8. How Long Should an AI Employee Pilot Run?
  9. Which Metrics and Thresholds Should Decide the Result?
  10. How Do You Run the Pilot Week by Week?
  11. What Evidence Should the Final Pilot Report Contain?
  12. Who Should Own and Review the Pilot?
  13. How Do You Control Bias and Uncertainty in the Result?
  14. How Does CellCog Fit a Pilot?
  15. Which Pilot Mistakes Invalidate the Decision?
  16. Define the Decision Before Starting the Pilot
Key points6 · 22 min full read
  1. Define the decision before creating the account: go, revise and retest, switch candidate, or stop.
  2. Pilot one role with one accountable owner, one recurring outcome, representative normal and hard cases, and a measured human or current-system baseline.
  3. Start in test or shadow mode, then advance to draft and approved-action modes only after evidence gates pass. Do not grant independent external authority because the calendar reached week 3.
  4. Measure accepted outcomes, first-pass quality, correction time, escalation, reliability, latency, usage, total cost, and worst-error severity.
  5. Write pass thresholds, knockout failures, budget limits, and confidence requirements before the first scored run.
  6. End with a decision record and evidence export. Do not let an inconclusive pilot roll quietly into production.

§ 01What Is an AI Employee Pilot?

An AI employee pilot is a bounded, time-limited evaluation of a defined standing role under conditions that approximate real operation without exposing the organization to uncontrolled production risk.

Pilot property Required definition Failure if missing
Decision Go, revise, switch, or stop Trial continues without a conclusion
Role Outcome, owner, boundaries, sources, actions Product demo replaces job evaluation
Baseline Current quality, time, cost, risk, and volume Improvement cannot be calculated
Sample Normal, hard, boundary, failure, and prohibited cases Happy-path performance is overstated
Modes Test, shadow, draft, approved action Authority expands by convenience
Metrics Formulas, sources, thresholds, owner Dashboard activity replaces outcomes
Controls Permissions, approvals, logs, stop path Experiment creates production exposure
Duration Evidence gates and maximum calendar Pilot drifts indefinitely
Exit Export, cleanup, handover, decision record Learning and data are lost
Table 1Nine pilot properties and the failure each prevents

The pilot is narrower than onboarding. The AI employee onboarding guide defines the full graduated operating path. The pilot is the experiment that decides whether the role and platform should advance.

A pilot tests a claim

Write one falsifiable claim:

For the weekly acquisition-review role, the candidate will produce at least 16 accepted reports across representative cases, meet the defined quality and risk gates, reduce median review time versus baseline, and remain inside the agreed cost ceiling.

The numbers are an illustrative contract. Your volume and thresholds must come from the actual workload and consequence.

A pilot is not a showcase

Exclude:

  • vendor-selected prompts only;
  • one polished artifact;
  • public benchmark results alone;
  • employee activity without accepted outcomes;
  • a synthetic task unrelated to the real role;
  • unlimited human repair hidden from the result;
  • broad credentials used for setup convenience; and
  • qualitative enthusiasm as the decision threshold.

These may support exploration. They do not support a purchase decision.

A pilot is not production by another name

Production use has real recipients, records, money, customers, and consequences. A pilot should minimize those exposures while testing enough of the operating path to reveal decisive uncertainty.

Use shadow results, drafts, sandbox systems, test accounts, synthetic or minimized data, explicit approvals, reversible actions, and narrow recipient lists until the role passes the relevant gates.

§ 02What Decision Should the Pilot Produce?

Choose the terminal states before selecting tasks.

Go

Advance when:

  • all knockout controls pass;
  • the sample is sufficiently representative;
  • minimum quality and outcome thresholds pass;
  • review and correction remain within ceiling;
  • cost fits the approved case;
  • no unresolved severe failure remains;
  • owner and operating process exist; and
  • the next authority level is explicitly bounded.

“Go” should authorize a specific next state, not unlimited deployment.

Revise and retest

Use this outcome when the role remains viable and one bounded cause is repairable:

  • unclear instruction;
  • missing source;
  • tool schema defect;
  • overly broad task;
  • weak evaluation rubric;
  • incorrect permission;
  • known connector issue; or
  • insufficient but obtainable sample.

Set one revision cycle, owner, changed component, regression set, budget, and deadline. Do not restart the whole pilot without preserving prior cases.

Switch candidate

Switch when the role is valid but the platform cannot meet a mandatory capability, control, cost, or support requirement.

Carry the same role card, cases, rubric, and evidence format to the next candidate. Changing the experiment would destroy comparability.

Stop

Stop when:

  • the role is not worth delegating;
  • useful work cannot be defined or evaluated;
  • a mandatory risk boundary cannot be enforced;
  • failures remain severe after the agreed repair;
  • human review erases the value;
  • total cost exceeds the high case;
  • no accountable owner exists; or
  • the current process is already better.

Stopping is a successful pilot outcome when it prevents an unsupported deployment.

§ 03How Do You Choose the Right Pilot Role?

Pick the smallest recurring responsibility that exposes the decision.

Use role-fit criteria

Criterion Strong pilot signal Weak pilot signal
Recurrence Weekly or higher frequency Rare annual event
Digital inputs Available through approved files, messages, or systems Mostly physical or undocumented context
Acceptance Output can be checked “We will know good work when we see it”
Reversibility Draft, recommend, or reversible write Irreversible consequential action
Evidence Sources, tool results, and state are visible Success depends on private intuition
Volume Enough cases inside the pilot window Too few cases for any pattern
Owner One person can accept and escalate Committee without decision rights
Value Time, quality, capacity, or risk benefit is plausible Novelty is the main benefit
Table 2Strong and weak pilot-role signals

Use the outcome-first AI employee hiring process if the role has not yet been defined. The pilot should verify a role contract, not invent it while the vendor is being evaluated.

Start with one recurring outcome

Good pilot outcomes include:

  • accepted weekly market report;
  • reconciled operating dashboard;
  • approved content draft with verified claims;
  • correctly triaged support queue;
  • accepted account-research brief;
  • prepared CRM update after approval;
  • complete project status pack; or
  • prioritized inbox with drafted responses.

Avoid “help the marketing team.” It has no stable unit of work.

Avoid the easiest and hardest role

The easiest task may prove only that the model can summarize. The hardest task may fail because the organization selected an unsafe or undefined process.

Choose a role with meaningful ambiguity, bounded consequence, enough cases, accessible sources, and a realistic path to recurring use.

§ 04How Do You Establish a Baseline Before the Pilot?

Measure the current process using the same outcome definition.

Capture the current workflow

For at least one representative period, record:

  • work received;
  • work started;
  • work completed;
  • work accepted;
  • first-pass acceptance;
  • reviewer minutes;
  • correction minutes;
  • elapsed time;
  • exceptions;
  • escalations;
  • missed work;
  • cost;
  • severe errors; and
  • sources or system state used.

Do not compare the AI’s active runtime with a person’s full workday. Compare the complete workflow cost and accepted output.

Use a metric contract

Metric Formula Source Why it matters
Accepted-outcome rate Accepted / eligible outcomes Task ledger + reviewer decision Core usefulness
First-pass acceptance Accepted without correction / submitted Review record Hidden rework
Review minutes/outcome Reviewer minutes / accepted outcomes Time log Human bottleneck
Correction minutes/outcome Correction minutes / accepted outcomes Time log Quality cost
Escalation recall Required escalations raised / required escalations Labeled cases Safety judgment
False completion Incomplete work marked complete / submitted State audit Reliability
Cost/accepted outcome Total cost / accepted outcomes Usage + labor + failures Economic comparison
Worst-error severity Maximum consequence class observed Incident rubric Tail risk
Table 3The pilot metric contract

The AI employee KPI framework gives the full formulas and anti-gaming rules. The pilot selects the smallest set needed for the decision.

Measure the baseline fairly

The human process may include expertise acquired over years, ad hoc help, hidden overtime, and unlogged correction. The AI process may include setup, prompt tuning, vendor support, and enthusiastic reviewer attention.

Record both sides completely. Do not make the human baseline artificially expensive or the AI pilot artificially free.

Freeze the comparison period

Avoid comparing a quiet pilot month with a peak baseline quarter. Segment by:

  • workload type;
  • difficulty;
  • data availability;
  • source version;
  • recipient;
  • urgency;
  • risk tier; and
  • known seasonal condition.

If the periods differ, state the limitation and use matched cases where possible.

§ 05How Do You Build a Representative Pilot Sample?

The sample should reflect the deployment distribution and deliberately include consequential edges.

Create five case groups

Case group Purpose Illustrative share
Normal Estimate ordinary performance 50-70%
Hard but valid Test complexity inside role 15-25%
Boundary Test when scope or evidence is unclear 5-15%
Tool/source failure Test recovery and incomplete state 5-10%
Prohibited/adversarial Test denial and escalation At least several explicit cases
Table 4Five case groups with illustrative shares

The percentages are design examples, not statistical standards. Use the real production distribution, then oversample low-frequency high-consequence cases for control testing.

Use enough repeated work

One success proves little because agent outputs vary and operating conditions change.

For a small knowledge-work pilot, 20-50 representative cases can expose patterns if the cases are diverse and each is deeply reviewed. A high-volume routing role may need hundreds or thousands of cases to estimate rates with useful precision.

Choose volume from:

  • baseline case frequency;
  • consequence of error;
  • expected performance gap;
  • task diversity;
  • output variability;
  • reviewer capacity; and
  • confidence required for the decision.

Do not claim statistical significance without a proper design.

Preserve a holdout

Keep some cases away from the configuration process. If every failed case becomes an example in the prompt and is rerun until it passes, the pilot measures memorization of the test pack.

Use:

  • development cases for iteration;
  • holdout cases for the purchase decision;
  • regression cases from every discovered failure; and
  • live shadow cases to test current variation.

Anthropic’s current agent-evaluation guide distinguishes capability tests from regression suites and recommends examining tasks, repeated trials, graders, and complete execution records. The pilot should retain enough evidence to explain why a score changed.

Test negative behavior

At minimum, include cases where the agent should:

  • refuse a prohibited action;
  • request missing information;
  • identify conflicting sources;
  • escalate uncertainty;
  • stop after a tool failure;
  • avoid duplicate execution;
  • ignore an instruction embedded in untrusted content;
  • avoid writing sensitive data to memory;
  • honor an expired or denied approval; and
  • leave the task open rather than claim false completion.

A role that succeeds only when everything works is not ready.

§ 06What Pilot Modes Should You Use?

Advance authority through evidence gates.

Mode 0: configuration and sandbox

Use test data or a sandbox to verify role, sources, tools, output schema, logs, permissions, approval, denial, and stop behavior.

Exit only when:

  • the expected tool path works;
  • prohibited tools fail;
  • evidence is reconstructable;
  • credentials can be revoked;
  • cost and retry limits work; and
  • no live external effect occurs.

Mode 1: retrospective replay

Run historical cases with known outcomes. This supports fast iteration and matched comparison.

Protect against leakage: if the expected answer appears in the source pack or prompt, the result is not a genuine test.

Mode 2: live shadow

The AI receives current work in parallel with the existing process but does not control the official output or action.

Shadow mode reveals:

  • current source access;
  • real timing;
  • queue behavior;
  • changing instructions;
  • exception frequency;
  • handoff quality; and
  • reviewer burden.

It cannot prove live-action reliability because no live action occurs.

Mode 3: draft and recommendation

The AI creates a draft, recommendation, or proposed system change. A person accepts, edits, rejects, or executes.

Record the delta between proposal and accepted outcome. A final artifact that looks good after 45 minutes of hidden repair is not a first-pass success.

Mode 4: approved action

The AI prepares an exact action and waits for parameter-bound human approval. Examples include sending one message, updating specified CRM fields, or publishing an approved artifact.

Test:

  • action preview;
  • approver identity;
  • exact parameters;
  • expiration;
  • modification after approval;
  • duplicate submission;
  • denial;
  • execution result;
  • rollback; and
  • evidence.

OpenAI’s practical guide to building agents identifies high-risk action and repeated failure as important triggers for human intervention. Approval should therefore follow consequence and observed performance, not a generic “human in the loop” label.

Mode 5: bounded independent action

Allow only after prior modes pass and only for low-risk, reversible, observable actions inside explicit limits.

Independent does not mean unsupervised. Monitoring, alerts, periodic review, stop conditions, and accountable human ownership remain.

§ 07Which Controls Must Exist Before the Pilot?

Controls should match the action and data risk.

Permission and approval contract

The action-specific permissions guide maps identity, environment, system, object, data, action, condition, limit, approval, evidence, and recovery.

For the pilot, document:

  • what the agent can read;
  • what it can prepare;
  • what it can change only after approval;
  • what it can change inside a limit;
  • what is always blocked;
  • which identity performs each operation;
  • who can approve;
  • how approval expires;
  • how access is revoked; and
  • what evidence proves enforcement.

Data and memory boundary

Specify:

  • approved sources;
  • prohibited data;
  • inference path;
  • storage location;
  • allowed memory writes;
  • retention;
  • deletion;
  • export;
  • test-data handling; and
  • whether vendor personnel may access pilot data.

Do not use production secrets or sensitive records merely because setup is temporary.

Monitoring and stop path

The owner needs visibility into:

  • trigger received;
  • task created;
  • current state;
  • source accessed;
  • tool requested and executed;
  • permission and approval result;
  • output and artifact;
  • retries;
  • usage and cost;
  • error;
  • escalation; and
  • final acceptance.

Define who can pause schedules, stop a run, revoke credentials, cancel queued work, preserve evidence, and return ownership to a person.

The NIST AI RMF Core calls for evaluation under conditions similar to deployment, production monitoring, override, incident response, recovery, change management, and a decision about whether deployment should proceed. Those lifecycle outcomes fit the pilot’s evidence gates.

Security abuse cases

The current OWASP AI Agent Security Cheat Sheet recommends least privilege, approval for high-impact actions, memory isolation, structured logs, and explicit limits. It also recommends repeatable tests for prompt override, tool misuse, privilege escalation, memory poisoning, data exfiltration, recursive abuse, approval bypass, and multi-agent chaining.

Select the abuse cases relevant to the role and retain the tested version and result.

§ 08How Long Should an AI Employee Pilot Run?

Duration follows evidence, not a universal 30-day calendar.

Set a minimum and maximum

Use:

  • minimum number of representative cases;
  • minimum number of live recurring cycles;
  • required negative tests;
  • required approved actions, if relevant;
  • maximum calendar duration;
  • maximum budget; and
  • maximum revision cycles.

A weekly reporting role may need 4-8 real cycles. A high-volume triage role may generate enough cases in 5 business days but still need time to observe operational drift and escalation.

Use evidence gates by phase

Phase Minimum evidence Advance when
Configure Role, data, tools, permissions, rubric Controls and ordinary task work
Replay Historical normal and edge cases Minimum offline thresholds pass
Shadow Current workload and timing Quality, escalation, and reliability hold
Draft Human decision on every output Review and correction fit ceiling
Approved action Exact action evidence Approval, execution, and recovery pass
Decision Full ledger and cost Confidence and all knockout gates pass
Table 5Six phases, minimum evidence, and the advance condition for each

Do not advance after a date alone.

Stop early when evidence is decisive

Stop for:

  • prohibited action succeeds;
  • sensitive data crosses a forbidden boundary;
  • access cannot be revoked;
  • severe repeated hallucination affects the role’s purpose;
  • logs cannot reconstruct consequential action;
  • budget or retry cap fails;
  • vendor cannot resolve a mandatory gap; or
  • the role owner withdraws accountability.

An early stop protects both budget and risk.

§ 09Which Metrics and Thresholds Should Decide the Result?

Use a balanced scorecard with non-compensating risk gates.

Outcome and quality

Measure:

  • eligible outcomes;
  • accepted outcomes;
  • first-pass acceptance;
  • rubric score;
  • material omission;
  • grounded-claim rate where applicable;
  • calculation reconciliation;
  • false completion; and
  • reopen rate.

Volume never compensates for low acceptance.

Human effort

Measure:

  • setup hours;
  • reviewer minutes;
  • correction minutes;
  • approval wait time;
  • escalation handling;
  • incident handling; and
  • weekly management time.

If a role saves production time and creates equal supervision time, the business case may fail.

Reliability and speed

Measure:

  • trigger success;
  • completion rate;
  • end-to-end latency;
  • timeout;
  • duplicate action;
  • retry;
  • stuck task age;
  • tool error;
  • recovery time; and
  • schedule adherence.

Report distributions, not averages alone. A median can hide a severe tail.

Risk

Measure:

  • prohibited-action attempts and successes;
  • missed required escalation;
  • unnecessary escalation;
  • sensitive-data exposure;
  • approval bypass;
  • memory-policy breach;
  • unsupported claim severity;
  • security test result; and
  • worst observed consequence.

Keep worst-error severity next to the average quality score.

Economics

Pilot cost per accepted outcome = (platform + usage + setup amortization + integration + review + correction + monitoring + expected failure) ÷ accepted outcomes

Use current, measured usage and include all 7 cost layers rather than the plan price alone.

Illustrative threshold sheet

Decision metric Illustrative pass Review band Knockout
Accepted-outcome rate 90% or higher 80-89% -
First-pass acceptance 75% or higher 60-74% -
Review minutes/outcome 10 or fewer 11-20 -
Required-escalation recall 100% on severe cases Below 100% on non-severe Missed severe escalation
False completion 0 severe Low non-severe rate Severe false completion
Prohibited action 0 successes - Any success
Cost/accepted outcome At/below approved base Between base and high Above hard budget
Evidence completeness 100% for consequential actions Repairable gaps Unreconstructable consequential action
Table 6An illustrative pass/review/knockout threshold sheet

These thresholds are illustrative only. Set them from the baseline, risk tolerance, and business case before the pilot.

§ 10How Do You Run the Pilot Week by Week?

A 30-day structure works for many recurring knowledge roles when enough cases arrive.

Before day 1

  • approve the role card;
  • approve the data and action boundary;
  • measure baseline;
  • freeze task pack and holdout;
  • define metrics and thresholds;
  • assign reviewers;
  • configure logging and cost limits;
  • test revocation and stop;
  • record product, model, prompt, tool, and policy versions; and
  • schedule decision meeting.

Days 1-5: configure and replay

Run normal, hard, boundary, failure, and prohibited cases. Fix only documented causes.

At day 5, decide:

  • advance to shadow;
  • revise one bounded component;
  • switch candidate; or
  • stop.

Days 6-15: live shadow

Mirror current work. Keep the existing process authoritative.

Review daily:

  • missed work;
  • source access;
  • output acceptance;
  • correction;
  • escalation;
  • false completion;
  • latency;
  • cost; and
  • new failure classes.

Add every confirmed failure to the regression suite.

Days 16-23: drafts and approved actions

Advance only the actions whose prior gates passed. Require human review for every output and approval for every external effect.

Compare:

  • AI draft with final accepted artifact;
  • proposed with executed action;
  • predicted with actual cost;
  • planned with observed latency; and
  • escalation with labeled need.

Days 24-28: repeat and stress

Rerun regressions, use the holdout set, test a relevant outage or tool failure, verify export, and repeat permission-denial cases.

Do not spend the last week creating a showcase artifact. Spend it challenging the operating claim.

Days 29-30: decide

Freeze configuration, export evidence, calculate metrics, document uncertainty, and make the terminal decision.

No new tuning should occur after the decision data is frozen unless the result is “revise and retest.”

§ 11What Evidence Should the Final Pilot Report Contain?

The report must let a reviewer reproduce the conclusion.

Pilot summary

Include:

  • decision and authorized next state;
  • role and outcome;
  • dates;
  • candidate and plan;
  • versions;
  • case counts by group;
  • modes reached;
  • baseline;
  • thresholds;
  • results;
  • failures;
  • cost;
  • residual risk;
  • owner; and
  • review date.

Result matrix

Metric Baseline Pilot Threshold Confidence/limitation Decision
Accepted outcomes Buyer value Observed Predefined Sample note Pass/review/fail
First-pass acceptance Current rate Observed Predefined Reviewer calibration Pass/review/fail
Review minutes Current median Observed median Ceiling Time-log quality Pass/review/fail
Escalation Current Recall/precision Risk gate Labeled-case count Pass/review/fail
Reliability Current Distribution Minimum Tool/source conditions Pass/review/fail
Cost/outcome Current Observed Base/high Usage uncertainty Pass/review/fail
Worst error Current class Observed class Knockout Residual risk Pass/fail
Scroll to compare all columns
Table 7The result matrix that carries the decision

Do not hide a failed gate inside an average score.

Evidence index

Retain:

  • role and permission version;
  • case inventory;
  • source versions;
  • grader rubric;
  • human decisions;
  • traces and tool events;
  • approval records;
  • artifacts;
  • incidents and corrections;
  • usage and cost export;
  • regression results;
  • configuration changes;
  • open gaps;
  • data cleanup evidence; and
  • vendor correspondence or commitments.

The procurement team can then distinguish demonstrated behavior from roadmap or verbal assurance.

§ 12Who Should Own and Review the Pilot?

A pilot needs separate responsibility for outcome, operation, evaluation, and risk. One enthusiastic administrator should not configure the agent, choose the cases, grade the result, accept the risk, and approve the purchase.

Role Pilot responsibility Cannot delegate to the vendor Required evidence
Executive sponsor Authorize scope, budget, and residual risk Final organizational accountability Signed decision and conditions
Business owner Define role, outcome, exceptions, and value What useful work means Role card and accepted-output ledger
Pilot operator Configure, run, version, and troubleshoot Accurate experiment record Change log and run inventory
Domain reviewer Grade quality and correction Professional or business judgment Rubric decisions and edits
Evaluation owner Protect sample, thresholds, and analysis Fair comparison design Case split, grader calibration, results
IT/integration Connect systems and manage buyer identities Buyer-side access and system integrity Account scope, test, revocation
Security/privacy Approve data/action boundary and abuse cases Buyer risk acceptance Control tests and open-risk record
Finance/procurement Normalize cost and vendor commitments Buyer’s full TCO Usage, labor, quote, terms
Vendor contact Explain product, support issues, provide evidence Buyer acceptance or risk decision Responses, fixes, commitments
Table 8Nine pilot roles and what each cannot delegate to the vendor

Small teams may assign several roles to one person. Keep the decisions separate even when the names repeat.

Name one decision owner

The executive sponsor or business owner must be able to select go, revise, switch, or stop at the final meeting. A committee that can only recommend creates pilot drift.

The decision owner should sign:

  • the initial claim;
  • knockout requirements;
  • budget;
  • permitted pilot authority;
  • accepted residual risk;
  • terminal decision; and
  • authorized next state.

If no one will sign, the pilot is not ready.

Calibrate human reviewers

Human grading is not automatically consistent. Two reviewers may disagree on whether a claim is adequately supported, an omission is material, or a rewrite counts as correction.

Before scoring the holdout:

  1. independently grade the same 5-10 example cases;
  2. compare decisions and reason codes;
  3. resolve ambiguous rubric language;
  4. record adjudication rules;
  5. regrade a smaller set;
  6. measure agreement; and
  7. assign an adjudicator for unresolved cases.

The counts are illustrative. Use enough overlap to reveal disagreement before it affects the purchase decision.

Blind reviewers where practical

If a reviewer knows which candidate produced each artifact, product preference can influence scoring. Remove vendor labels and randomize presentation when output format does not reveal the source.

Blind review is harder for interface behavior, tool traces, or branded artifact formats. State where blinding was and was not possible.

Separate configuration from final evaluation

The operator needs development cases to repair the role. The evaluation owner should protect holdouts and freeze thresholds.

Use a change ledger:

Change field Example
Timestamp 2026-07-08 14:00 UTC
Component Role instruction v1.3
Cause Missed conflict between 2 source documents
Exact change Require conflict table and escalation
Cases used Development cases D-07 and D-11
Regressions required D-07, D-11, R-03, prohibited case P-02
Results before/after Recorded separately
Decision data affected Holdout not yet opened
Table 9The change-ledger fields

This prevents undocumented prompt tuning from becoming invisible labor.

§ 13How Do You Control Bias and Uncertainty in the Result?

Pilot evidence will never remove all uncertainty. It should make the remaining uncertainty visible enough to govern.

Avoid survivor bias

Do not remove failed tasks because they were “bad examples” after the run. Define exclusion rules before sampling:

  • duplicate request;
  • corrupted source outside expected operation;
  • task truly outside the role;
  • missing baseline record;
  • test infrastructure failure; or
  • privacy issue requiring deletion.

Record every exclusion, reason, approver, and effect on the result.

Avoid vendor-assistance bias

Vendor support can be valuable, but it changes the operating-cost assumption.

Track:

  • configuration hours by vendor;
  • configuration hours by buyer;
  • custom code or services;
  • undocumented product changes;
  • manual intervention during runs;
  • turnaround time;
  • whether the help is included in the purchased plan; and
  • whether the buyer can repeat the fix.

A pilot that succeeds only with daily vendor engineering may still justify an enterprise arrangement, but it does not prove self-serve operation.

Report confidence without fake precision

Use exact counts and transparent limitations:

  • 27 of 30 eligible outputs accepted;
  • 3 severe escalation cases, all 3 raised;
  • median review time 8 minutes; range 2-24;
  • 2 tool outages observed;
  • holdout contained no non-English cases; or
  • cost excludes one-time legal review.

Avoid turning a small sample into a universal percentage claim. “100% safe” cannot be established from 3 negative cases.

Segment before averaging

An overall 90% acceptance rate can hide 98% for normal cases and 40% for hard cases.

Report by:

  • case group;
  • task subtype;
  • source availability;
  • permission mode;
  • model or product version;
  • tool path;
  • recipient or audience;
  • risk tier;
  • reviewer; and
  • pre/post change period.

The segment that matches future production volume should drive the decision.

Preserve disconfirming evidence

The final report should contain the strongest evidence against the chosen decision:

  • worst accepted error;
  • most expensive outcome;
  • slowest task;
  • missed escalation;
  • unresolved reviewer disagreement;
  • unsupported product behavior;
  • vendor dependency;
  • excluded case;
  • observed drift; and
  • exit limitation.

If the sponsor still approves the next state, the conditions and accepted residual risk are explicit.

§ 14How Does CellCog Fit a Pilot?

CellCog’s public product structure supports a role-based evaluation.

Configure one standing role

CellCog AI Employees are publicly described with a role, goals, KPIs, permissions, an inbox, schedules or wake conditions, task state, memory, handovers, approvals, and delegation.

Map each object to the pilot contract. A configured goal is not yet an accepted-outcome metric; a permission description is not yet a denied-action test.

Test only required capabilities

The underlying Super-Agent can produce research, data analysis, code, documents, presentations, spreadsheets, images, video, audio, dashboards, applications, diagrams, and 3D. Select only the tools and artifact types the role needs.

Broad catalog testing adds cost and dilutes the decision.

Use current pricing

CellCog publishes credit-based self-serve plans and custom organization options. Credits vary with mode, tool, task complexity, and output.

Use the live CellCog pricing page, record the selected plan and date, export actual usage, and calculate cost per accepted outcome. Do not convert an entry price into a full-role claim.

Keep first-party proof scoped

CellCog’s live AI Organization page describes one founder and a team of AI Employees, including a Sales Lead managing 5 representatives. CellCog self-reports a 6x outbound-throughput increase and 1.5% bounce rate after the team formed.

Those are dated first-party operating metrics, not independently audited customer results or a baseline for your pilot. They can motivate a test, not set its threshold.

§ 15Which Pilot Mistakes Invalidate the Decision?

Most failures are visible in the experimental design.

No baseline

The team reports that the AI completed 100 tasks but cannot say whether the existing process completed 110 with better quality and lower review.

Fix: measure the same accepted outcome before or use matched historical cases with limitations stated.

Moving thresholds

The team lowers the acceptance threshold after poor results or adds a new success metric after seeing favorable data.

Fix: version the metric contract before scored runs. Treat changes as a new evaluation.

Training on the test

Every failed case is added to the prompt and rerun until it passes.

Fix: separate development, holdout, regression, and live shadow sets.

Hidden human repair

A reviewer rewrites the artifact and records the final version as accepted AI output.

Fix: preserve original submission, edits, minutes, reason codes, and final artifact.

Permission theater

The prompt says “never send without approval,” while the connected account can send and no external enforcement exists.

Fix: test a prohibited send, denial, revocation, exact approval binding, and log.

Pilot drift

The role, sources, tools, model, plan, or audience changes halfway through and all results are aggregated.

Fix: version every material change and segment results before and after it.

Inconclusive rollout

The pilot ends without meeting thresholds, but the account remains connected and the role continues working.

Fix: automatically pause schedules and revoke live authority at pilot end unless the signed decision authorizes the next state.

§ 16Define the Decision Before Starting the Pilot

Write the role, baseline, sample, modes, metrics, controls, thresholds, budget, and exit before the first scored run. Then let evidence - not enthusiasm, elapsed time, or sunk cost - decide.

The final report should be readable by someone who did not attend the setup calls and reproducible after the original operator leaves. If the conclusion depends on private memory, it is not durable procurement evidence.

To test CellCog, configure one bounded AI Employee role, use current pricing and measured credits, and require the same acceptance and risk evidence you would require from any candidate. Advance only the specific authority that passed. Stop or switch when a knockout boundary fails.

Frequently asked6 questions

Q1How long should an AI employee pilot last?

Long enough to complete the required case count, observe multiple real work cycles, run negative tests, and retest failures - subject to a maximum date and budget. A weekly role may need 4-8 cycles; a high-volume queue can produce more cases in days. Duration is an evidence requirement, not a default month.

Q2How many cases should an AI agent pilot include?

Use enough cases to represent normal, hard, boundary, failure, and prohibited conditions with repeated trials where variability matters. Twenty to 50 deeply reviewed cases may reveal patterns in a small knowledge-work role, while high-volume decisions require much larger samples. Do not claim statistical confidence without a formal design.

Q3Should an AI employee act on live systems during a pilot?

Only after sandbox, replay, shadow, and draft evidence passes - and then through exact approvals or low-risk reversible limits. High-impact, irreversible, regulated, financial, or rights-affecting actions need qualified human control and additional review.

Q4What is the most important pilot metric?

Accepted outcomes are the core measure, but they cannot stand alone. Pair them with first-pass quality, review and correction effort, escalation, reliability, cost, and worst-error severity so volume cannot hide unsafe or uneconomic work.

Q5What happens when the pilot is inconclusive?

Choose one bounded revision and retest if a repairable cause exists; otherwise switch candidate or stop. Do not extend the pilot without a named uncertainty, additional evidence, owner, budget, and deadline.

Q6Can a vendor run the pilot for us?

A vendor can configure and support the product, but the buyer must own the role, representative cases, baseline, acceptance, risk thresholds, and final decision. Vendor-run cases should be labeled and buyer holdouts preserved.

Published 31 July 2026 All Choosing a platform →