Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentPricingBlogStoryContact

AI Employee KPIs: Measure Outcomes, Quality, and Escalation

Napkin-style sketch of a seven-layer KPI scorecard pyramid with accepted outcomes at the top and activity diagnostics at the base
Fig 0Measure the work the business accepts, not the activity the AI produces.

The best AI employee KPI is not tasks completed. It is the rate and value of accepted outcomes produced inside the role’s quality, escalation, risk, time, and cost boundaries.

Use a seven-part scorecard:

  1. business outcome;
  2. acceptance and quality;
  3. escalation and human control;
  4. flow and reliability;
  5. cost;
  6. risk; and
  7. activity diagnostics.

Activity explains what the worker did. It does not prove the work was useful. One hundred generated drafts can produce zero value if none are accepted. A high completion rate can hide missed exceptions, silent corrections, reopened work, or one severe incident.

A complete metric contract covers one recurring role. Goals, risk assessment, and complete ROI remain separate decisions.

On this page · 17 sectionsOpen
  1. AI Employee KPIs at a Glance
  2. What Is an AI Employee KPI?
  3. Why Are Tasks Completed a Weak Primary KPI?
  4. Which Outcome KPIs Should an AI Employee Have?
  5. Which Quality KPIs Matter?
  6. How Do You Measure Escalation Quality?
  7. Which Reliability and Flow KPIs Matter?
  8. Which Cost KPIs Should You Track?
  9. Which Risk KPIs Should You Track?
  10. Which Activity Metrics Are Useful?
  11. How Do You Write an AI Employee Metric Contract?
  12. What Should an AI Employee Dashboard Show?
  13. How Often Should AI Employee KPIs Be Reviewed?
  14. How Does CellCog Support AI Employee KPIs?
  15. What Does a Worked AI Employee KPI Scorecard Look Like?
  16. Common AI Employee KPI Mistakes
  17. Final Recommendation
Key points6 · 21 min full read
  1. Define an accepted outcome before measuring the AI employee; “task completed” is only a system state.
  2. Use outcome, quality, escalation, reliability, cost, and risk KPIs together. Keep prompts, tokens, shifts, and tool calls as diagnostics.
  3. Track accepted without correction, accepted after correction, rejected, correctly escalated, falsely escalated, missed escalation, and reopened.
  4. Preserve error severity and worst-case failure beside averages; one high-impact miss can outweigh a strong acceptance rate.
  5. Calculate total cost per accepted outcome using platform, setup, integration, review, correction, monitoring, and failure exposure.
  6. In CellCog, map role goals, KPIs, task state, approvals, shifts, handovers, and dashboards into the organization’s metric contract; verify the current product implementation.

§ 01AI Employee KPIs at a Glance

KPI layer Primary question Example metric Why it matters
1. Outcome Did the business receive the intended result? Accepted outcomes/week Connects work to role purpose
2. Quality Was the result usable and evidence-complete? First-pass acceptance Exposes hidden correction
3. Escalation Did the right cases reach a person? Escalation recall/precision Measures control quality
4. Flow/reliability Did work finish consistently? Cycle time, reopen rate Detects queues and brittle completion
5. Cost What did usable work cost? Total cost/accepted outcome Normalizes economics
6. Risk What harm or policy failure occurred? Incidents by severity Preserves tail risk
7. Activity What resources and actions explain results? Runs, credits, tool calls Diagnoses, not proves, performance
Table 1The seven KPI layers

A launch dashboard should show at least one measure from the first six layers. Activity metrics support diagnosis after a result changes.

§ 02What Is an AI Employee KPI?

An AI employee KPI is a decision-linked measure that shows whether a standing AI role is producing its accepted outcome within defined quality, control, time, cost, and risk boundaries.

A complete KPI has:

Field Question
Role/outcome Which responsibility is being measured?
Metric What exactly is counted or calculated?
Numerator Which events/values are included?
Denominator Received, completed, accepted, customer, period, or cost?
Source Which system or review record supplies the data?
Owner Who is accountable for accuracy and action?
Cadence Per task, shift, week, or month?
Threshold What triggers advance, hold, narrow, pause, or stop?
Segment Which task subtype, channel, risk class, or customer group?
Anti-gaming rule How could the number improve while work worsens?
Table 2The ten fields of a complete KPI

“Quality score” without a rubric, denominator, source, and decision is not a KPI contract.

Goal versus KPI

A goal states the outcome and constraint:

Deliver a source-backed competitor-change briefing every Friday by 2:00 p.m., covering the approved 12 companies, with material changes and unresolved conflicts visible.

KPIs measure whether the role achieves it: accepted brief delivered on time; required-source completeness; material-change recall; unsupported-claim rate; reviewer/correction minutes; escalation precision; cost per accepted brief; and worst-error severity.

Define the objective separately, then use the measurement system to test whether the role achieves it.

KPI versus metric

Every KPI is a metric; not every metric is a KPI.

KPI: accepted briefs delivered on time. Supporting quality metric: claims with required source. Diagnostic: web tool calls per brief. Log field: model/mode version.

Do not fill the executive dashboard with every available counter.

KPI versus evaluation

An evaluation tests performance on a defined case or set. A KPI tracks live operating performance over time.

Use evaluations before launch; before permission expansion; after material change; for regression testing; and after a failure. Use KPIs per shift; weekly; monthly; by task subtype; and at operating reviews.

The two should connect. If a live KPI falls, rerun or extend the relevant evaluation cases.

§ 03Why Are Tasks Completed a Weak Primary KPI?

Completion is self-reported by the work system unless a downstream user or postcondition verifies it.

A task can be marked complete when output exists but is unusable; required sources are missing; the reviewer repairs it silently; an action call succeeded but the business state did not change; the customer reopens the issue; an exception should have escalated; duplicate work was produced; the wrong task was completed; or the result arrived too late.

Separate seven outcome states

Track:

  1. Accepted without correction
  2. Accepted after correction
  3. Rejected
  4. Correctly escalated
  5. Falsely escalated
  6. Missed escalation
  7. Reopened

A role may also have canceled, duplicate, expired, or out-of-scope states. Keep them outside completed/accepted unless the contract says otherwise.

Define accepted

An accepted outcome satisfies the role’s completion contract; meets the quality rubric; uses required evidence; stays inside permissions; reaches the downstream user/system; receives the required approval; passes its postcondition; and does not require undisclosed material repair.

For a draft role, acceptance may mean an editor approves the draft - not that the content was published.

Use the correct denominator

Metric Formula What it answers
Completion rate Completed / received Did the queue move?
First-pass acceptance Accepted without correction / completed Was output usable immediately?
Total acceptance (Accepted + accepted after correction) / completed Did work become usable eventually?
Rejection rate Rejected / completed How much output had no usable value?
Reopen rate Reopened / accepted Did “done” remain done?
Accepted yield Accepted outcomes / received What share of demand became usable work?
Table 3Six outcome metrics and their denominators

Report received, completed, and accepted counts beside rates. A 100% acceptance rate based on one easy case should not look like mature performance.

Record correction separately

If the reviewer changes facts, reasoning, policy application, audience, action, or structure, record the correction category; minutes; severity; who corrected it; whether the issue was detectable; whether a rerun was required; and whether the fix became a new evaluation case.

Silent correction makes weak automation appear strong.

§ 04Which Outcome KPIs Should an AI Employee Have?

Choose one primary accepted-outcome KPI and one business-result measure.

Primary accepted-outcome KPI

Examples: accepted research briefs delivered by deadline; accepted support drafts per low-risk ticket received; reconciled reports with verified postconditions; qualified research packets accepted by sales; approved content drafts that pass source and brand review; accurately routed exceptions resolved within policy; and task handovers accepted by the next owner.

The count must represent a useful unit, not any artifact.

Business-result measure

Connect the role to a downstream result without claiming causation prematurely:

Role Accepted outcome KPI Downstream measure
Research Accepted decision briefs Decisions supported or analyst capacity released
Support Accepted triage/draft packets Resolution time, recontact, satisfaction
Sales research Accepted account packets Meetings assisted, conversion by qualified segment
Content Accepted source-backed drafts Qualified traffic, assisted pipeline, correction rate
Operations Accepted exception packets Resolution time, reopened cases, leakage avoided
Data Reconciled report refreshes Decision latency, data defects found
Table 4Accepted outcomes and their downstream measures by role

Keep the accepted outcome as the worker’s direct KPI. Treat downstream metrics as shared or influenced when people, market conditions, or other systems also determine the result.

Use contribution boundaries

Do not assign total revenue, customer satisfaction, or company growth to one AI employee unless a credible causal design supports it.

Use labels: owned (role directly produces the accepted unit); shared (role and human/team jointly determine result); influenced (role contributes evidence or capacity); diagnostic (metric explains operation).

This stops the role from optimizing an outcome it cannot control.

§ 05Which Quality KPIs Matter?

Quality is multidimensional. One subjective 1-5 score hides the failure.

Quality metric Formula/rule Failure detected
First-pass acceptance Accepted without correction / completed General usability
Evidence completeness Required evidence present / required evidence Missing support
Factual error rate Outputs with material factual error / reviewed outputs Incorrect claims
Policy adherence Cases following required policy / applicable cases Process violation
Format completeness Required fields present / required fields Incomplete handoff
Correction minutes Total correction minutes / accepted outcomes Hidden human labor
Reviewer agreement Same-score cases / double-reviewed cases Rubric ambiguity
Reopen rate Reopened / accepted Fragile completion
Postcondition pass Intended verified results / actions Tool/business mismatch
Table 5Nine quality metrics and the failures they detect

Build a rubric

Use criteria with pass/fail or anchored levels: outcome completeness; factual/evidentiary correctness; source authority/freshness; policy compliance; reasoning/decision validity; format/channel; uncertainty disclosure; permissions; escalation; and downstream usability.

Define material and cosmetic corrections. Do not let grammar edits count like a wrong customer or unauthorized action.

Sample strategically

Review every high-impact item; every escalation; every new task subtype; every new source/tool/action; all failures; a random sample of routine accepted work; and a targeted sample of cases near thresholds.

Sampling only the best outputs creates a marketing dashboard, not an operating control.

§ 06How Do You Measure Escalation Quality?

Escalation is a prediction and routing problem. The worker must raise cases that need a person without sending every case to the human queue.

Use a four-cell table:

Actual need for escalation Worker escalates Worker does not escalate
Yes Correct escalation Missed escalation
No False escalation Correct independent handling
Table 6The escalation four-cell table

Then calculate: escalation recall = correct escalations / all cases that required escalation. Escalation precision = correct escalations / all escalations raised.

Why recall and precision both matter

Low recall means the worker misses important cases. Low precision means people spend time reviewing routine cases.

Set different thresholds by severity. Missing a high-impact case can be unacceptable even if overall recall is high.

Measure the escalation packet

Score whether the packet includes the case and current state; the trigger; sources; the policy/decision boundary; uncertainty; attempted steps; options and tradeoffs; the recommended next action; the deadline; the requested decision; and whether work is paused or safely continuing.

An escalation that says “I’m unsure” without evidence may be safe but operationally weak.

Measure response and closure

Track time from trigger to escalation; time waiting for a human; approver response time; decisions returned with enough clarity; cases resumed correctly; repeated escalations for the same unresolved rule; and handover of open escalation state.

Some escalation failures belong to the organization, not the AI employee. A missing owner or slow approval queue should not be mislabeled as model failure.

§ 07Which Reliability and Flow KPIs Matter?

Use flow metrics to detect queues, repeated work, and unreliable completion.

Metric Formula/definition Use
Cycle time Accepted time - received time End-to-end speed
Active work time Time actually processing Execution efficiency
Waiting time Time blocked on input/approval/system Coordination bottleneck
Queue age Current time - received time Stale work
On-time rate Accepted before deadline / due outcomes Cadence reliability
Retry rate Repeated attempts / initiated tasks Tool/instruction instability
Duplicate rate Duplicate outputs/actions / outcomes Trigger/state failure
Handover reconciliation Correctly continued open items / open items Cross-shift continuity
Postcondition pass Verified intended state / actions Execution correctness
Table 7Nine flow and reliability metrics

Report median and a tail percentile such as p90 or p95 when volume supports it. An average can hide a small set of very late tasks.

The percentile choice and minimum sample should be predeclared. Do not report a p95 from a handful of observations as if it were stable.

Separate worker time from system waiting

If a task takes 18 hours because it waited 17 hours for approval, optimize the approval process before buying a faster model.

Segment cycle time into queue; active processing; tool/system wait; human approval; correction; retry; and downstream acceptance.

Track durable completion

Reopened tasks, corrected reports, bounced messages, rolled-back changes, and repeated customer contact show that “done” did not hold.

Use a durability window appropriate to the outcome. A support case may reopen within days; a monthly report can fail when the next refresh cannot reproduce it.

§ 08Which Cost KPIs Should You Track?

Measure total cost per accepted outcome, not credits or plan price alone.

Use: total cost per accepted outcome = total role cost / accepted outcomes.

Total role cost includes platform/subscription; variable usage and tools; setup amortization; integrations and administration; human review; correction and rework; monitoring/governance; and expected failure exposure.

The AI employee cost framework defines each layer and worked examples.

Track direct and operating cost separately

Cost metric Formula Decision
Direct usage/outcome Platform + tools / accepted Compare execution consumption
Review cost/outcome Review hours x loaded rate / accepted Detect oversight burden
Correction cost/outcome Repair/rerun cost / accepted Detect hidden quality cost
Total cost/outcome All cost layers / accepted Compare complete economics
Unused-capacity rate Unused eligible capacity / purchased Right-size plan
Cost variance Actual - budget Detect drift
Table 8Six cost metrics and the decisions they inform

Do not bury reviewer labor inside “existing payroll.” Record capacity released separately from cash savings.

When several AI roles share the same supervisors, the human span-of-control workload model converts review, exception, approval, correction, and peak-queue demand into a portfolio capacity limit.

Segment cost by task subtype

A role may look economic overall while one subtype consumes most review and retries.

Segment by normal versus exception; channel; output format; source/tool; customer/account tier; permission state; model/mode; accepted/rejected; severity; and new versus familiar case.

Use the segment to narrow or redesign the role rather than average expensive failure into cheap routine work.

§ 09Which Risk KPIs Should You Track?

Risk metrics preserve consequence that averages hide.

Severity model

Use an organization-specific scale. An illustrative model:

Severity Description Example response
S0 No issue Record pass
S1 Cosmetic/low-cost correction Fix and trend
S2 Material error contained before external impact Correct and add regression
S3 External, financial, privacy, security, or reputation impact Incident response
S4 Severe rights, safety, legal, or irreversible impact Immediate containment/escalation
Table 9An illustrative severity scale

These labels are an illustrative operating scale, not a universal standard.

Risk measures

Track incidents by severity; policy violations; prohibited-action attempts; approval bypasses; sensitive-data exposure; wrong-target actions; rollback/recovery events; undetected errors found downstream; time to contain; repeated failure after correction; and maximum plausible loss/exposure.

Always show the highest severity beside incident rate.

Measure controls, not just failures

Control metric What it shows
Required approvals obtained Gate execution
Denied out-of-scope attempts Boundary behavior
Revocation test passed Containment readiness
Recovery test passed Reversibility
Access review completed Permission freshness
Regression set passed Change safety
Incident action completed Closure discipline
Table 10Seven control metrics

The permission and approval matrix defines the action boundary. KPI reporting should show whether it worked in practice.

Keep risk as a gate

Do not create one weighted score in which fast, cheap output compensates for a severe control failure.

Use mandatory gates such as zero unresolved S3/S4 incidents; no missed escalation on defined critical cases; no approval bypass; no cross-tenant/customer access; recovery demonstrated before authority expansion; and immediate pause on prohibited action.

Thresholds must match the organization’s tolerance and applicable obligations.

§ 10Which Activity Metrics Are Useful?

Activity metrics explain performance but should not lead the scorecard.

Useful diagnostics include tasks received/started/completed; shifts/runs; messages; tokens/credits; tool calls; sources retrieved; context size; retries; steps; model/mode; media/file volume; execution time; memory reads/writes; approvals requested; and delegated subtasks.

Use activity to investigate change

Examples: accepted outcomes fall while tool calls rise; correction increases after a context/source change; cost rises because retries doubled; false escalations rise after a policy update; cycle time rises because human wait grew; duplicate actions follow a new trigger; or one mode consumes more credits without improving acceptance.

The diagnostic explains where to look. It is not the business result.

Reject vanity metrics

Do not use as a primary value claim: prompts answered; hours “worked”; tasks attempted; words generated; messages sent; dashboards created; agents delegated; tools connected; or outputs produced.

More activity can increase cost and risk while reducing accepted value.

§ 11How Do You Write an AI Employee Metric Contract?

Write the contract before the pilot so the role cannot redefine success after seeing results. The AI employee pilot guide turns that contract into a baseline, representative case pack, staged operating modes, evidence gates, and a precommitted go/revise/switch/stop decision.

Metric contract template

Field Example
Role Competitive Intelligence AI Employee
Outcome Accepted weekly competitor-change brief
Received unit One scheduled weekly cycle
Accepted unit Brief passes editor rubric and deadline
Primary KPI Accepted briefs delivered on time
Quality Evidence completeness; material-error rate
Escalation Recall/precision on material/conflicting changes
Flow Cycle time; waiting-on-owner
Cost Total cost/accepted brief
Risk Worst severity; prohibited action
Diagnostics Credits, web calls, retries
Segments Competitor, change type, source
Owners Strategy lead; reviewer; system owner
Cadence Per brief + monthly role review
Gates Advance/hold/narrow/pause rules
Table 11An example metric contract

Define data events

Create consistent events: received; started; blocked; escalated; draft completed; reviewed; corrected; accepted; rejected; action attempted; postcondition passed/failed; reopened; incident opened/closed; and handover.

Each event needs role, task, subtype, timestamp, current state, owner, and relevant version.

Define attribution

If several agents and people contribute, record the originator; current owner; delegated worker; reviewer; approver; executor; accepted by; and final business owner.

Avoid crediting every participant with the whole outcome.

Define decision rules

For every KPI, state the threshold; minimum sample; severity override; trend window; alert; owner response; advance/hold/narrow/pause/stop; and required retest.

A dashboard without a decision rule is reporting, not management.

§ 12What Should an AI Employee Dashboard Show?

Use a small decision dashboard with drill-down.

Executive view

Section Current Target/gate Trend Owner action
Accepted outcomes 18 >=16 Up Continue
First-pass acceptance 78% >=80% Flat Repair top correction
Escalation recall 100% 100% critical cases Stable Continue
Escalation precision 67% >=70% Down Clarify exception rule
Median correction 8 min <=10 min Stable Continue
On-time 94% >=95% Up Review late subtype
Cost/accepted $42 <=$45 Down Continue
Worst severity S2 No unresolved S3/S4 Stable Close regression
Table 12An illustrative executive view

All numbers in this table are illustrative.

Drill-down

Allow filtering by task subtype; time period; source/tool; model/mode; customer/account segment; reviewer; accepted/corrected/rejected; escalation state; permission state; and severity.

Keep representative examples behind each number. A team should be able to open the corrected and failed cases, not only see the rate.

Anti-gaming checks

KPI How it can be gamed Countermeasure
Acceptance Lower rubric or review less Fixed rubric + sampling
Cycle time Close incomplete work Postcondition + reopen rate
Escalation recall Escalate everything Track precision and wait
Cost/outcome Exclude review/correction Total-cost ledger
On-time Reject difficult cases Received-to-accepted yield
Error rate Review only easy work Risk-weighted sampling
Output count Split one outcome into many Fixed unit definition
Table 13Seven KPIs, how they get gamed, and the countermeasures

Metrics shape behavior. Review the loophole before rewarding the number.

§ 13How Often Should AI Employee KPIs Be Reviewed?

Use layered cadences.

Cadence Review
Per action/task Acceptance, evidence, escalation, postcondition, severity
Per shift Queue, blockers, retries, handover, cost
Weekly during pilot Outcome, quality, review/correction, exceptions, incidents
Monthly when stable Role scope, trend, economics, access, source/tool drift
After material change Regression tests plus affected KPI baseline
After incident Containment, root cause, controls, threshold, re-entry
Before expansion Representative evidence and severity gate
Table 14Review cadences by layer

The onboarding guide uses the same evidence to decide whether authority advances. Do not expand because a calendar milestone arrived.

NIST’s AI RMF Core calls for ongoing monitoring and periodic review, defined human-AI oversight roles, measurable improvement, incident response, recovery, change management, and decommissioning. The exact KPI cadence should follow the role’s impact and change rate.

Use control limits, not constant reaction

For stable high-volume metrics, define the normal range; warning threshold; pause threshold; severity override; and minimum sample.

Do not change instructions after every routine fluctuation. Investigate sustained or severe changes.

Rebaseline carefully

Rebaseline when the role scope changes; the task subtype changes; a source/tool changes; the model/mode changes; approval/permission changes; the reviewer/rubric changes; the business process changes; or the outcome value changes.

Preserve the prior baseline and explain the break. Otherwise improvement can be manufactured by changing the denominator.

Retire metrics that no longer drive decisions

A metric creates maintenance work: events must be captured, definitions governed, edge cases resolved, and dashboards reviewed. Remove or demote it when nobody owns the response; the measure duplicates another KPI; the denominator cannot be reproduced; the role no longer controls or influences it; the metric has become a target that degrades real work; a changed process makes historical comparison invalid; or the diagnostic has not informed a decision across several relevant review periods.

Do not delete the historical definition. Mark the metric retired, record the reason and date, preserve the prior series where required, and identify any replacement.

Before adding a KPI, ask:

  1. Which decision changes when it crosses the threshold?
  2. Who will make that decision?
  3. Can the underlying cases be inspected?
  4. What behavior might the KPI reward unintentionally?
  5. Which higher-priority measure prevents that gaming?

If there is no decision, the number belongs in exploration or diagnostics - not the management scorecard.

Review the scorecard with the people who receive, correct, approve, and act on the work. A metric can be technically reproducible and still omit the burden experienced downstream.

§ 14How Does CellCog Support AI Employee KPIs?

CellCog’s AI Employees page describes goals, KPIs, task lists, shifts, approvals, memory, handovers, and live dashboards as part of its employee model.

Map those surfaces into the metric contract:

CellCog surface KPI use Buyer verification
Goal/KPI configuration State intended outcome Formula, threshold, owner
Task list/board Received, state, blocker, completion Accepted state exists outside self-report
Shifts Period/cadence Start/end and task attribution
Approvals Human-control events Request, decision, expiry, result
Memory/context Version/change segment Source and memory change visible
Handover Open-work continuity Next owner accepts state
Dashboard Decision view Definitions and drill-down available
Credits/usage Cost diagnostic Role/time attribution export
Table 15CellCog surfaces mapped to KPI uses

Use the current AI Employees guide to verify live behavior. Public product descriptions do not prove that every formula, segment, export, or custom dashboard field exists in the current account.

Do not use CellCog’s company metrics as your role KPI

CellCog publishes first-party organizational and product metrics elsewhere. Those may illustrate how CellCog describes its own operation, but they do not establish your baseline; your accepted outcome; your reviewer burden; your error rate; your economics; or your risk tolerance.

Build the role dashboard from your process evidence.

Keep an organization-owned ledger

If the platform does not expose a required metric directly, maintain the task ID and subtype; received/completed/reviewed timestamps; acceptance state; correction; escalation; source/tool/version; usage/cost; postcondition; severity; and owner.

The organization must be able to compare the role even if the vendor, plan, or interface changes.

§ 15What Does a Worked AI Employee KPI Scorecard Look Like?

The following examples are illustrative. They show formulas and decisions, not CellCog results or industry benchmarks.

Example 1: weekly research role

The role owns one accepted competitor-change brief per week.

During one four-week period: 4 briefs received; 4 completed; 3 accepted without correction; 1 accepted after 15 minutes of correction; 48 required claim-source pairs; 47 present and valid; 2 cases required escalation; both were escalated; 1 additional false escalation; no briefs reopened; $320 total role cost; and worst observed issue S1.

KPI Calculation Result Decision
Accepted yield 4 accepted / 4 received 100% Pass
First-pass acceptance 3 / 4 completed 75% Improve
Evidence completeness 47 / 48 97.9% Review missing-source cause
Escalation recall 2 / 2 required 100% Pass
Escalation precision 2 / 3 raised 66.7% Clarify one false trigger
Correction/outcome 15 minutes / 4 3.75 minutes Pass if below gate
Cost/accepted brief $320 / 4 $80 Compare with baseline/value
Reopen rate 0 / 4 0% Pass
Worst severity Maximum S1 Inside illustrative tolerance
Table 16Worked example 1: the research role scorecard

The headline is not “100% completed.” The decision is more specific: the role produced every accepted brief and escalated every required case, but first-pass quality and one false escalation still need repair.

Use the best-task suitability framework before treating this scorecard as a reason to expand the role. A strong score on one bounded research outcome does not qualify public publishing, customer outreach, or unrelated analysis.

Example 2: support triage and draft role

The role receives 200 allowlisted low-risk tickets. It classifies each ticket and prepares a response draft. A person handles refunds, security issues, account changes, and exceptions.

Observed period: 200 received; 180 completed; 135 accepted without correction; 30 accepted after correction; 15 rejected; 20 still open at cutoff; 24 cases required escalation; 22 were correctly escalated; 8 other cases were falsely escalated; 9 accepted cases reopened; $2,475 total role cost; and one S2 issue contained before external impact.

KPI Calculation Result Interpretation
Completion 180 / 200 90% Queue movement only
Accepted yield 165 / 200 82.5% Usable output/demand
First-pass acceptance 135 / 180 75% Material review burden remains
Total acceptance 165 / 180 91.7% Eventual usability
Rejection 15 / 180 8.3% Diagnose by subtype
Escalation recall 22 / 24 91.7% Two required cases missed
Escalation precision 22 / 30 73.3% Eight unnecessary escalations
Reopen 9 / 165 5.5% Completion durability issue
Cost/accepted $2,475 / 165 $15 Compare with current process
Worst severity Maximum S2 Add regression and contain
Table 17Worked example 2: the support role scorecard

This role should not receive broader authority merely because total acceptance is 91.7%. The two missed escalations and the S2 case determine the next decision. Depending on their subtype and consequence, the team may hold, narrow, or return to shadow/draft mode.

Use the outcome-first hiring contract to confirm that low-risk classification and draft preparation - not final exception resolution - remain the role’s responsibility.

Diagnose the support example by subtype

Suppose the 15 rejections break down as:

Subtype Received Rejected Rejection rate
Order-status question 70 1 1.4%
Product how-to 60 2 3.3%
Billing clarification 35 5 14.3%
Account/access issue 15 7 46.7%
Table 18Rejections by subtype

The aggregate rejection rate does not justify treating all tickets alike. Keep order-status and product how-to in the pilot; narrow, repair, or human-route account/access issues. If billing clarification follows stable, explicit rules, compare deterministic workflow automation with an AI employee instead of solving every failure with more model autonomy.

Convert the scorecard into a decision

Result Action
Outcome, quality, escalation, cost, and risk pass Continue at current authority
Quality close; no severe control failure Hold and repair
One subtype passes; another repeatedly fails Narrow the role
Cost fails but quality/risk pass Right-size mode, process, or plan
Escalation or permission gate fails Reduce authority and retest
Severe unresolved incident Pause and contain
Stable rule explains most failures Route that path to workflow automation
Table 19The decision table

The six-stage onboarding method defines how evidence changes authority. KPI review supplies the evidence; it does not automatically grant the next permission.

§ 16Common AI Employee KPI Mistakes

Mistake Why it misleads Better measure
Tasks completed Self-reported system state Accepted yield
Output volume More can mean more waste Accepted outcomes
Accuracy alone Hides severity and coverage Rubric + severity
Acceptance alone Hides silent correction First-pass + correction time
Escalation count High can be noise or safety Recall + precision
Average cycle time Hides tail delay Median + p90/p95
Platform cost only Excludes human work/failure Total cost/accepted
One composite score Benefits offset severe risk Separate risk gates
Company revenue Attribution is weak Owned outcome + influenced result
Benchmark score Different task/distribution Role-specific cases/live KPIs
Dashboard without samples Cannot inspect failures Linked evidence
Threshold after results Encourages moving goalposts Precommitted contract
No segmentation Easy cases hide hard failures Task/risk segments
Constant rebaseline Erases regression Versioned baseline
Table 20Fourteen KPI mistakes and better measures

§ 17Final Recommendation

Measure the work the business accepts, not the activity the AI produces.

Use seven layers: outcome + quality + escalation + reliability + cost + risk + diagnostics.

Start with accepted outcomes. Add correction burden, escalation recall and precision, durable completion, total cost per accepted outcome, and worst-error severity. Use tasks, shifts, credits, tool calls, and tokens only to explain why those results changed.

For CellCog, connect goals, KPIs, task state, shifts, approvals, handovers, dashboards, and usage into the same buyer-owned metric contract. Verify which fields and exports the current product exposes.

A role should expand only when its outcomes are accepted, its exceptions reach the right people, its cost remains viable, and its worst failures stay inside tolerance.

Frequently asked6 questions

Q1What is the best KPI for an AI employee?

Use accepted outcomes per period, paired with first-pass quality, escalation, total cost, and worst-error severity. The exact accepted unit depends on the role.

Q2Should tasks completed be an AI employee KPI?

It can be a flow metric, but not the primary value KPI. Completion may hide rejection, correction, missed escalation, failed postconditions, or reopened work. Track accepted yield instead.

Q3How do you measure AI employee quality?

Use a written rubric covering completeness, facts/evidence, policy, decision validity, format, uncertainty, permission, escalation, and downstream usability. Track first-pass acceptance and correction minutes.

Q4How do you measure AI employee escalation?

Calculate recall - correct escalations divided by all cases that required escalation - and precision - correct escalations divided by all escalations raised. Preserve severity and measure packet quality, response time, and closure.

Q5How do you calculate AI employee cost per outcome?

Add platform, usage, setup, integration, review, correction, monitoring, and expected failure exposure, then divide by accepted outcomes. Do not divide by attempts or generated outputs.

Q6How many KPIs should an AI employee have?

Use the smallest set that covers outcome, quality, escalation, reliability, cost, and risk, with activity diagnostics behind it. One primary outcome KPI plus roughly one or two measures per control layer is often clearer than a large unprioritized dashboard.

Published 31 July 2026 All Hiring & onboarding →