Use an AI employee role scorecard before you configure a worker, connect a tool, or compare vendors.
The scorecard tests whether a proposed standing role has a recurring outcome, observable acceptance, coherent scope, approved context, feasible tool path, bounded authority, evaluable behavior, human ownership, and enough value to justify its operating cost.
It does not predict model accuracy from a meeting-room estimate. It identifies what is known, what must be redesigned, and what evidence a pilot must produce.
If you are still choosing the business outcome, begin with the outcome-first AI employee hiring process. If the outcome is chosen but the operating contract is incomplete, write the AI employee job description first. Then bring that document, current work samples, and process evidence into this scorecard.
On this page · 13 sectionsOpen
- The Role Scorecard at a Glance
- What Is an AI Employee Role Scorecard?
- How to Use the Scorecard
- Score Criteria 1-5: Outcome and Role Fit
- Score Criteria 6-8: Context, Tools, and Authority
- Score Criteria 9-10: Evaluation and Ownership
- Score Criteria 11-12: Economics and Operating Load
- Calculate the Decision Without Hiding the Weaknesses
- Three Worked Examples
- Copy-and-Use AI Employee Role Scorecard
- Use the Scorecard With CellCog
- Avoid Common Role-Scorecard Mistakes
- Final Recommendation
- Score the role, not the job title or vendor demo - use 12 criteria covering outcome, scope, work pattern, context, tools, authority, evaluation, ownership, economics, and operating load.
- Give each criterion 0-3 points: unknown, weak, workable, or strong - and record evidence beside every score. An unsupported opinion is not a 3.
- Apply non-compensating gates: no total score can excuse missing ownership, untestable outcomes, prohibited data, uncontrolled high-impact action, or an absent pause path.
- Treat 30-36 as a candidate for controlled testing, 23-29 as redesign, and 0-22 as a reason to choose a different role or system shape. These bands are a planning convention, not a universal benchmark.
- Re-score after the job description, evaluation set, permission design, and cost baseline improve.
- A high score authorizes a pilot decision - not production autonomy.
§ 01The Role Scorecard at a Glance
| Criterion | Core question | 0 | 1 | 2 | 3 |
|---|---|---|---|---|---|
| 1. Recurring outcome | Does one result persist across cycles? | Unknown | Broad aspiration | Outcome exists but spans several loops | One observable recurring outcome |
| 2. Accepted unit | Can a downstream user accept the result? | Unknown | Activity only | Unit exists; rubric is incomplete | Unit, schema, reviewer, rubric |
| 3. Coherent scope | Can the role state what it will not do? | Unknown | Department-wide title | Some boundaries; mixed risk/review | Bounded queue, non-goals, exceptions |
| 4. Work recurrence | Is there enough stable demand to justify a standing role? | Unknown | Rare or unpredictable | Recurs but volume/baseline is weak | Measured queue and cadence |
| 5. Agentic fit | Does the path require contextual choice? | Unknown | Fixed rules or purely human judgment | Hybrid path | Variable path with bounded judgment |
| 6. Context readiness | Are approved sources available and governable? | Unknown | Tacit/conflicting/no owner | Sources exist; gaps remain | Source hierarchy, owners, freshness |
| 7. Tool feasibility | Can the role do the work through scoped systems? | Unknown | Broad/shared access required | Feasible with integration work | Minimum action path is testable |
| 8. Authority and recovery | Can actions be limited, approved, verified, and reversed? | Unknown | Unbounded/irreversible | Some controls; gaps remain | Per-action control and recovery |
| 9. Evaluation readiness | Can representative behavior be tested? | Unknown | “Looks good” review | Rubric or cases incomplete | Cases, graders, thresholds, evidence |
| 10. Human ownership | Who decides, reviews, escalates, pauses, and retires? | Unknown | No accountable owner | Owner exists; response path weak | Named owners and response contracts |
| 11. Value and economics | Can benefit be compared with full operating cost? | Unknown | Vague productivity claim | Baseline or value estimate incomplete | Measured baseline and accepted-unit value |
| 12. Operating load | Will the role reduce total work after review and coordination? | Unknown | Creates more hidden work | Net benefit uncertain | Review, exception, and maintenance load fit |
The maximum is 36. The total is useful only after the gate review. A role with 34 points and an uncontrolled payment action is not a better candidate than a 29-point internal research role whose weaknesses can be repaired safely.
§ 02What Is an AI Employee Role Scorecard?
An AI employee role scorecard is a pre-configuration decision record. It tests whether a business responsibility is ready to become a standing, governed software role.
It sits between role discovery and implementation:
Candidate outcome, then job description, then role scorecard, then pilot contract, then onboarding.
The scorecard answers:
- Is this a coherent role?
- Is an AI employee the right system shape?
- What makes the role safe and evaluable enough to test?
- Which missing artifacts block a pilot?
- Which evidence must the pilot produce?
It should not answer: Which vendor has the longest feature list? Will the role definitely deliver ROI? Is the worker safe in every context? What production permission should it receive? Can it replace a person?
Those conclusions require platform evidence, representative evaluation, current controls, and live operating data.
Score the work system, not the model
The same model can succeed in one role and fail in another because the surrounding conditions differ: source quality; task mix; tool design; permission boundary; output schema; reviewer standard; escalation behavior; system latency; error detectability; and consequence.
The role scorecard therefore evaluates the complete assignment. A benchmark, polished answer, or one successful tool run cannot substitute for role-specific evidence.
Use a cross-functional review
The score should include the people who know the work and its consequences: the business owner; the current process owner; the downstream user; the domain reviewer; the implementation lead; the security or system owner; and privacy, legal, compliance, finance, or another qualified function when relevant.
One person can fill several roles in a small team. Preserve the decisions even if the meeting is short.
NIST’s AI RMF Core treats AI risk as contextual and calls for clear roles, multidisciplinary input, measurement, ongoing monitoring, and documented risk response. The scorecard applies that principle at the moment a proposed role is still cheap to narrow or reject.
§ 03How to Use the Scorecard
Apply the Six Gates Before Adding Points
Run the gates first. A failed gate produces Stop or Redesign, regardless of the total.
| Gate | Pass condition | Stop or redesign when |
|---|---|---|
| Accountable ownership | A named person owns scope, review, escalation, pause, and retirement | Responsibility is assigned to “the AI” or a committee with no operator |
| Observable acceptance | A downstream user can accept or reject a defined unit | Success is activity, style, or a vague goal |
| Lawful and approved inputs | Required sources and data can be used for the stated purpose | The role depends on prohibited, unavailable, or unowned data |
| Controlled consequence | High-impact actions remain human-led or have enforceable action controls | The role needs broad credentials or irreversible authority without a gate |
| Evaluation | Representative normal, edge, refusal, and failure behavior can be tested | Nobody can tell when work is correct or harmful |
| Intervention and exit | The organization can pause triggers, revoke access, preserve evidence, and retire the role | The worker can continue acting after ownership or controls fail |
Gate 1: accountable ownership
The software can produce work, maintain state, and escalate. It cannot become the legal or organizational owner of the role’s consequences.
Require names for the business outcome; daily supervision; acceptance; source policy; access; incidents; and retirement.
If nobody will review or answer escalations, the correct decision is not “more autonomy.” It is no launch.
Gate 2: observable acceptance
The role must produce a result another person or system can inspect: an accepted briefing; an approved draft; a correctly reconciled record; a resolved allowlisted case; a verified state change; or a complete escalation packet.
“Helpful,” “productive,” and “proactive” do not create a denominator.
Gate 3: lawful and approved inputs
Identify the required information before selecting a tool. Stop when the outcome depends on data the organization may not collect, disclose, combine, infer, or retain for that purpose.
Do not assume that technical access equals approved use.
Gate 4: controlled consequence
The AI employee permissions and approvals model should be feasible for every consequential action.
Some roles can remain useful in observe or prepare mode. A support employee can classify and draft while a person owns refunds, account changes, security incidents, and policy exceptions.
Gate 5: evaluation
The organization needs representative work samples, a rubric, qualified graders, and boundary cases. If expert reviewers cannot agree on a correct result, narrow the role or improve the policy before asking software to own it.
Anthropic’s agent-evaluation guidance emphasizes that agents act across turns, tools, and changing state. Score the trajectory, permissions, escalation, and durable result - not only the final prose.
Gate 6: intervention and exit
The role needs pause control; trigger disablement; credential revocation; task-state visibility; approval invalidation; evidence preservation; recovery ownership; and a retirement procedure.
A system that works only while everything is healthy is not ready for standing work.
How to Score: 0, 1, 2, or 3
Use the same scale for every criterion.
| Score | Meaning | Evidence expectation | Decision effect |
|---|---|---|---|
| 0 | Unknown or not assessed | No reliable artifact or owner answer | Treat as a blocker |
| 1 | Weak | Assertion, broad draft, or known critical gap | Redesign before a pilot |
| 2 | Workable | Partial artifact and credible repair plan | Can advance only if gates pass and gaps enter pilot contract |
| 3 | Strong | Current artifact, observed baseline, tested control, or accountable decision | Supports controlled testing |
Do not use 0 to mean “low risk.” It means the team does not know.
Add an evidence code
Record one code beside each score:
- O - Observed: current work sample, system record, measured baseline, or reproduced test.
- D - Documented: approved policy, process map, contract, rubric, or access design.
- A - Assumed: owner judgment without current evidence.
- U - Unknown: unanswered or contradictory.
A score of 3 should usually have observed or documented evidence. An assumed 3 is a warning that discussion confidence has outrun proof.
Score the current design
Do not award points for controls someone plans to create later. Score the current role, then record the artifact needed to improve it.
| Criterion | Current score | Evidence | Missing artifact | Proposed owner |
|---|---|---|---|---|
| Context readiness | 1A | Team says sources exist | Source register, hierarchy, freshness | Research lead |
| Evaluation readiness | 0U | No cases collected | 30-case representative set | Domain reviewer |
| Authority and recovery | 2D | Draft read/write matrix | Postcondition and revocation test | System owner |
This makes the scorecard a work queue rather than a ceremonial approval.
Do not average reviewers silently
If a business owner gives tool feasibility a 3 and the system owner gives it a 1, record the disagreement. Resolve which action, identity, data, or environment caused the gap.
Averaging to 2 hides the information that matters.
§ 04Score Criteria 1-5: Outcome and Role Fit
The first four criteria determine whether the proposal is a real standing role.
1. Recurring outcome
Ask: what result should still be owned after the current task ends?
Score 3 when the outcome identifies the accepted unit; the downstream user; the queue or population; the cadence or trigger; the starting scope; and the exception destination.
Examples: maintain one accepted competitor-change brief each week; prepare policy-grounded responses for three allowlisted support categories; reconcile the operating dashboard and explain material changes every Friday.
Score 1 for “manage marketing,” “help operations,” or “be an AI chief of staff.” These may be role-discovery labels, but they are not operating outcomes.
2. Accepted unit
The accepted unit should be countable, reviewable, and tied to a downstream need.
Score 3 when the role has a deliverable or state schema; a completion rule; an acceptance owner; a rubric; a rejection or correction path; and a denominator.
The accepted unit protects the economic calculation. If a worker starts 100 tasks, completes 80, and only 40 are usable, cost per task started hides the operating result.
3. Coherent scope
A coherent role shares one outcome; source hierarchy; acceptance rubric; risk class; reviewer; authority pattern; and escalation path.
Score 1 when the proposal combines research, strategic decisions, publication, customer outreach, CRM administration, and performance reporting because they all sit under “marketing.”
Score 2 when several subtypes share an outcome but need separate test cases or approvals.
Score 3 when included work, non-goals, edge cases, and handoff destinations are explicit.
4. Work recurrence
A standing role adds setup and governance overhead. It needs enough recurring demand to justify an identity; schedules or wake conditions; persistent context; task state; access review; evaluations; monitoring; and handovers.
Score 3 when current records show queue volume, cadence, seasonality, task mix, elapsed time, work time, and exceptions.
Score 1 when the work happened once or each request has a different purpose. An on-demand agent may fit better.
Use the best-tasks framework for AI employees when the role contains several possible starting tasks. That guide evaluates task suitability; this scorecard decides whether those tasks form one standing role.
5. Is the Work Actually Agentic?
An AI employee should not replace a simpler deterministic system.
| Work pattern | Default system shape | Agentic-fit score |
|---|---|---|
| Fixed trigger and exact rules | Workflow, query, or conventional software | 0-1 |
| One-off ambiguous request | On-demand agent or assistant | 1-2 |
| Recurring variable-path responsibility | AI employee candidate | 2-3 |
| Consequential professional judgment | Qualified human with AI support | 0-2 for autonomous ownership |
| Stable core with variable exceptions | Hybrid workflow + AI + human | 2-3 when boundaries are explicit |
OpenAI’s practical guide to building agents recommends agents where contextual decisions, difficult-to-maintain rules, and unstructured data make deterministic approaches inadequate. It also recommends starting incrementally rather than maximizing autonomy at the beginning.
Anthropic’s building-effective-agents guide distinguishes workflows with predefined paths from agents that direct their own process and tool use. Apply that distinction inside the role: exact calculations remain code; policy thresholds remain deterministic; permissions remain enforced controls; the agent handles bounded interpretation and next-step choice; and a person handles decisions whose consequence or authority remains human.
Score 0-1 when an API query already returns the answer; the same steps always apply; a rules engine can route the case; no persistent responsibility exists; the outcome depends entirely on human professional judgment; or the role is an excuse to avoid fixing a broken source process.
Score 2 when part of the path benefits from agentic interpretation; deterministic and human-owned segments are identified; the exact split still needs testing; and a narrow draft or prepare mode remains valuable.
Score 3 when the outcome recurs; the path changes with context; the worker must choose among approved tools or sources; the decision can be bounded and reviewed; progress must persist across cycles; and the role creates value without owning prohibited judgment.
§ 05Score Criteria 6-8: Context, Tools, and Authority
These criteria test whether the role can operate inside a real environment.
6. Context readiness
Score 3 when every required source has a purpose; an authority level; an owner; a location; an access scope; a freshness rule; conflict behavior; retention treatment; and a correction path.
Score 1 when “the shared drive” or “our CRM” is the context plan.
The AI employee memory guide explains why stored context needs provenance, freshness, correction, retention, and deletion. The scorecard does not require a complete context pack, but it does require evidence that one can be built from owned sources.
Ask the current process owner: Which documents do people actually use? Which source wins when systems disagree? Which knowledge remains tacit? Which records are stale? Which information may not enter model context? Who corrects an error?
Score 2 if the sources exist but ownership or conflict policy remains incomplete.
7. Tool feasibility
List the minimum operations required to produce the accepted unit:
| Operation | System/object | Interface | Identity | Environment | Proof needed |
|---|---|---|---|---|---|
| Read assigned requests | Research board/project | API or controlled UI | Role identity | Production read | Correct queue and fields |
| Retrieve approved sources | Source registry/domains | Browser/API | Role identity | Read | Scope and provenance |
| Create draft | Project folder | Document interface | Role identity | Staging | Correct destination |
| Update state | Role task | API/UI | Role identity | Production | State and actor recorded |
| Request approval | Approval queue | Workflow | Role identity | Production | Exact action and expiry |
Score 3 when the minimum path can be tested with a distinct identity and scoped objects.
Score 1 when a person must share broad credentials; the system exposes no appropriate object or action boundary; expected state cannot be verified; failures cannot be detected; required interfaces are too brittle for the role; or the integration effort exceeds the value.
Computer use can extend reach into interfaces without APIs, but it does not remove the need for access, action, evidence, and recovery controls.
8. Authority and recovery
Score each action separately:
| Authority state | Meaning |
|---|---|
| Blocked | The role cannot perform the action |
| Observe | It can read approved data |
| Prepare | It can create a proposed output without external effect |
| Act after approval | It can execute the exact approved action |
| Act within policy | It can execute an allowlisted, limited, tested action |
Score 3 when the job description identifies the action and target; data scope; conditions; a volume, spend, time, or frequency limit; approval; audit evidence; the postcondition; recovery; and revocation.
Score 1 when the role needs “full access” or when a prompt is the only restriction.
High-impact, irreversible, sensitive, public, financial, security, legal, employment, healthcare, or rights-affecting work should remain human-led unless qualified owners approve a defensible, enforceable design. A high score in other dimensions does not override that boundary.
§ 06Score Criteria 9-10: Evaluation and Ownership
9. Evaluation readiness
Score 3 when the team can build a representative evaluation set that includes normal cases; subtypes and segments; boundary cases; missing inputs; stale and conflicting sources; out-of-scope requests; unsafe or adversarial instructions; tool failure; approval expiry; required escalation; required refusal; and a plausible high-severity failure.
The rubric should judge the outcome; evidence; accuracy; completeness; policy; authority; escalation; durable state; format; and downstream usability.
An evaluation also needs a grader. Name the person or deterministic check that can judge each criterion.
Score 2 when cases exist but cover only the happy path, or when the rubric has not separated cosmetic correction from material failure.
Score 1 when “we will know it when we see it” is the review plan.
Test the trajectory
An agent can arrive at acceptable prose after reading a prohibited source; exposing sensitive content; taking an unauthorized action; retrying excessively; creating duplicate work; or leaving the system in the wrong state.
Score tool choices, permissions, intermediate state, escalation, and postconditions as well as the final artifact.
10. Human ownership
Score 3 when the role names an owner for every decision:
| Decision | Required owner |
|---|---|
| Business outcome and scope | Business owner |
| Daily queue and escalation | Supervisor |
| Acceptance and correction | Domain/downstream reviewer |
| Source authority | Knowledge or policy owner |
| Access grant and revocation | System/security owner |
| High-impact approval | Authorized decision owner |
| Incident containment and recovery | Incident owner |
| Expansion, pause, and retirement | Business owner with required reviewers |
Add response expectations. An escalation queue with no service expectation can leave urgent work unresolved or encourage the worker to infer permission from silence.
Score 2 when names exist but backups, decision authority, or response behavior is unclear.
Score 1 when ownership is a department name.
Microsoft’s current responsible-agent guidance recommends deciding grounding sources, allowed actions, and human-approval points early and using risk-scaled release gates. The scorecard turns those decisions into pre-hire evidence.
§ 07Score Criteria 11-12: Economics and Operating Load
11. Value and economics
Estimate value against the complete operating cost.
Start with a baseline:
| Baseline field | What to record |
|---|---|
| Queue | Cases per period and subtype mix |
| Human work | Active minutes per case |
| Elapsed time | Intake to accepted outcome |
| Quality | Acceptance, correction, reopen, severe error |
| Cost | Labor, software, data, contractors, delay |
| Constraint | Backlog, skill bottleneck, response time, forgone capacity |
| Outcome value | Avoided cost, enabled capacity, protected value, or incremental contribution |
Then estimate: expected role value = accepted outcomes multiplied by value per accepted outcome. Expected role cost = platform + usage + data + integration + reviewer + correction + supervision + incident allowance.
Score 3 when the baseline uses current records and the value unit does not depend on unsupported attribution.
Score 2 when the role clearly relieves a constraint but the baseline or price model remains incomplete.
Score 1 when the business case is “AI is cheaper than a salary.” An AI employee is software capacity, not a like-for-like legal or organizational substitute for a person.
Use the AI employee total-cost framework to build the full ledger. The pre-hire score should expose assumptions rather than pretend the estimate is a realized ROI.
OpenAI’s use-case prioritization guidance uses impact and effort to separate quick wins from high-effort, low-impact ideas. The scorecard goes one level deeper by requiring an accepted unit, controls, and ongoing operating cost.
12. Operating load
The role creates human work: preparing and maintaining sources; reviewing outputs; answering escalations; approving actions; correcting errors; monitoring cost and KPIs; managing access; investigating incidents; updating evaluations; and changing the role when the business changes.
Score 3 when the organization has estimated review minutes per accepted outcome; escalation rate and response time; correction burden; access and source maintenance; peak queue; backup coverage; incident workload; and change cadence.
Score 1 when the proposal assumes “the AI handles everything.”
Include coordination tax
If the role hands work to people or other agents, measure assignments rejected; clarification loops; duplicate tasks; waiting time; handover correction; unresolved ownership; and final integration work.
A role can complete its own output while increasing total system work. Score the net operating effect.
§ 08Calculate the Decision Without Hiding the Weaknesses
After the gate review:
- Score all 12 criteria from 0-3.
- Add an evidence code to each.
- Record disagreements and missing artifacts.
- Add the points.
- Apply the planning band.
- Inspect the lowest criteria and worst plausible failure.
- Make one decision: Test, Redesign, Choose another system shape, or Stop.
Planning bands
| Total | Default interpretation | Next move |
|---|---|---|
| 30-36 | Strong candidate for a bounded test | Write a pilot contract; do not expand authority |
| 23-29 | Promising but not ready | Repair the lowest criteria, then re-score |
| 0-22 | Poor standing-role candidate today | Narrow, use workflow/assistant/human path, or stop |
These bands are an internal planning convention. They are not research-validated thresholds, safety certifications, or guarantees.
Add decision rules
Use rules that the total cannot override:
- Any failed gate: Stop or Redesign.
- Any unknown in authority, ownership, evaluation, or required data: do not configure live action.
- Any score of 1: name the repair before pilot approval.
- High score with low-confidence evidence: gather evidence before relying on it.
- Severe failure outside tolerance: narrow or stop even if averages pass.
- Net operating load greater than plausible value: choose a different design.
Preserve the shape of the score
Two roles can both score 28: one may have strong outcome and evaluation but weak integration; another may have easy tools but no coherent scope.
Do not compare totals without the criterion profile. The weak dimensions determine the repair.
§ 09Three Worked Examples
Worked Example 1: Weekly Competitor-Change Brief
The proposed role monitors an approved 12-company registry and produces one cited change brief each Monday for the strategy lead.
| Criterion | Score | Evidence | Note |
|---|---|---|---|
| Recurring outcome | 3D | Approved weekly brief charter | One outcome and audience |
| Accepted unit | 3D | Current brief schema and reviewer | Acceptance rubric exists |
| Coherent scope | 3D | Company/change allowlist | No publication or outreach |
| Work recurrence | 3O | 16 weeks of current briefs | Stable cadence and subtype history |
| Agentic fit | 3D | Variable sources and materiality rubric | Agent interprets; human approves strategy |
| Context readiness | 2D | Source registry exists | Freshness owner missing for 2 sources |
| Tool feasibility | 3D | Read-only browser + project folder | Minimum path testable |
| Authority and recovery | 3D | Draft-only; task-state update | No external communication |
| Evaluation readiness | 2D | 24 cases | Needs conflicting-source and injection cases |
| Human ownership | 3D | Research lead and backup | Response contract defined |
| Value and economics | 2O | Work-time baseline | Usage and reviewer cost need pilot data |
| Operating load | 2A | Review estimate | No observed AI correction burden yet |
All six gates pass. The decision is Test, not deploy autonomously.
The pilot contract should require source-owner completion; adverse cases; reviewer minutes; the factual-correction rate; escalation quality; cost per accepted brief; and zero publication or external contact.
Worked Example 2: “AI Marketing Manager”
The proposal asks one worker to research markets, write strategy, produce content, publish, update CRM, run outbound email, and report revenue.
| Criterion | Score | Main weakness |
|---|---|---|
| Recurring outcome | 1 | Several outcomes hidden under a title |
| Accepted unit | 1 | No single downstream acceptance |
| Coherent scope | 0 | Mixed sources, risk, reviewers, and authority |
| Work recurrence | 2 | Tasks recur, but no measured queue |
| Agentic fit | 2 | Some variable work; some fixed automation |
| Context readiness | 1 | Broad “all marketing docs” plan |
| Tool feasibility | 1 | CRM, CMS, email, analytics, and browser access |
| Authority and recovery | 0 | Publication and outreach boundary absent |
| Evaluation readiness | 1 | Output review only |
| Human ownership | 1 | “Marketing team” |
| Value and economics | 1 | Salary-substitution claim |
| Operating load | 0 | No review or coordination estimate |
The total is 11/36, and it fails the scope, consequence, evaluation, and ownership gates. The decision is Redesign.
Split the proposal into candidate outcomes: a weekly market-evidence brief; an approved content-outline queue; a CRM enrichment draft; campaign-performance reconciliation; and allowlisted email-draft preparation.
Score each separately. Some may fit workflow automation; some may fit an on-demand assistant; one or two may justify a standing role.
Worked Example 3: Support Triage and Refund Resolution
The initial request asks the role to classify tickets, answer customers, change accounts, and issue refunds.
The team redesigns it: classify three low-risk categories; retrieve the current policy; prepare a response draft; request approval for every send; route refunds, security issues, account changes, and exceptions to people; and update only the internal task state.
| Criterion group | Score | Why |
|---|---|---|
| Outcome and scope | 10/12 | Accepted draft/triage unit; queue baseline still developing |
| Agentic fit | 3/3 | Variable language and policy-grounded routing |
| Context, tools, authority | 7/9 | Policy sources ready; approval interface needs expiry test |
| Evaluation and ownership | 5/6 | Reviewers named; adverse set needs more security cases |
| Economics and load | 4/6 | Value plausible; correction and escalation cost unknown |
The narrower role may be useful without autonomous refund or customer-send authority. The scorecard makes that hybrid design visible.
§ 10Copy-and-Use AI Employee Role Scorecard
A. Role identity. Proposed role; version/date; business owner; supervisor; downstream user/reviewer; current process owner; proposed system shape (workflow/assistant/AI employee/hybrid/human); one-sentence outcome (outcome, queue, cadence, sources, escalation).
B. Gate review. For each of the six gates: pass/fail; evidence; owner; required repair. Then the gate decision: Pass, Redesign, or Stop.
C. Twelve-criterion score. For each criterion: score 0-3; evidence code O/D/A/U; the evidence or gap; owner; due date. Then the total out of 36.
D. Decision. The decision (Test/Redesign/Workflow/Assistant/Human/Stop); failed or conditional gates; the three lowest criteria; the worst plausible failure; human-only decisions; the required pilot mode (observe/shadow/draft/approved action); the evidence the pilot must produce; maximum initial authority; the pause condition; the re-score date or trigger; and approvers.
§ 11Use the Scorecard With CellCog
CellCog’s AI Employees product page describes persistent roles with goals, KPIs, an inbox, task board, memory, schedules, wake conditions, permissions, approvals, shifts, and handovers around a general-purpose agent.
The scorecard should decide which of those capabilities the role needs:
| Scorecard evidence | CellCog capability to verify |
|---|---|
| Recurring outcome and accepted unit | Role, goals, KPIs, dashboard artifacts |
| Queue and trigger | Inbox, task list, schedule, wake condition |
| Approved sources | Context and memory configuration |
| Minimum tool path | Browse/Cowork and connected systems |
| Authority boundary | Permissions and approvals |
| Continuity | Task state, shifts, and handovers |
| Evaluation and ownership | Observable work, supervisor review, pause behavior |
Do not raise a score because the platform advertises a feature. Ask for proof against the exact role: Can the source scope be limited? Can actions be separated into read, prepare, approved, and blocked states? Does approval bind to the exact action? Can the worker preserve evidence and task state? Can a supervisor pause work and revoke access? Can output, correction, escalation, and cost be measured?
Use CellCog’s role pages as examples, not pre-approved job descriptions. A title such as AI Research Assistant or AI Operations Manager becomes a real candidate only after the organization defines its outcome, queue, inputs, authority, and acceptance.
After a role passes the gates and reaches a test-ready score, use the AI employee pilot framework to precommit the sample, modes, metrics, budget, failure rules, and go/redesign/stop decision. Then use graduated onboarding to increase context and authority only when the evidence gate passes.
§ 12Avoid Common Role-Scorecard Mistakes
Scoring the title. “AI executive assistant” may describe ten different outcomes. Score the actual queue, deliverable, sources, tools, and authority.
Treating unknown as average. An unanswered permission or data question is a 0, not a 2. Unknowns become surprises after connection.
Letting a high total cancel a failed gate. Value and recurrence cannot compensate for prohibited data, missing ownership, or uncontrolled consequence.
Awarding points for planned artifacts. “We will write a rubric” is a gap. Create it, review it, and re-score.
Starting from vendor capability. “The platform can browse and email” does not prove that your role has approved sources or sending authority.
Using precise bands as scientific truth. The thresholds organize an internal decision. Calibrate them with portfolio evidence over time, and keep non-compensating gates.
Ignoring the current alternative. Compare the role with workflow automation, an on-demand assistant, process repair, another specialist product, or human-led work. “Do nothing” is not the only alternative.
Hiding reviewer disagreement. Disagreement about consequence, source authority, or correctness is evidence. Resolve it or narrow the scope.
Equating a high score with production readiness. The scorecard approves the next learning step. Evaluation, pilot, security review, onboarding, monitoring, and change control still remain.
§ 13Final Recommendation
Use the scorecard to force an explicit pre-hire decision:
- Pass the six gates.
- Score the current role across 12 criteria.
- Attach evidence and owners.
- Identify the lowest dimensions.
- Compare alternative system shapes.
- Define the maximum safe pilot mode.
- Precommit the evidence needed to proceed.
The best first role is not the one with the grandest title. It is the one whose outcome matters, boundaries hold, evidence can be tested, human ownership is real, and total operating work can plausibly fall.
Q1What is an AI employee role scorecard?
It is a pre-configuration decision tool that tests whether a proposed standing AI role has a recurring outcome, observable acceptance, coherent scope, approved context, feasible tools, bounded authority, evaluable behavior, accountable owners, and a credible economic case.
Q2What is a good score for an AI employee role?
As a planning convention, 30-36 can indicate a candidate for bounded testing, 23-29 calls for redesign, and 0-22 suggests another role or system shape. Any failed gate overrides the total. These bands are not universal or safety-validated thresholds.
Q3Should a role with a high score go live?
No. A high score supports a controlled pilot decision. The role still needs representative evaluation, approved access, evidence-gated onboarding, monitoring, incident response, and change control before any bounded production authority.
Q4Who should complete the scorecard?
Include the business owner, current process owner, downstream reviewer, implementation lead, and the owners of relevant systems, security, privacy, legal, compliance, finance, or domain policy. One person may cover several functions in a small team, but every decision needs an accountable owner.
Q5How often should the role be re-scored?
Re-score after material changes to the outcome, task mix, source set, retained context, model, tools, permissions, approval rules, output destination, consequence, pricing, or reviewer capacity. Also re-score after the pilot replaces assumptions with observed evidence.
Q6Can the scorecard be used for any AI agent?
Most criteria apply to agents generally. The recurrence, standing ownership, continuity, task-state, schedule, and handover criteria are especially important when the agent is being placed into an ongoing AI employee role rather than used for a single request.
