Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentPricingBlogStoryContact

AI Employee Role Scorecard: A Pre-Hire Template

Napkin-style sketch of a pre-hire scorecard sheet with six gate checkboxes above twelve scored criterion rows and a total band
Fig 0Gates first, points second: no total score can excuse missing ownership, untestable outcomes, or uncontrolled consequence.

Use an AI employee role scorecard before you configure a worker, connect a tool, or compare vendors.

The scorecard tests whether a proposed standing role has a recurring outcome, observable acceptance, coherent scope, approved context, feasible tool path, bounded authority, evaluable behavior, human ownership, and enough value to justify its operating cost.

It does not predict model accuracy from a meeting-room estimate. It identifies what is known, what must be redesigned, and what evidence a pilot must produce.

If you are still choosing the business outcome, begin with the outcome-first AI employee hiring process. If the outcome is chosen but the operating contract is incomplete, write the AI employee job description first. Then bring that document, current work samples, and process evidence into this scorecard.

On this page · 13 sectionsOpen
  1. The Role Scorecard at a Glance
  2. What Is an AI Employee Role Scorecard?
  3. How to Use the Scorecard
  4. Score Criteria 1-5: Outcome and Role Fit
  5. Score Criteria 6-8: Context, Tools, and Authority
  6. Score Criteria 9-10: Evaluation and Ownership
  7. Score Criteria 11-12: Economics and Operating Load
  8. Calculate the Decision Without Hiding the Weaknesses
  9. Three Worked Examples
  10. Copy-and-Use AI Employee Role Scorecard
  11. Use the Scorecard With CellCog
  12. Avoid Common Role-Scorecard Mistakes
  13. Final Recommendation
Key points6 · 23 min full read
  1. Score the role, not the job title or vendor demo - use 12 criteria covering outcome, scope, work pattern, context, tools, authority, evaluation, ownership, economics, and operating load.
  2. Give each criterion 0-3 points: unknown, weak, workable, or strong - and record evidence beside every score. An unsupported opinion is not a 3.
  3. Apply non-compensating gates: no total score can excuse missing ownership, untestable outcomes, prohibited data, uncontrolled high-impact action, or an absent pause path.
  4. Treat 30-36 as a candidate for controlled testing, 23-29 as redesign, and 0-22 as a reason to choose a different role or system shape. These bands are a planning convention, not a universal benchmark.
  5. Re-score after the job description, evaluation set, permission design, and cost baseline improve.
  6. A high score authorizes a pilot decision - not production autonomy.

§ 01The Role Scorecard at a Glance

Criterion Core question 0 1 2 3
1. Recurring outcome Does one result persist across cycles? Unknown Broad aspiration Outcome exists but spans several loops One observable recurring outcome
2. Accepted unit Can a downstream user accept the result? Unknown Activity only Unit exists; rubric is incomplete Unit, schema, reviewer, rubric
3. Coherent scope Can the role state what it will not do? Unknown Department-wide title Some boundaries; mixed risk/review Bounded queue, non-goals, exceptions
4. Work recurrence Is there enough stable demand to justify a standing role? Unknown Rare or unpredictable Recurs but volume/baseline is weak Measured queue and cadence
5. Agentic fit Does the path require contextual choice? Unknown Fixed rules or purely human judgment Hybrid path Variable path with bounded judgment
6. Context readiness Are approved sources available and governable? Unknown Tacit/conflicting/no owner Sources exist; gaps remain Source hierarchy, owners, freshness
7. Tool feasibility Can the role do the work through scoped systems? Unknown Broad/shared access required Feasible with integration work Minimum action path is testable
8. Authority and recovery Can actions be limited, approved, verified, and reversed? Unknown Unbounded/irreversible Some controls; gaps remain Per-action control and recovery
9. Evaluation readiness Can representative behavior be tested? Unknown “Looks good” review Rubric or cases incomplete Cases, graders, thresholds, evidence
10. Human ownership Who decides, reviews, escalates, pauses, and retires? Unknown No accountable owner Owner exists; response path weak Named owners and response contracts
11. Value and economics Can benefit be compared with full operating cost? Unknown Vague productivity claim Baseline or value estimate incomplete Measured baseline and accepted-unit value
12. Operating load Will the role reduce total work after review and coordination? Unknown Creates more hidden work Net benefit uncertain Review, exception, and maintenance load fit
Scroll to compare all columns
Table 1The 12 criteria, scored 0 to 3 each

The maximum is 36. The total is useful only after the gate review. A role with 34 points and an uncontrolled payment action is not a better candidate than a 29-point internal research role whose weaknesses can be repaired safely.

§ 02What Is an AI Employee Role Scorecard?

An AI employee role scorecard is a pre-configuration decision record. It tests whether a business responsibility is ready to become a standing, governed software role.

It sits between role discovery and implementation:

Candidate outcome, then job description, then role scorecard, then pilot contract, then onboarding.

The scorecard answers:

  1. Is this a coherent role?
  2. Is an AI employee the right system shape?
  3. What makes the role safe and evaluable enough to test?
  4. Which missing artifacts block a pilot?
  5. Which evidence must the pilot produce?

It should not answer: Which vendor has the longest feature list? Will the role definitely deliver ROI? Is the worker safe in every context? What production permission should it receive? Can it replace a person?

Those conclusions require platform evidence, representative evaluation, current controls, and live operating data.

Score the work system, not the model

The same model can succeed in one role and fail in another because the surrounding conditions differ: source quality; task mix; tool design; permission boundary; output schema; reviewer standard; escalation behavior; system latency; error detectability; and consequence.

The role scorecard therefore evaluates the complete assignment. A benchmark, polished answer, or one successful tool run cannot substitute for role-specific evidence.

Use a cross-functional review

The score should include the people who know the work and its consequences: the business owner; the current process owner; the downstream user; the domain reviewer; the implementation lead; the security or system owner; and privacy, legal, compliance, finance, or another qualified function when relevant.

One person can fill several roles in a small team. Preserve the decisions even if the meeting is short.

NIST’s AI RMF Core treats AI risk as contextual and calls for clear roles, multidisciplinary input, measurement, ongoing monitoring, and documented risk response. The scorecard applies that principle at the moment a proposed role is still cheap to narrow or reject.

§ 03How to Use the Scorecard

Apply the Six Gates Before Adding Points

Run the gates first. A failed gate produces Stop or Redesign, regardless of the total.

Gate Pass condition Stop or redesign when
Accountable ownership A named person owns scope, review, escalation, pause, and retirement Responsibility is assigned to “the AI” or a committee with no operator
Observable acceptance A downstream user can accept or reject a defined unit Success is activity, style, or a vague goal
Lawful and approved inputs Required sources and data can be used for the stated purpose The role depends on prohibited, unavailable, or unowned data
Controlled consequence High-impact actions remain human-led or have enforceable action controls The role needs broad credentials or irreversible authority without a gate
Evaluation Representative normal, edge, refusal, and failure behavior can be tested Nobody can tell when work is correct or harmful
Intervention and exit The organization can pause triggers, revoke access, preserve evidence, and retire the role The worker can continue acting after ownership or controls fail
Table 2The six non-compensating gates

Gate 1: accountable ownership

The software can produce work, maintain state, and escalate. It cannot become the legal or organizational owner of the role’s consequences.

Require names for the business outcome; daily supervision; acceptance; source policy; access; incidents; and retirement.

If nobody will review or answer escalations, the correct decision is not “more autonomy.” It is no launch.

Gate 2: observable acceptance

The role must produce a result another person or system can inspect: an accepted briefing; an approved draft; a correctly reconciled record; a resolved allowlisted case; a verified state change; or a complete escalation packet.

“Helpful,” “productive,” and “proactive” do not create a denominator.

Gate 3: lawful and approved inputs

Identify the required information before selecting a tool. Stop when the outcome depends on data the organization may not collect, disclose, combine, infer, or retain for that purpose.

Do not assume that technical access equals approved use.

Gate 4: controlled consequence

The AI employee permissions and approvals model should be feasible for every consequential action.

Some roles can remain useful in observe or prepare mode. A support employee can classify and draft while a person owns refunds, account changes, security incidents, and policy exceptions.

Gate 5: evaluation

The organization needs representative work samples, a rubric, qualified graders, and boundary cases. If expert reviewers cannot agree on a correct result, narrow the role or improve the policy before asking software to own it.

Anthropic’s agent-evaluation guidance emphasizes that agents act across turns, tools, and changing state. Score the trajectory, permissions, escalation, and durable result - not only the final prose.

Gate 6: intervention and exit

The role needs pause control; trigger disablement; credential revocation; task-state visibility; approval invalidation; evidence preservation; recovery ownership; and a retirement procedure.

A system that works only while everything is healthy is not ready for standing work.

How to Score: 0, 1, 2, or 3

Use the same scale for every criterion.

Score Meaning Evidence expectation Decision effect
0 Unknown or not assessed No reliable artifact or owner answer Treat as a blocker
1 Weak Assertion, broad draft, or known critical gap Redesign before a pilot
2 Workable Partial artifact and credible repair plan Can advance only if gates pass and gaps enter pilot contract
3 Strong Current artifact, observed baseline, tested control, or accountable decision Supports controlled testing
Table 3The scoring scale and its decision effect

Do not use 0 to mean “low risk.” It means the team does not know.

Add an evidence code

Record one code beside each score:

  • O - Observed: current work sample, system record, measured baseline, or reproduced test.
  • D - Documented: approved policy, process map, contract, rubric, or access design.
  • A - Assumed: owner judgment without current evidence.
  • U - Unknown: unanswered or contradictory.

A score of 3 should usually have observed or documented evidence. An assumed 3 is a warning that discussion confidence has outrun proof.

Score the current design

Do not award points for controls someone plans to create later. Score the current role, then record the artifact needed to improve it.

Criterion Current score Evidence Missing artifact Proposed owner
Context readiness 1A Team says sources exist Source register, hierarchy, freshness Research lead
Evaluation readiness 0U No cases collected 30-case representative set Domain reviewer
Authority and recovery 2D Draft read/write matrix Postcondition and revocation test System owner
Table 4Example: turning low scores into a work queue

This makes the scorecard a work queue rather than a ceremonial approval.

Do not average reviewers silently

If a business owner gives tool feasibility a 3 and the system owner gives it a 1, record the disagreement. Resolve which action, identity, data, or environment caused the gap.

Averaging to 2 hides the information that matters.

§ 04Score Criteria 1-5: Outcome and Role Fit

The first four criteria determine whether the proposal is a real standing role.

1. Recurring outcome

Ask: what result should still be owned after the current task ends?

Score 3 when the outcome identifies the accepted unit; the downstream user; the queue or population; the cadence or trigger; the starting scope; and the exception destination.

Examples: maintain one accepted competitor-change brief each week; prepare policy-grounded responses for three allowlisted support categories; reconcile the operating dashboard and explain material changes every Friday.

Score 1 for “manage marketing,” “help operations,” or “be an AI chief of staff.” These may be role-discovery labels, but they are not operating outcomes.

2. Accepted unit

The accepted unit should be countable, reviewable, and tied to a downstream need.

Score 3 when the role has a deliverable or state schema; a completion rule; an acceptance owner; a rubric; a rejection or correction path; and a denominator.

The accepted unit protects the economic calculation. If a worker starts 100 tasks, completes 80, and only 40 are usable, cost per task started hides the operating result.

3. Coherent scope

A coherent role shares one outcome; source hierarchy; acceptance rubric; risk class; reviewer; authority pattern; and escalation path.

Score 1 when the proposal combines research, strategic decisions, publication, customer outreach, CRM administration, and performance reporting because they all sit under “marketing.”

Score 2 when several subtypes share an outcome but need separate test cases or approvals.

Score 3 when included work, non-goals, edge cases, and handoff destinations are explicit.

4. Work recurrence

A standing role adds setup and governance overhead. It needs enough recurring demand to justify an identity; schedules or wake conditions; persistent context; task state; access review; evaluations; monitoring; and handovers.

Score 3 when current records show queue volume, cadence, seasonality, task mix, elapsed time, work time, and exceptions.

Score 1 when the work happened once or each request has a different purpose. An on-demand agent may fit better.

Use the best-tasks framework for AI employees when the role contains several possible starting tasks. That guide evaluates task suitability; this scorecard decides whether those tasks form one standing role.

5. Is the Work Actually Agentic?

An AI employee should not replace a simpler deterministic system.

Work pattern Default system shape Agentic-fit score
Fixed trigger and exact rules Workflow, query, or conventional software 0-1
One-off ambiguous request On-demand agent or assistant 1-2
Recurring variable-path responsibility AI employee candidate 2-3
Consequential professional judgment Qualified human with AI support 0-2 for autonomous ownership
Stable core with variable exceptions Hybrid workflow + AI + human 2-3 when boundaries are explicit
Table 5Work patterns and their default system shapes

OpenAI’s practical guide to building agents recommends agents where contextual decisions, difficult-to-maintain rules, and unstructured data make deterministic approaches inadequate. It also recommends starting incrementally rather than maximizing autonomy at the beginning.

Anthropic’s building-effective-agents guide distinguishes workflows with predefined paths from agents that direct their own process and tool use. Apply that distinction inside the role: exact calculations remain code; policy thresholds remain deterministic; permissions remain enforced controls; the agent handles bounded interpretation and next-step choice; and a person handles decisions whose consequence or authority remains human.

Score 0-1 when an API query already returns the answer; the same steps always apply; a rules engine can route the case; no persistent responsibility exists; the outcome depends entirely on human professional judgment; or the role is an excuse to avoid fixing a broken source process.

Score 2 when part of the path benefits from agentic interpretation; deterministic and human-owned segments are identified; the exact split still needs testing; and a narrow draft or prepare mode remains valuable.

Score 3 when the outcome recurs; the path changes with context; the worker must choose among approved tools or sources; the decision can be bounded and reviewed; progress must persist across cycles; and the role creates value without owning prohibited judgment.

§ 05Score Criteria 6-8: Context, Tools, and Authority

These criteria test whether the role can operate inside a real environment.

6. Context readiness

Score 3 when every required source has a purpose; an authority level; an owner; a location; an access scope; a freshness rule; conflict behavior; retention treatment; and a correction path.

Score 1 when “the shared drive” or “our CRM” is the context plan.

The AI employee memory guide explains why stored context needs provenance, freshness, correction, retention, and deletion. The scorecard does not require a complete context pack, but it does require evidence that one can be built from owned sources.

Ask the current process owner: Which documents do people actually use? Which source wins when systems disagree? Which knowledge remains tacit? Which records are stale? Which information may not enter model context? Who corrects an error?

Score 2 if the sources exist but ownership or conflict policy remains incomplete.

7. Tool feasibility

List the minimum operations required to produce the accepted unit:

Operation System/object Interface Identity Environment Proof needed
Read assigned requests Research board/project API or controlled UI Role identity Production read Correct queue and fields
Retrieve approved sources Source registry/domains Browser/API Role identity Read Scope and provenance
Create draft Project folder Document interface Role identity Staging Correct destination
Update state Role task API/UI Role identity Production State and actor recorded
Request approval Approval queue Workflow Role identity Production Exact action and expiry
Scroll to compare all columns
Table 6The minimum tool path: one row per operation

Score 3 when the minimum path can be tested with a distinct identity and scoped objects.

Score 1 when a person must share broad credentials; the system exposes no appropriate object or action boundary; expected state cannot be verified; failures cannot be detected; required interfaces are too brittle for the role; or the integration effort exceeds the value.

Computer use can extend reach into interfaces without APIs, but it does not remove the need for access, action, evidence, and recovery controls.

8. Authority and recovery

Score each action separately:

Authority state Meaning
Blocked The role cannot perform the action
Observe It can read approved data
Prepare It can create a proposed output without external effect
Act after approval It can execute the exact approved action
Act within policy It can execute an allowlisted, limited, tested action
Table 7Authority states per action

Score 3 when the job description identifies the action and target; data scope; conditions; a volume, spend, time, or frequency limit; approval; audit evidence; the postcondition; recovery; and revocation.

Score 1 when the role needs “full access” or when a prompt is the only restriction.

High-impact, irreversible, sensitive, public, financial, security, legal, employment, healthcare, or rights-affecting work should remain human-led unless qualified owners approve a defensible, enforceable design. A high score in other dimensions does not override that boundary.

§ 06Score Criteria 9-10: Evaluation and Ownership

9. Evaluation readiness

Score 3 when the team can build a representative evaluation set that includes normal cases; subtypes and segments; boundary cases; missing inputs; stale and conflicting sources; out-of-scope requests; unsafe or adversarial instructions; tool failure; approval expiry; required escalation; required refusal; and a plausible high-severity failure.

The rubric should judge the outcome; evidence; accuracy; completeness; policy; authority; escalation; durable state; format; and downstream usability.

An evaluation also needs a grader. Name the person or deterministic check that can judge each criterion.

Score 2 when cases exist but cover only the happy path, or when the rubric has not separated cosmetic correction from material failure.

Score 1 when “we will know it when we see it” is the review plan.

Test the trajectory

An agent can arrive at acceptable prose after reading a prohibited source; exposing sensitive content; taking an unauthorized action; retrying excessively; creating duplicate work; or leaving the system in the wrong state.

Score tool choices, permissions, intermediate state, escalation, and postconditions as well as the final artifact.

10. Human ownership

Score 3 when the role names an owner for every decision:

Decision Required owner
Business outcome and scope Business owner
Daily queue and escalation Supervisor
Acceptance and correction Domain/downstream reviewer
Source authority Knowledge or policy owner
Access grant and revocation System/security owner
High-impact approval Authorized decision owner
Incident containment and recovery Incident owner
Expansion, pause, and retirement Business owner with required reviewers
Table 8Decision ownership: one named owner per decision

Add response expectations. An escalation queue with no service expectation can leave urgent work unresolved or encourage the worker to infer permission from silence.

Score 2 when names exist but backups, decision authority, or response behavior is unclear.

Score 1 when ownership is a department name.

Microsoft’s current responsible-agent guidance recommends deciding grounding sources, allowed actions, and human-approval points early and using risk-scaled release gates. The scorecard turns those decisions into pre-hire evidence.

§ 07Score Criteria 11-12: Economics and Operating Load

11. Value and economics

Estimate value against the complete operating cost.

Start with a baseline:

Baseline field What to record
Queue Cases per period and subtype mix
Human work Active minutes per case
Elapsed time Intake to accepted outcome
Quality Acceptance, correction, reopen, severe error
Cost Labor, software, data, contractors, delay
Constraint Backlog, skill bottleneck, response time, forgone capacity
Outcome value Avoided cost, enabled capacity, protected value, or incremental contribution
Table 9The pre-hire baseline: what to record before scoring value

Then estimate: expected role value = accepted outcomes multiplied by value per accepted outcome. Expected role cost = platform + usage + data + integration + reviewer + correction + supervision + incident allowance.

Score 3 when the baseline uses current records and the value unit does not depend on unsupported attribution.

Score 2 when the role clearly relieves a constraint but the baseline or price model remains incomplete.

Score 1 when the business case is “AI is cheaper than a salary.” An AI employee is software capacity, not a like-for-like legal or organizational substitute for a person.

Use the AI employee total-cost framework to build the full ledger. The pre-hire score should expose assumptions rather than pretend the estimate is a realized ROI.

OpenAI’s use-case prioritization guidance uses impact and effort to separate quick wins from high-effort, low-impact ideas. The scorecard goes one level deeper by requiring an accepted unit, controls, and ongoing operating cost.

12. Operating load

The role creates human work: preparing and maintaining sources; reviewing outputs; answering escalations; approving actions; correcting errors; monitoring cost and KPIs; managing access; investigating incidents; updating evaluations; and changing the role when the business changes.

Score 3 when the organization has estimated review minutes per accepted outcome; escalation rate and response time; correction burden; access and source maintenance; peak queue; backup coverage; incident workload; and change cadence.

Score 1 when the proposal assumes “the AI handles everything.”

Include coordination tax

If the role hands work to people or other agents, measure assignments rejected; clarification loops; duplicate tasks; waiting time; handover correction; unresolved ownership; and final integration work.

A role can complete its own output while increasing total system work. Score the net operating effect.

§ 08Calculate the Decision Without Hiding the Weaknesses

After the gate review:

  1. Score all 12 criteria from 0-3.
  2. Add an evidence code to each.
  3. Record disagreements and missing artifacts.
  4. Add the points.
  5. Apply the planning band.
  6. Inspect the lowest criteria and worst plausible failure.
  7. Make one decision: Test, Redesign, Choose another system shape, or Stop.

Planning bands

Total Default interpretation Next move
30-36 Strong candidate for a bounded test Write a pilot contract; do not expand authority
23-29 Promising but not ready Repair the lowest criteria, then re-score
0-22 Poor standing-role candidate today Narrow, use workflow/assistant/human path, or stop
Table 10Planning bands: what the total suggests

These bands are an internal planning convention. They are not research-validated thresholds, safety certifications, or guarantees.

Add decision rules

Use rules that the total cannot override:

  • Any failed gate: Stop or Redesign.
  • Any unknown in authority, ownership, evaluation, or required data: do not configure live action.
  • Any score of 1: name the repair before pilot approval.
  • High score with low-confidence evidence: gather evidence before relying on it.
  • Severe failure outside tolerance: narrow or stop even if averages pass.
  • Net operating load greater than plausible value: choose a different design.

Preserve the shape of the score

Two roles can both score 28: one may have strong outcome and evaluation but weak integration; another may have easy tools but no coherent scope.

Do not compare totals without the criterion profile. The weak dimensions determine the repair.

§ 09Three Worked Examples

Worked Example 1: Weekly Competitor-Change Brief

The proposed role monitors an approved 12-company registry and produces one cited change brief each Monday for the strategy lead.

Criterion Score Evidence Note
Recurring outcome 3D Approved weekly brief charter One outcome and audience
Accepted unit 3D Current brief schema and reviewer Acceptance rubric exists
Coherent scope 3D Company/change allowlist No publication or outreach
Work recurrence 3O 16 weeks of current briefs Stable cadence and subtype history
Agentic fit 3D Variable sources and materiality rubric Agent interprets; human approves strategy
Context readiness 2D Source registry exists Freshness owner missing for 2 sources
Tool feasibility 3D Read-only browser + project folder Minimum path testable
Authority and recovery 3D Draft-only; task-state update No external communication
Evaluation readiness 2D 24 cases Needs conflicting-source and injection cases
Human ownership 3D Research lead and backup Response contract defined
Value and economics 2O Work-time baseline Usage and reviewer cost need pilot data
Operating load 2A Review estimate No observed AI correction burden yet
Table 11Scoring the competitor-brief role: 32 of 36

All six gates pass. The decision is Test, not deploy autonomously.

The pilot contract should require source-owner completion; adverse cases; reviewer minutes; the factual-correction rate; escalation quality; cost per accepted brief; and zero publication or external contact.

Worked Example 2: “AI Marketing Manager”

The proposal asks one worker to research markets, write strategy, produce content, publish, update CRM, run outbound email, and report revenue.

Criterion Score Main weakness
Recurring outcome 1 Several outcomes hidden under a title
Accepted unit 1 No single downstream acceptance
Coherent scope 0 Mixed sources, risk, reviewers, and authority
Work recurrence 2 Tasks recur, but no measured queue
Agentic fit 2 Some variable work; some fixed automation
Context readiness 1 Broad “all marketing docs” plan
Tool feasibility 1 CRM, CMS, email, analytics, and browser access
Authority and recovery 0 Publication and outreach boundary absent
Evaluation readiness 1 Output review only
Human ownership 1 “Marketing team”
Value and economics 1 Salary-substitution claim
Operating load 0 No review or coordination estimate
Table 12Scoring the “AI marketing manager”: 11 of 36

The total is 11/36, and it fails the scope, consequence, evaluation, and ownership gates. The decision is Redesign.

Split the proposal into candidate outcomes: a weekly market-evidence brief; an approved content-outline queue; a CRM enrichment draft; campaign-performance reconciliation; and allowlisted email-draft preparation.

Score each separately. Some may fit workflow automation; some may fit an on-demand assistant; one or two may justify a standing role.

Worked Example 3: Support Triage and Refund Resolution

The initial request asks the role to classify tickets, answer customers, change accounts, and issue refunds.

The team redesigns it: classify three low-risk categories; retrieve the current policy; prepare a response draft; request approval for every send; route refunds, security issues, account changes, and exceptions to people; and update only the internal task state.

Criterion group Score Why
Outcome and scope 10/12 Accepted draft/triage unit; queue baseline still developing
Agentic fit 3/3 Variable language and policy-grounded routing
Context, tools, authority 7/9 Policy sources ready; approval interface needs expiry test
Evaluation and ownership 5/6 Reviewers named; adverse set needs more security cases
Economics and load 4/6 Value plausible; correction and escalation cost unknown
Table 13Scoring the redesigned support role: 29 of 36

The narrower role may be useful without autonomous refund or customer-send authority. The scorecard makes that hybrid design visible.

§ 10Copy-and-Use AI Employee Role Scorecard

A. Role identity. Proposed role; version/date; business owner; supervisor; downstream user/reviewer; current process owner; proposed system shape (workflow/assistant/AI employee/hybrid/human); one-sentence outcome (outcome, queue, cadence, sources, escalation).

B. Gate review. For each of the six gates: pass/fail; evidence; owner; required repair. Then the gate decision: Pass, Redesign, or Stop.

C. Twelve-criterion score. For each criterion: score 0-3; evidence code O/D/A/U; the evidence or gap; owner; due date. Then the total out of 36.

D. Decision. The decision (Test/Redesign/Workflow/Assistant/Human/Stop); failed or conditional gates; the three lowest criteria; the worst plausible failure; human-only decisions; the required pilot mode (observe/shadow/draft/approved action); the evidence the pilot must produce; maximum initial authority; the pause condition; the re-score date or trigger; and approvers.

§ 11Use the Scorecard With CellCog

CellCog’s AI Employees product page describes persistent roles with goals, KPIs, an inbox, task board, memory, schedules, wake conditions, permissions, approvals, shifts, and handovers around a general-purpose agent.

The scorecard should decide which of those capabilities the role needs:

Scorecard evidence CellCog capability to verify
Recurring outcome and accepted unit Role, goals, KPIs, dashboard artifacts
Queue and trigger Inbox, task list, schedule, wake condition
Approved sources Context and memory configuration
Minimum tool path Browse/Cowork and connected systems
Authority boundary Permissions and approvals
Continuity Task state, shifts, and handovers
Evaluation and ownership Observable work, supervisor review, pause behavior
Table 14Mapping scorecard evidence to CellCog capabilities

Do not raise a score because the platform advertises a feature. Ask for proof against the exact role: Can the source scope be limited? Can actions be separated into read, prepare, approved, and blocked states? Does approval bind to the exact action? Can the worker preserve evidence and task state? Can a supervisor pause work and revoke access? Can output, correction, escalation, and cost be measured?

Use CellCog’s role pages as examples, not pre-approved job descriptions. A title such as AI Research Assistant or AI Operations Manager becomes a real candidate only after the organization defines its outcome, queue, inputs, authority, and acceptance.

After a role passes the gates and reaches a test-ready score, use the AI employee pilot framework to precommit the sample, modes, metrics, budget, failure rules, and go/redesign/stop decision. Then use graduated onboarding to increase context and authority only when the evidence gate passes.

§ 12Avoid Common Role-Scorecard Mistakes

Scoring the title. “AI executive assistant” may describe ten different outcomes. Score the actual queue, deliverable, sources, tools, and authority.

Treating unknown as average. An unanswered permission or data question is a 0, not a 2. Unknowns become surprises after connection.

Letting a high total cancel a failed gate. Value and recurrence cannot compensate for prohibited data, missing ownership, or uncontrolled consequence.

Awarding points for planned artifacts. “We will write a rubric” is a gap. Create it, review it, and re-score.

Starting from vendor capability. “The platform can browse and email” does not prove that your role has approved sources or sending authority.

Using precise bands as scientific truth. The thresholds organize an internal decision. Calibrate them with portfolio evidence over time, and keep non-compensating gates.

Ignoring the current alternative. Compare the role with workflow automation, an on-demand assistant, process repair, another specialist product, or human-led work. “Do nothing” is not the only alternative.

Hiding reviewer disagreement. Disagreement about consequence, source authority, or correctness is evidence. Resolve it or narrow the scope.

Equating a high score with production readiness. The scorecard approves the next learning step. Evaluation, pilot, security review, onboarding, monitoring, and change control still remain.

§ 13Final Recommendation

Use the scorecard to force an explicit pre-hire decision:

  1. Pass the six gates.
  2. Score the current role across 12 criteria.
  3. Attach evidence and owners.
  4. Identify the lowest dimensions.
  5. Compare alternative system shapes.
  6. Define the maximum safe pilot mode.
  7. Precommit the evidence needed to proceed.

The best first role is not the one with the grandest title. It is the one whose outcome matters, boundaries hold, evidence can be tested, human ownership is real, and total operating work can plausibly fall.

Frequently asked6 questions

Q1What is an AI employee role scorecard?

It is a pre-configuration decision tool that tests whether a proposed standing AI role has a recurring outcome, observable acceptance, coherent scope, approved context, feasible tools, bounded authority, evaluable behavior, accountable owners, and a credible economic case.

Q2What is a good score for an AI employee role?

As a planning convention, 30-36 can indicate a candidate for bounded testing, 23-29 calls for redesign, and 0-22 suggests another role or system shape. Any failed gate overrides the total. These bands are not universal or safety-validated thresholds.

Q3Should a role with a high score go live?

No. A high score supports a controlled pilot decision. The role still needs representative evaluation, approved access, evidence-gated onboarding, monitoring, incident response, and change control before any bounded production authority.

Q4Who should complete the scorecard?

Include the business owner, current process owner, downstream reviewer, implementation lead, and the owners of relevant systems, security, privacy, legal, compliance, finance, or domain policy. One person may cover several functions in a small team, but every decision needs an accountable owner.

Q5How often should the role be re-scored?

Re-score after material changes to the outcome, task mix, source set, retained context, model, tools, permissions, approval rules, output destination, consequence, pricing, or reviewer capacity. Also re-score after the pilot replaces assumptions with observed evidence.

Q6Can the scorecard be used for any AI agent?

Most criteria apply to agents generally. The recurrence, standing ownership, continuity, task-state, schedule, and handover criteria are especially important when the agent is being placed into an ongoing AI employee role rather than used for a single request.

Published 31 July 2026 All Hiring & onboarding →