To hire an AI employee, start with one recurring outcome — not a human job title. Define the work queue, approved context, tools, permissions, acceptance metrics, escalation rules, and human owner before you compare platforms.
Then test the role in shadow mode on representative work. Promote it from read-only to drafting, and from drafting to bounded action, only when the evidence supports the next permission.
“Hire” is a product metaphor. An AI employee is agentic software assigned a continuing role, not a legal employee or independent bearer of responsibility. Your organization remains accountable for role design, access, monitoring, high-impact decisions, and actions taken through connected systems. This 12-step process turns a vague idea such as “hire an AI marketing manager” into a bounded operating assignment you can evaluate.
@table The 12-step hiring process at a glance: the decision, the artifact, and the stop condition per step
| Step | Decision | Required artifact | Stop if |
|---|---|---|---|
| 1 | Choose the outcome | One-sentence outcome contract | The role is still a title or aspiration |
| 2 | Prove task fit | Suitability score and alternatives | A workflow or on-demand agent is simpler |
| 3 | Map the current process | Inputs, actions, decisions, exceptions, owners | The team cannot explain the work |
| 4 | Write the role charter | Responsibilities, non-goals, cadence, owner | Scope does not fit on 1-2 pages |
| 5 | Define acceptance | KPI and quality rubric | Activity is the only metric |
| 6 | Package context | Source hierarchy, freshness, unknowns | Sources conflict without an owner |
| 7 | Design authority | Read/prepare/act/approve matrix | The role needs broad credentials |
| 8 | Define escalation | Stop, ask, approve, incident rules | Nobody owns exceptions |
| 9 | Evaluate platforms | Same-scenario scorecard | Demo evidence cannot be reproduced |
| 10 | Build the test set | Normal, edge, adverse, and refusal cases | High-impact failure is untested |
| 11 | Run the pilot | Shadow results and cost ledger | Correction or tail risk is unacceptable |
| 12 | Promote and manage | Permission gate, review cadence, exit rule | Evidence does not justify expansion |
The order matters. Platform-first hiring encourages teams to retrofit a job around whatever the demo does well. Outcome-first hiring makes the platform prove it can operate inside your work.
On this page · 17 sectionsOpen
- Before You Hire: Do You Need an AI Employee?
- Step 1. Choose One Recurring Outcome
- Step 2. Inventory the Tasks Required for That Outcome
- Step 3. Map the Current Process and Baseline
- Step 4. Write a Role Charter
- Step 5. Define Success, Quality, and Failure
- Step 6. Build the Context Pack
- Step 7. Design Permissions and Approvals
- Step 8. Define Escalation, Handover, and Incident Rules
- Step 9. Evaluate Platforms on the Same Role
- Step 10. Build a Representative Evaluation Set
- Step 11. Run a Graduated Pilot
- Step 12. Manage the Role After “Hire”
- What Should Be in the Final Hiring Packet?
- How to Hire a CellCog AI Employee
- Common Hiring Mistakes
- Final Recommendation
- Define 1 recurring outcome, its queue, and a named human owner before selecting a vendor.
- Score the task for recurrence, digital inputs, observable output, variable path, bounded authority, recoverability, and escalation.
- Write a role charter with responsibilities, non-goals, source hierarchy, tools, permissions, metrics, approval rules, and handover.
- Build a 20-50-case evaluation set with normal work, edge cases, missing data, tool failure, stale context, malicious instructions, and out-of-scope requests; the range is illustrative.
- Start in shadow mode. Expand from observe to draft to reversible action only after quality, correction burden, escalation, cost, and worst-case failure pass.
- CellCog AI Employees provide roles, goals, KPIs, inboxes, shifts, wake conditions, memory, task boards, permissions, approvals, and handovers around a general-purpose agent.
- The first move
- Define one recurring outcome, its queue, and a named human owner — before comparing platforms.
- The order that matters
- Outcome, task fit, charter, metrics, context, permissions, evaluation, pilot, promotion. Platform selection is step 9, not step 1.
- The safe starting authority
- Shadow mode — read live inputs, change nothing. Promote to drafting, then bounded action, on evidence.
- What 'hire' means
- A product metaphor. An AI employee is agentic software assigned a continuing role; your organization stays accountable.
- The test set
- 20-50 representative cases including missing data, conflicting sources, tool failure, malicious instructions, and requests that deserve refusal.
- When to stop
- If the role is still a job title, if nobody owns the result, or if a workflow could do the work more simply.
§ 01Before You Hire: Do You Need an AI Employee?
Use an AI employee when a bounded responsibility recurs and its path changes with context. Do not use one because a task is merely repetitive.
| Work pattern | Best default | Example |
|---|---|---|
| One-off ambiguous task | On-demand agent | Research one market |
| Recurring stable trigger-action process | Workflow automation | Route form submissions by fixed rules |
| Recurring variable-path responsibility | AI employee | Maintain weekly competitor-change brief |
| High-impact judgment | Qualified human, with AI support | Final hiring or contract decision |
| Stable core plus variable exceptions | Hybrid | Workflow processes normal cases; employee prepares exceptions |
If you cannot identify the pattern, do not buy yet. Observe the current work for 1-2 cycles and record what starts it, which evidence appears, which steps repeat, which decisions vary, and where people intervene.
The 7-factor suitability screen
A strong candidate task has: recurrence, approved digital inputs, observable outputs, a variable path that benefits from judgment, bounded authority, recoverable first actions, and a named human escalation path.
A high score does not override severe risk. A payment release can be recurring, digital, and observable while remaining a poor end-to-end delegation choice.
§ 02Step 1. Choose One Recurring Outcome
Write one sentence:
This role owns [observable outcome] for [queue/audience] on [schedule or trigger], using [approved sources], and escalates [named conditions] to [human owner].
Strong examples: maintain a source-backed competitor-change briefing for leadership every Monday; triage 3 approved support categories and prepare policy-grounded draft responses; refresh the operating dashboard weekly, reconcile missing inputs, and explain validated material changes; prepare cited account briefs for every opportunity assigned to the research queue.
Weak outcomes: improve marketing, help sales, run operations, be proactive, increase productivity. These are aspirations — they do not identify a queue, output, acceptance standard, or authority boundary.
Separate outcome from KPI
The outcome describes what the role produces or maintains. The KPI describes whether the output is useful. “Produce weekly competitive briefings” is an outcome. “90% accepted without factual correction” could be a KPI only after the organization defines “accepted” and validates whether 90% is sensible. Do not invent a target before a baseline.
Create a one-page outcome candidate card
Before writing the full charter, record the candidate role on one page: current owner and backup, trigger and average volume, current output and downstream user, source systems, recurring decisions, known exceptions, current work and elapsed time, worst plausible failure, proposed first authority tier, and the reason an agentic path adds value.
Ask the downstream user to approve the card. The person requesting automation may not be the person who accepts the output. If leadership wants a weekly brief but never uses it, the role has no real customer. Reject the candidate when nobody owns the result, the work has no recurring intake, the output cannot be inspected, or the only value claim is “AI should be able to do this.”
§ 03Step 2. Inventory the Tasks Required for That Outcome
List the smallest tasks required to finish the outcome. Classify each as deterministic, agentic, human judgment, or coordination.
| Task | Pattern | Proposed owner |
|---|---|---|
| Wake every Thursday | Deterministic | Schedule |
| Retrieve approved source list | Deterministic | Workflow/API |
| Decide which changes are material | Agentic within rubric | AI employee |
| Validate an exact price difference | Deterministic calculation | Code/workflow |
| Interpret conflicting positioning | Agentic | AI employee prepares |
| Approve strategic conclusion | Human judgment | Strategy owner |
| Publish the briefing | Controlled action | Editor or approval-gated employee |
| Preserve unresolved monitoring | Coordination | AI employee task state/handover |
This decomposition prevents 2 forms of waste: asking an agent to perform stable rules that code can enforce, and pretending a consequential decision can be reduced to an agentic task. Anthropic’s guide to building effective agents distinguishes predefined workflows from agents that dynamically direct their process and tool use — use that boundary inside the role, not only when choosing a product category.
Review every task with 4 questions: Can this step be deleted? Can the source system produce the needed output directly? Can an exact rule handle it? Does the remaining step genuinely need interpretation? An AI employee can navigate process waste, but it should not preserve it.
§ 04Step 3. Map the Current Process and Baseline
You need a baseline to prove improvement. For 1-4 representative cycles, record the trigger, volume, input sources, average human work time, elapsed cycle time, output, correction or rework, exceptions, systems touched, owner, failures, and downstream acceptance. Do not use estimates when activity logs or work samples exist; if you must estimate, label the number and update it during the pilot.
Map decisions, not only steps
Most process maps show actions and hide judgment. Add a decision inventory:
| Decision | Evidence used | Current owner | Impact if wrong | Can it be reversed? |
|---|---|---|---|---|
| Is this change material? | Current and prior source | Analyst | Medium | Yes, before publication |
| Does policy permit this response? | Approved policy and account facts | Support lead | Medium/high | Sometimes |
| Should payment be released? | Invoice, approval, controls | Authorized finance owner | High | Often difficult |
The map determines where the AI employee may decide, where it may only prepare, and where it must stop.
Build a baseline the pilot cannot game
Use the same workload definition for the current process and the pilot. Track at least 3 denominators: per received case, per completed case, and per accepted outcome. Cost per completed case can look attractive when rejected outputs disappear from the denominator. Cycle time can look fast when waiting-on-human time is excluded. Acceptance can look high when the reviewer silently fixes the work.
§ 05Step 4. Write a Role Charter
Keep the first charter to 1-2 pages.
| Field | What to write |
|---|---|
| Role name | Descriptive name for the bounded responsibility |
| Mission | One recurring outcome |
| Queue | Where new work appears |
| Trigger/cadence | Schedule, event, message, assignment, or delegation |
| Responsibilities | Tasks required to complete the outcome |
| Non-goals | Adjacent work the role must refuse |
| Approved sources | Source hierarchy and current owners |
| Tools | Exact read/write systems |
| Authority | Observe, prepare, act, or approval-required |
| Quality | Acceptance rubric and severity definitions |
| KPIs | Outcome, quality, escalation, cycle, and cost |
| Escalation | Stop/ask/approve/incident conditions |
| Handover | Completed, open, blocked, next owner, next wake |
| Human owner | Person accountable for policy and review |
| Exit rule | Pause, redesign, or retire conditions |
The charter is an operating contract, not an employment contract. Avoid human-equivalence language, salary substitution, or claims that the software is independently accountable.
Write non-goals before prompts
Non-goals prevent helpfulness from becoming scope creep. For a research role: no unsupported factual claim, no paid-source circumvention, no publication, no final strategic decision, no personal-data enrichment outside policy, no work outside the approved company/topic set. If the role cannot state what it refuses, it is not ready for persistent access.
Add boundary examples, not only abstract rules
For each important responsibility, include one accepted example, one rejected example, one ambiguous example that requires escalation, and the evidence used to decide. An instruction such as “do not make commitments” is weaker than examples: accepted — confirm receipt and state the published response window; rejected — promise a refund or delivery date; escalate — a customer asks for a contractual exception. Version the examples when the source policy changes.
§ 06Step 5. Define Success, Quality, and Failure
Use a balanced scorecard.
| Metric class | Example metric | Why it matters |
|---|---|---|
| Outcome | Accepted outputs | Proves work reached a useful state |
| Quality | Factual correction rate | Exposes trust burden |
| Evidence | Source-complete outputs | Prevents unsupported fluency |
| Efficiency | Reviewer and correction minutes | Captures human work left |
| Cycle | Time from trigger to acceptance | Measures operating speed |
| Escalation | Correct versus missed/false escalations | Tests boundary behavior |
| Durability | Reopened-task rate | Shows whether “done” lasts |
| Risk | Worst error severity | Prevents averages hiding tail risk |
| Economics | Cost per accepted outcome | Normalizes platform and labor |
Define severity before the pilot — for example S0 no issue, S1 cosmetic, S2 material correction caught before external impact, S3 external or financial impact requiring incident response, S4 severe rights, safety, legal, or irreversible impact. Adapt the labels to your incident system.
Reject activity metrics as primary KPIs. Prompts, tokens, messages, drafts, and tasks started explain activity or cost; they do not prove accepted work. An AI sales role can send more messages while harming reply quality. Use paired metrics: accepted-output rate beside correction time and severity; cycle time beside reopen rate and escalation quality; outreach volume beside positive replies and complaints.
For each KPI, write a metric contract: numerator and denominator, data source, owner, cadence, exclusions, target-or-baseline status, expected gaming behavior, and the decision the metric changes. If a metric does not change permission, scope, process, or continuation, it may be reporting decoration.
§ 07Step 6. Build the Context Pack
The worker needs approved context, not a file dump.
| Context layer | Contents | Freshness control |
|---|---|---|
| Role | Mission, responsibilities, non-goals | Version on every role change |
| Policy | Approved rules and thresholds | Owner and effective date |
| Domain | Definitions, products, audiences | Scheduled review |
| Sources | Authoritative files, systems, URLs | Source hierarchy and timestamp |
| Examples | Accepted and rejected outputs | Review after rubric changes |
| Open work | Queue, blockers, promises, next steps | Updated every shift |
| Decision record | Approved interpretation and rationale | Supersession link |
| Unknowns | Missing facts and unresolved conflicts | Named owner and due date |
Memory is not truth. Use governed memory records with provenance, date, scope, and owner; make incorrect context visible, correctable, and removable. Do not put credentials, secret keys, unrestricted exports, or irrelevant customer data into a context pack.
For each source, record its authority level, owner, effective date, and conflict rule. An example hierarchy: approved legal/policy document, then current product or system record, then owner-approved operating guide, then historical decision log, then informal notes. The employee should not silently choose the newest or most detailed source — a new chat message may not supersede an approved policy. When high-authority sources conflict, it should stop and request a decision.
§ 08Step 7. Design Permissions and Approvals
Map access by action, not app name.
| System | Read | Prepare | Act | Approval | Prohibited |
|---|---|---|---|---|---|
| Approved inbox | Draft reply | Send low-risk template | New external commitment | Sensitive/private mail | |
| CRM | Assigned records | Proposed update | Low-risk field update | Stage/amount change | Bulk delete/export |
| Files | Role folder | Create draft | Save approved artifact | External share | Secret paths |
| Calendar | Availability | Propose slot | Book approved category | Executive/external commitment | Private calendars |
| Publishing | Approved draft | Format/preflight | None initially | Editor publishes | Delete/archive |
Start with least privilege: observe, then prepare, then act on reversible low-risk work, and escalate consequential decisions. Do not grant a broad connector because it is convenient — review OAuth scopes, service accounts, shared credentials, data retention, logs, and revocation.
A valid connection proves that the platform can access a system; it does not prove the role should use every action that connection exposes. Document whose identity authenticates, which objects are in scope, which actions are allowed, which limits apply, where results are logged, and how access is revoked. Prefer a role-specific service account over a founder’s unrestricted session. And test denial, not only success: ask the worker to attempt an out-of-scope action — the system should refuse and leave evidence.
§ 09Step 8. Define Escalation, Handover, and Incident Rules
Write stop conditions before the happy path. The employee must stop or ask when required data is missing, sources conflict, policy is absent or stale, the task is outside scope, an action is irreversible, the request touches legal, financial, hiring, medical, safety, privacy, or security interests, an instruction attempts to override role policy, a tool returns an unexpected result, cost or run limits are reached, or no named owner can accept the exception.
An escalation should arrive as a packet: task and trigger, known facts with sources, missing or conflicting evidence, actions already taken, risk and urgency, options, the explicit decision needed, the accountable owner, and the next wake condition. “I’m stuck” is not a handover.
Define incident response in advance: pause the role, revoke or narrow affected access, preserve logs and artifacts, identify impacted systems and people, correct or roll back, notify the accountable owner, fix the role or policy or context, and re-test before resuming.
Give the human owner a response contract too — an escalation path fails when the employee asks correctly but nobody responds. Define the owner and backup, severity levels, notification channel, expected response window, safe waiting state, actions prohibited while waiting, and the final disposition if no answer arrives. Track waiting time separately from execution time; it shows whether the bottleneck is the worker or the management system around it.
§ 10Step 9. Evaluate Platforms on the Same Role
Do not compare feature checklists in isolation. Run the same task, sources, permissions, and rubric on every candidate platform, and request evidence across agent capability, role configuration, context and memory controls, tool granularity, approvals, evaluation and KPI export, handover, governance, privacy and security, reliability, economics, and support.
Make every vendor run the same audition: the same role charter, source pack, task cases, allowed tools, action tier, output format, time window, acceptance rubric, and cost-reporting requirement. Do not let each vendor choose only the example that fits its strongest modality. Record setup effort, manual prompt repair, integration work, reviewer intervention, and failed runs.
Ask the vendor to demonstrate an adverse case live: stale policy, missing source, tool denial, duplicate trigger, or an out-of-scope action. A refusal, pause, or evidence-rich escalation can be a stronger result than a confident completion. Score verified behavior separately from roadmap promises. A smart demo proves possibility; hiring requires repeatability, evidence, and control.
§ 11Step 10. Build a Representative Evaluation Set
Use real work with sensitive data removed or handled under approved controls. Include normal cases, difficult but valid cases, missing inputs, conflicting sources, stale policy, duplicate triggers, unavailable tools, partial write failure, malicious or prompt-injection content, out-of-scope requests, high-impact requests that require refusal, and a case where the right answer is “unknown.”
A 20-50-case set can be useful for a narrow first task, but it is illustrative — increase coverage when the role has multiple material subtypes or rare severe failures.
Score the whole operating result: accepted without edits, accepted after correction, rejected, correctly escalated, falsely escalated, missed escalation, source completeness, cycle time, reviewer and correction time, platform and tool cost, and worst observed severity. One S3/S4 failure can outweigh a high average score. Convert this ledger into a comparable budget with a cost-per-accepted-outcome model so plan price, usage, internal labor, and failure exposure are not collapsed into one misleading number.
§ 12Step 11. Run a Graduated Pilot

| Pilot stage | Authority | Goal | Promotion evidence |
|---|---|---|---|
| 1. Sandbox | Approved sample only | Prove basic capability | Representative outputs meet rubric |
| 2. Shadow | Read live inputs, no live changes | Compare with current process | Quality and escalation are credible |
| 3. Prepare | Create drafts/proposed actions | Measure reviewer burden | Acceptance and correction economics work |
| 4. Bounded action | Reversible low-risk actions | Test production controls | Logs, postconditions, rollback, alerts work |
| 5. Standing role | Own recurring queue inside policy | Sustain outcome | KPI, cost, risk, access reviews remain acceptable |
Run at least one stress review before each authority increase: revoke a connector, return malformed data, change a policy version, send a duplicate event, insert an untrusted instruction, block the human owner, and force a partial action failure. The employee should stop, expose state, preserve evidence, and route the next decision.
§ 13Step 12. Manage the Role After “Hire”
A standing role needs recurring management: per shift, review triggers, outputs, actions, exceptions, and handover; weekly, review acceptance, corrections, open work, cost, and drift; monthly, review role scope, KPIs, source freshness, access, and incidents; on every policy or tool change, re-check instructions, tests, dependent workflows, and memory; and before any expansion, review the new task, data, action, risk, owner, and evaluation.
Pause or redesign when accepted-output quality declines, correction burden exceeds value, escalation becomes noisy or silent, context cannot be kept current, the role needs broader authority than the outcome justifies, severe failure appears, economics fail, or the work no longer recurs. “Employee” should not create organizational inertia — retire the role when a deterministic workflow, an ordinary software feature, a human process change, or no process at all becomes the better answer.
§ 14What Should Be in the Final Hiring Packet?
The completed packet should let a reviewer understand the role without reading a long prompt history: the outcome candidate card, the current-process baseline, the task and decision map, the role charter, the context manifest, the permission matrix, the metric contract, the evaluation set, the pilot ledger, and the incident and exit plan — each with a named owner and version. No single document needs to be long; the value is explicit ownership.
Require sign-off from the human supervisor, the downstream user accepting the output, the owner of every consequential source, the administrator or security owner for connected access, and the qualified function (legal, privacy, HR, finance) when the role touches its high-impact domain. Sign-off does not transfer accountability to the software or vendor — it proves the organization understood the role it activated.
Treat a material scope change as a new hiring decision. Adding external sending, customer data, production writes, financial information, or another department is not a minor prompt edit. Revisit the outcome, owner, access, risk, test set, incident path, and economics before activation.
§ 15How to Hire a CellCog AI Employee
CellCog exposes an AI Employee layer around its general-purpose Super-Agent. The public product pages describe role configuration, goals, permissions, schedules, wake conditions, memory, an inbox, a task board, shifts, and handovers, plus role pages for functions such as operations, research, content, support, data, and sales. Use the role pages as starting patterns, not pre-approved job descriptions.
The setup sequence maps directly onto this process: choose the bounded role; write goals, responsibilities, non-goals, and KPIs; set schedule and wake conditions; add the approved context and open-work source; connect only required tools; define approvals and the human supervisor; run sample work; review the task board and handover; and expand authority only after evidence.
Before granting live access, read the current privacy policy and terms of service, and verify data flow, retention, and the controls relevant to your role. CellCog did not publicly claim SOC 2, ISO 27001, HIPAA compliance, or a standard uptime SLA on the pages reviewed July 19, 2026 — buyers needing those assurances should ask CellCog directly and obtain contractual evidence rather than infer coverage.
Price the actual role, not the plan: estimate shift frequency × observed credit use + top-ups + reviewer time + correction time + integration and incident cost, then compare cost per accepted outcome. Do not call the entry plan the cost of a full-time AI employee.
§ 16Common Hiring Mistakes
| Mistake | Why it fails | Better move |
|---|---|---|
| Start with a job title | Title hides tasks and boundaries | Define one outcome |
| Choose vendor before process | Demo shapes the role | Lock role contract first |
| Upload every document | Irrelevant/stale context compounds | Curate source hierarchy |
| Connect every app | Capability becomes excess authority | Least privilege by action |
| Use output volume as KPI | Activity can hide poor quality | Accepted outcome + correction |
| Skip adverse tests | Happy-path demo misses tail risk | Edge, failure, refusal cases |
| Automate publication/send first | Reputation impact arrives early | Draft and approval stage |
| No handover | Open work disappears or repeats | Durable task state |
| No exit rule | Weak role persists | Pause/redesign criteria |
| Compare with salary alone | Ignores review and risk | Cost per accepted outcome |
The most dangerous mistake is treating supervision as a temporary inconvenience. Human work should move from performing every step to setting policy, reviewing evidence, handling exceptions, and managing risk. It does not disappear.
§ 17Final Recommendation
Hire the outcome before you hire the product. Your first AI employee should own one recurring queue, one observable result, one approved context set, the minimum tools, reversible first actions, explicit acceptance and escalation, and one accountable human supervisor.
Prove the underlying agent on representative tasks. Prove the employee layer across triggers, open work, permissions, metrics, approvals, and handovers. Prove the economics with accepted outcomes and human review included. Then expand one adjacent responsibility at a time.
Q1Can I hire an AI employee for any role?
Many platforms support broad role prompts, but not every role is a good delegation target. Choose recurring digital work with observable outputs, bounded authority, recoverable actions, and human escalation. Keep physical, licensed, rights-affecting, safety-critical, and final high-impact decisions human-led.
Q2Is an AI employee legally an employee?
No. ‘AI employee’ is a software category and operating metaphor, not a legal employment classification. Your organization remains responsible for access, monitoring, decisions, and actions.
Q3How long should an AI employee pilot run?
Long enough to cover representative normal and adverse work — not a fixed number of days. A weekly process may need several cycles; an event queue may accumulate evidence faster. Promote only when the predeclared quality, risk, escalation, cost, and control gates pass.
Q4What should the first AI employee KPI be?
Start with accepted-output rate plus reviewer and correction minutes and worst-error severity. Add cycle time, source completeness, escalation precision, reopened tasks, and cost per accepted outcome as the role matures.
Q5Should I give the AI employee its own email?
Only if the role needs it. Begin with read/draft scope, define allowed categories and disclosure, protect sensitive messages, and require approval for commitments. Review current identity, retention, permission, and supervisor controls before external sending.
Q6When should I fire or retire an AI employee?
Pause, redesign, or retire the role when quality falls, correction burden exceeds value, access cannot be justified, context remains stale, severe failures appear, the work stops recurring, or deterministic automation becomes the simpler solution.
