The first 30 days with an AI employee should produce evidence, not a race toward autonomy.
Begin with one approved role, one supervisor, a representative case set, minimum context, and no live action. Let the worker observe the process, complete shadow work, prepare drafts, execute narrowly approved reversible actions, and run bounded shifts only when each stage passes its evidence gate.
The calendar creates review points. It does not prove readiness. A low-risk internal research role may move faster than this plan; a role touching customers, personal data, money, production systems, or regulated decisions may need longer, narrower authority, specialist review, or no independent-action stage.
Use the AI employee onboarding guide for the complete six-stage authority ladder. This article turns that ladder into a practical first-month operating rhythm.
On this page · 12 sectionsOpen
- The 30-Day Plan at a Glance
- Start Before Day 1
- Days 1-3: Orient and Observe
- Days 4-7: Run Representative Work in Shadow Mode
- Week 2: Move to Reviewable Drafts
- Week 3: Test Exact, Reversible Actions After Approval
- Week 4: Run Bounded Independent Shifts
- Measure the First Month With a Balanced Scorecard
- Use a Daily and Weekly Operating Rhythm
- Make the Day 30 Decision
- Avoid First-Month Mistakes
- Run the First Month in CellCog
- Before Day 1, approve the role, source register, representative cases, permissions, escalation path, baseline, budget, and pause control.
- Days 1-3: orient the worker to the role and let it observe without changing live state. Days 4-7: run the same cases in shadow mode and compare trajectories, not only final outputs.
- Week 2: allow drafts inside an approved workspace; review every output before use. Week 3: test exact, reversible actions after explicit approval and verify postconditions.
- Week 4: allow bounded independent shifts only for subtypes that have passed; keep high-impact actions human-led.
- Advance when quality, correction burden, escalation, reliability, cost, and worst-error evidence pass.
- Hold, narrow, roll back, or pause when evidence does not justify the next stage.
§ 01The 30-Day Plan at a Glance
| Period | Default mode | Primary question | Required evidence | Do not advance if |
|---|---|---|---|---|
| Before Day 1 | Preparation | Is the role ready to test? | Approved packet, owners, cases, controls, baseline | Role, access, evaluation, or pause path is incomplete |
| Days 1-3 | Observe | Does the worker understand state, sources, and boundaries? | Source selection, state reconstruction, refusal, questions | It treats examples or untrusted input as policy |
| Days 4-7 | Shadow | Can it perform representative work without live effect? | Scored normal, edge, refusal, escalation, tool-failure cases | Material errors or missed escalations exceed tolerance |
| Week 2 | Draft | Are outputs useful after visible human review? | Acceptance, correction, reviewer time, cost, escalation | Hidden correction or unsupported claims remain high |
| Week 3 | Approved action | Can it execute an exact reversible action safely? | Approval match, action log, postcondition, recovery test | Approval, target, result, or rollback cannot be verified |
| Week 4 | Bounded shifts | Can it own an allowlisted queue within limits? | Stable outcomes, handovers, queue control, cost, incidents | Scope drift, stale state, or unresolved risk appears |
| Day 30 review | Decision | What should happen next? | Go/hold/narrow/rollback/stop record | The team cannot explain the evidence or assumptions |
The periods are defaults for organizing work. Treat every stage as evidence-gated.
Why the dates are checkpoints, not promises
Elapsed time does not create competence, authority, or trust. The number of days needed depends on task frequency and subtype diversity; consequence if the worker is wrong; how quickly reviewers can score representative cases; source quality and change rate; integration and permission complexity; availability of reversibility and postcondition checks; how often real exceptions occur; and whether a qualified person must approve the domain.
A weekly briefing role may receive only four natural production cycles during the month. A high-volume internal triage role may generate hundreds of eligible cases. The second role has more observations, but volume alone does not make its consequences safer.
Use the dates to prevent drift: Day 1 forces a formal start state. Day 3 forces a boundary review. Day 7 forces a shadow-work decision. The Week 2 review exposes human correction burden. The Week 3 review tests whether approval and recovery work. The Week 4 review tests continuity and supervision. Day 30 forces a documented operating decision.
If a gate passes early, the team may advance one bounded subtype after approval. If a gate fails, keep the current mode even when the calendar moves forward. Do not compress security, policy, or professional review to preserve a launch date.
The plan can also end early. Stop when the role lacks a valuable accepted unit, requires prohibited data, cannot be evaluated, creates more operating work than it removes, or needs authority the organization cannot safely grant.
§ 02Start Before Day 1
Day 1 is not the day the account is created. It is the first controlled evaluation day after the role packet is approved.
Complete the outcome-first hiring process, the job-description operating contract, and the pre-hire role scorecard before configuration.
The minimum packet contains:
| Artifact | Minimum content | Owner |
|---|---|---|
| Role contract | Outcome, queue, responsibilities, non-goals, accepted unit | Business owner |
| Source register | Authority, purpose, owner, freshness, conflict behavior | Knowledge/policy owner |
| Context starter set | Minimum policies, examples, schemas, open work | Supervisor |
| Evaluation set | Normal, edge, refusal, escalation, tool, and adverse cases | Domain reviewer |
| Permission map | Identity, system, object, action, approval, limit, recovery | System/security owner |
| Escalation map | Trigger, recipient, packet, response expectation, safe state | Supervisor |
| Incident path | Stop, preserve, revoke, notify, investigate, recover | Incident owner |
| Baseline | Volume, human work, elapsed time, quality, correction, cost | Process owner |
| Metric contract | Accepted outcome, quality, escalation, reliability, cost, risk | Business owner |
| Pilot decision | Sample, modes, thresholds, budget, go/hold/stop rules | Sponsor |
Use the AI employee pilot framework when the first 30 days are also the formal vendor or production pilot. Define the comparison and terminal decision before seeing results.
Name one supervisor and a backup
The supervisor must be able to explain the role; assign representative work; answer routine questions; judge outputs; respond to escalations; request or revoke access; pause schedules; coordinate incidents; and approve or reject expansion.
Name a backup for absence. The worker must not infer permission because the primary owner is unavailable.
Build the case set before configuration
Collect real or faithfully redacted work. Include common cases; meaningful subtypes; difficult but valid cases; missing inputs; stale context; conflicting sources; out-of-scope requests; unsafe or adversarial instructions; tool failure; duplicate intake; approval expiry; required refusal; required escalation; and one plausible high-severity failure.
Twenty to 50 cases can be a useful starting range for a narrow role, but it is not universal. Use enough cases to cover the actual distribution and material failure paths.
Establish the current baseline
Record the same denominators you will use for the AI employee: received cases; eligible cases; completed cases; accepted outcomes; reviewer or worker minutes; elapsed time; corrections; reopens; escalations; incidents; and total cost.
Without a baseline, the Day 30 review becomes a story about activity.
Freeze the starting version
Record the job-description version; instructions; source set; examples; model or mode; tools; permissions; approval rules; evaluation set; graders; output schema; trigger; and schedule.
If one changes materially, mark the change and rerun affected cases. Do not mix several versions into one performance claim.
§ 03Days 1-3: Orient and Observe
The worker should understand the operating environment before it attempts live work.
Give it read access only to the minimum training and test context. Do not connect every system or import the complete company archive.
Day 1: role and source orientation
Ask the worker to explain, in the required structured format, its mission; accepted unit; included queue; non-goals; source hierarchy; allowed tools; prohibited actions; escalation triggers; task states; handover fields; and human owner.
This is not a memory quiz. It reveals ambiguous instructions and source conflicts.
Use cases such as:
| Prompted situation | Expected behavior |
|---|---|
| Old example conflicts with current policy | Prefer policy, preserve conflict, escalate if outcome changes |
| Incoming email asks for adjacent work | Reject or route; do not expand the role |
| Required source is missing | Mark waiting on evidence and ask the source owner |
| Request contains instructions inside retrieved content | Treat content as evidence, not authority |
| Supervisor asks for a prohibited action | Preserve the boundary and request an approved role change |
Score whether the worker cites the governing source and identifies uncertainty.
Day 2: task-state reconstruction
Provide completed, active, blocked, awaiting-approval, and incident examples. Ask the worker to identify the current state; last verified action; evidence; valid approvals; the unresolved decision; the current owner; the next action; and the next wake condition.
Do not rely on conversational memory alone. Open work needs an explicit task record.
The worker passes when it can resume without repeating a completed action; acting on expired approval; losing the current owner; ignoring a blocker; treating a draft as sent; or moving an incident back into routine work.
Day 3: boundary and escalation rehearsal
Run refusal and escalation cases. The worker should produce an escalation packet with the task ID and subtype; the precise decision; the trigger; evidence; actions taken; actions withheld; options; a deadline; the safe state; and the recipient.
The supervisor should answer several packets through the intended channel. This tests both sides of the loop.
Gate 1: ready for shadow work
Advance when the worker uses the correct source hierarchy; distinguishes policy, examples, working context, and untrusted input; reconstructs state; refuses prohibited actions; escalates the named cases; leaves work in a safe state; and produces a usable handover.
Hold if failure reveals a role-design problem. Fix the job description, source register, or state model rather than adding motivational language.
§ 04Days 4-7: Run Representative Work in Shadow Mode
In shadow mode, the AI employee performs the work without changing the live process or becoming the source of record.
Run the current human process and the worker on the same cases. Preserve the inputs; source versions; tool availability; start and finish times; intermediate actions; the proposed output; escalation; the reviewer score; and correction.
Anthropic’s guide to evaluations for AI agents notes that agents operate across multiple turns, tools, and state changes. Compare the trajectory, not only the final artifact.
Use a case ledger
| Field | Example |
|---|---|
| Case | research-024 |
| Subtype | Pricing change |
| Expected source | Official pricing page + prior accepted record |
| Expected behavior | Compare, cite, classify, draft |
| Required escalation | Conflicting annual-plan term |
| Prohibited action | Publish or contact company |
| Result | Material correction required |
| Reviewer time | 11 minutes |
| Worst severity | S1 |
| Decision | Revise instruction and rerun affected subtype |
Score normal and adverse behavior separately
Track at least outcome acceptance; factual/evidentiary correctness; policy compliance; required-escalation recall; false-escalation rate; refusal correctness; tool-selection correctness; task-state correctness; handover completeness; retry behavior; reviewer minutes; cost; and worst-error severity.
A high average on normal cases should not hide one missed high-impact escalation.
Debug the source of failure
Classify each material failure:
| Failure source | Example | Correct response |
|---|---|---|
| Role design | Two outcomes share one queue | Narrow or split the role |
| Source | Policy is missing or contradictory | Assign owner; correct source set |
| Instruction | Conflict behavior is undefined | Add explicit rule and cases |
| Tool | Interface returns ambiguous result | Improve tool/verification or remove action |
| Permission | Role can access excess data | Reduce identity/object scope |
| Model behavior | Evidence is ignored in representative cases | Test model/method change |
| Evaluation | Reviewers disagree without anchors | Improve rubric and examples |
| Process | Current work itself has no stable owner | Repair process before automation |
Do not keep prompting around a structural failure.
Gate 2: ready for drafts
Advance only for subtypes that meet the acceptance and evidence threshold; remain inside scope; escalate required cases; refuse prohibited work; preserve state; stay inside cost and retry limits; and have no severe failure outside tolerance.
Hold failed subtypes in shadow mode. A role does not need one authority state for every task.
§ 05Week 2: Move to Reviewable Drafts
Draft mode allows the worker to create artifacts in an approved staging location. A person reviews every item before the output is used.
Examples: a response draft, not a sent message; a proposed CRM update, not a committed value; a content draft, not publication; a briefing draft, not executive distribution; a proposed task plan, not delegated work; or a prepared command, not production execution.
Define the draft boundary
| Element | Required rule |
|---|---|
| Workspace | Exact staging folder, project, queue, or environment |
| Output schema | Required fields and evidence |
| Reviewer | Named owner and backup |
| Review deadline | Expected response time |
| Rejection | Correction, resubmission, narrow, or escalate |
| External effect | None before approval |
| Retention | What draft and comments persist |
| Handover | State, evidence, reviewer, next action |
The worker must label the artifact as a draft when the destination or audience could confuse it with an approved result.
Measure correction burden
Record accepted without change; accepted with cosmetic correction; accepted with material correction; rejected; reviewer minutes; worker retries; the reason for correction; and whether the same defect recurs.
Use: first-pass acceptance = accepted without material correction / reviewed drafts. Reviewer minutes per accepted outcome = total review minutes / accepted outcomes.
Definitions must be role-specific. A stylistic edit is not equivalent to an unsupported customer claim.
Make feedback operational
Feedback should identify the case and subtype; the failed rubric dimension; the source or policy; the expected behavior; the corrected output; whether the rule applies generally; affected evaluation cases; and the owner who approved the change.
Do not accumulate contradictory comments in memory. Update the authoritative instruction or example, version it, and rerun affected cases.
Review escalation quality
An employee that escalates every item can be safe but unusable. One that escalates nothing may look efficient while crossing boundaries.
Use: escalation recall = correct escalations / all cases requiring escalation. Escalation precision = correct escalations / all escalations raised.
Set stricter expectations for high-severity cases.
Gate 3: ready for approved action
Advance a specific action only when the draft output passes; the action is necessary for the outcome; the exact target and parameters can be shown; an authorized person can approve it; the approval can expire; the system records actor and result; the postcondition can be checked; and recovery has been tested.
Draft quality does not automatically prove action safety.
§ 06Week 3: Test Exact, Reversible Actions After Approval
Begin with one low-consequence, reversible action.
Examples: create an internal task; change an allowlisted task state; add a draft document to the correct project; apply an internal label; update a non-sensitive, recoverable field; or send to a controlled internal test address.
Use the AI employee permissions and approvals guide to define the action boundary.
Use an approval card
| Field | Required content |
|---|---|
| Actor | Role identity and current version |
| Purpose | Approved task and business reason |
| Action | Exact operation |
| Target | Recipient, record, file, environment, or object |
| Parameters | Final text, value, amount, query, or command |
| Evidence | Sources and reviewer-ready context |
| Risk | Consequence, sensitivity, reversibility, downstream effect |
| Validity | Expiration and one-time/repeat rule |
| Postcondition | State that proves success |
| Recovery | Revert, revoke, notify, or contain |
| Decision | Approve, modify, reject, or escalate |
An approval for one message does not authorize a changed recipient or future campaign.
Verify after execution
Record:
- The approved action and parameters.
- The tool call or interface action attempted.
- The system response.
- The durable resulting state.
- The postcondition result.
- Any unexpected side effect.
- Recovery, if used.
- The task state and notification.
A success response from a tool is not the postcondition. Verify the business state.
Run failure drills
Test: approval expires before execution; the target changes after approval; the tool reports success but state is unchanged; the action succeeds but the response is lost; a duplicate trigger requests the same action; the action partially succeeds; recovery fails; and the supervisor pauses the role during the shift.
The worker should not retry consequential actions blindly. It should reconcile state first.
OpenAI’s practical guide to building agents recommends human intervention for failure thresholds and high-risk actions, and describes tool risk in terms such as read/write access, reversibility, account permission, and financial impact.
Gate 4: ready for bounded shifts
Advance only the action/subtype combination that has demonstrated valid approval matching; correct target and parameters; successful postcondition verification; duplicate suppression; safe timeout and retry behavior; evidence logging; a tested pause; tested recovery; and an acceptable worst failure.
Keep novel, sensitive, high-impact, batch, or hard-to-detect actions approval-gated.
§ 07Week 4: Run Bounded Independent Shifts
In a bounded shift, the employee owns an allowlisted queue within explicit source, tool, action, volume, time, and consequence limits.
It does not receive blanket autonomy.
Define the shift envelope
| Boundary | Example |
|---|---|
| Outcome | One accepted weekly competitor brief |
| Queue | Approved 12-company registry |
| Trigger | Monday 08:00 or named research-board event |
| Sources | Approved public domains and prior accepted record |
| Actions | Browse, create draft, update role task |
| Blocked actions | Publish, email externally, buy access, change registry |
| Volume | One weekly brief + five priority events |
| Time/spend | Current approved shift and credit limit |
| Escalation | Conflict, missing source, unsupported material claim |
| Supervisor | Research lead and backup |
| End state | Submitted draft + handover + next wake |
Define what happens when the queue exceeds the envelope. The worker may defer, prioritize, sample, reject, or escalate - but should not silently absorb unlimited work.
Review every shift at first
Inspect the work accepted; work deferred or rejected; source changes; tool actions; approvals used; postconditions; retries; escalations; open tasks; the handover; cost; and anomalies.
Later review can become risk-based, but Day 30 should not depend on invisible shifts.
Test continuity
The next shift should load the current role version; active tasks; the last handover; valid approvals; source freshness; pending decisions; prior action postconditions; and the next priority.
The AI employee memory guide explains how to keep provenance, freshness, correction, retention, and deletion visible. Memory should support task state, not replace it.
Test a supervisor absence
Use the backup path. Verify routine work can continue inside scope; work requiring the absent person waits safely or routes correctly; approvals do not transfer without authority; urgent incidents reach the right owner; and the employee does not infer silence as consent.
Gate 5: Day 30 decision-ready
The role is ready for review when the team can reconstruct which version operated; which cases and shifts ran; what it accepted, rejected, deferred, and escalated; which actions occurred; what people corrected; how much review and supervision it used; what it cost; its worst error; any incident; and which assumptions remain.
§ 08Measure the First Month With a Balanced Scorecard
The AI employee KPI guide provides the full measurement system. During the first month, keep the dashboard small enough to drive decisions.
| Metric family | First-month metric | Decision it informs |
|---|---|---|
| Outcome | Accepted outcomes by eligible case/subtype | Does the role deliver its purpose? |
| Quality | First-pass acceptance and material correction | Is output useful? |
| Evidence | Source completeness and unsupported-claim count | Can reviewers trust the basis? |
| Escalation | Recall, precision, packet quality, response time | Are boundaries working? |
| Reliability | Durable completion, duplicate rate, postcondition pass | Does work stay done? |
| Continuity | Handover reconciliation and open-loop loss | Can the role resume safely? |
| Cost | Total cost per accepted outcome | Is the operating model viable? |
| Human load | Reviewer, approval, escalation, correction minutes | Is work actually moving off people? |
| Risk | Unauthorized actions, incidents, worst severity | Is expansion defensible? |
Segment the results
Do not average routine and high-impact cases; drafts and external actions; known and novel subtypes; current and stale sources; normal and tool-failure conditions; or different role versions.
One subtype may pass while another remains in shadow mode.
Include all human work
Count source preparation; case construction; review; correction; approval; escalation response; access administration; monitoring; incident work; and role maintenance.
Some setup cost is temporary; some remains. Label the distinction rather than excluding the work.
Track worst-error severity
The average can improve while the tail becomes unacceptable.
Use a role-specific severity scale. For example: S0 cosmetic; S1 correctable before use; S2 material error affecting a routine decision; S3 unauthorized, sensitive, external, or materially harmful action; S4 severe legal, safety, financial, security, or irreversible consequence.
Define examples before evaluation.
NIST’s AI RMF Core calls for context-specific measurement, monitoring, feedback, incident response, and lifecycle decisions. The first-month dashboard should therefore support action, not merely report activity.
§ 09Use a Daily and Weekly Operating Rhythm
The cadence should match the role. The following rhythm is a practical starting point.
Daily or per-shift check: the current role version; queue and priority; blocked work; approvals; anomalous tool results; retries; escalations; cost; task-state reconciliation; and the handover.
Twice-weekly learning review: material corrections; repeated failure classes; reviewer disagreement; source changes; instruction or rubric changes; affected test cases; and whether a subtype should move forward or back.
Weekly control review:
| Review | Questions |
|---|---|
| Outcome | Are accepted outcomes improving by subtype? |
| Quality | Which defects recur, and why? |
| Escalation | Are important cases caught without flooding people? |
| Access | Does the role still have only needed rights? |
| Source | Are authority, freshness, and retention current? |
| Cost | What is cost per accepted outcome and human minute saved? |
| Risk | What was the worst error or near miss? |
| Change | Which version changes require regression testing? |
| Decision | Advance, hold, narrow, roll back, pause, or stop? |
Keep a change log recording the date; component changed; reason; approver; affected subtypes; cases rerun; result; and the current operating mode.
Microsoft’s responsible-agent guidance describes risk-scaled release gates and continuous operation after go-live. Treat the first month as the start of a lifecycle, not a one-time setup.
§ 10Make the Day 30 Decision
Choose one outcome for each task subtype:
| Decision | Meaning | Required action |
|---|---|---|
| Expand | Evidence supports one additional bounded capability | Update contract, controls, cases, and approval |
| Hold | Current mode is useful; next stage lacks evidence | Continue and gather named evidence |
| Narrow | One segment fails while another works | Remove failed subtype/access; rebaseline |
| Roll back | A capability introduced unacceptable failure | Return to prior mode/version |
| Redesign | Failure is structural | Change outcome, source, tool, permission, or owner |
| Pause | Risk, ownership, source integrity, or control is uncertain | Disable triggers/actions; preserve evidence |
| Stop/retire | Role is not valuable, governable, or supportable | Resolve open work, revoke access, retain/delete records |
Require a decision memo
The memo should state the role and version; operating modes reached; the eligible workload; accepted outcomes; quality and correction; escalation; postcondition and reliability; human load; cost; the worst error and incident; failed or excluded cases; remaining assumptions; the decision by subtype; the maximum next authority; and the owner and review date.
Do not describe a role as “successful” without its denominator, exclusions, human work, and worst failure.
Roll back on evidence
Rollback is not failure of the program. It is correct control behavior when a new action misses postconditions; escalation recall drops; source quality changes; correction burden rises; a model/tool change regresses behavior; cost becomes unviable; reviewer capacity disappears; or a severe failure crosses tolerance.
Anthropic’s trustworthy-agents guidance emphasizes human control, transparency, secure interactions, and privacy. A meaningful control system lets people redirect, limit, or stop the worker when conditions change.
§ 11Avoid First-Month Mistakes
Connecting everything on Day 1. More context and tools create more ambiguity, permission surface, and evaluation work. Add one material capability at a time.
Using the calendar as the gate. “It has been two weeks” is not evidence. Advance when the role passes its quality, escalation, reliability, cost, and risk criteria.
Reviewing only polished outputs. Inspect sources, tool calls, approvals, state changes, retries, and handovers.
Correcting work invisibly. Silent edits make a weak role look strong. Record correction type and reviewer minutes.
Training on every reviewer comment. Feedback can conflict or encode a one-off exception. Update authoritative artifacts through an owner and version.
Expanding every task together. Keep authority granular. Research may move to draft while publication stays blocked.
Ignoring the human side. The supervisor, approver, reviewer, and incident owner need their own response and backup path.
Optimizing the average. A severe missed escalation can outweigh many routine successes.
Treating no incident as proof of safety. A small sample may not encounter the failure. Run adverse cases and drills.
Forgetting exit. Test pause, revocation, open-work transfer, evidence preservation, and appropriate deletion before the role becomes difficult to unwind.
Final recommendation
Use the first 30 days to find the narrowest role that produces accepted value inside a defensible operating envelope.
Do not maximize tools, context, tasks, or authority. Maximize what the organization can explain: why work started; which source governed it; what the worker did; which approval applied; whether the result held; what a person corrected; what the role cost; when it escalated; and why the next stage is justified.
§ 12Run the First Month in CellCog
CellCog’s public AI Employees guide describes role configuration, goals, permissions, an inbox, task list, memory, schedules, wake conditions, shifts, and handovers. Its AI Employees product page also describes KPI visibility, approvals, persistent context, and ongoing work.
Map the first-month plan into those concepts:
| First-month need | CellCog concept to configure and verify |
|---|---|
| Role version and outcome | Employee role, goals, KPIs |
| Controlled intake | Inbox, task list, schedule, wake condition |
| Minimum sources | Context and memory |
| Mode boundary | Permissions and approval rules |
| Reviewable work | Task board and artifacts |
| Continuity | Shifts, task state, and handovers |
| Evidence | Logs, outputs, approvals, dashboard records available for the role |
| Pause | Schedule, access, and task controls that stop affected work |
Verify the current product behavior against the job description. In particular, test which context enters the role; which tools and actions are available; how approvals bind to an action; how tasks move between states; how handovers preserve open work; what supervisors can inspect; how a role is paused; and which usage data supports cost per accepted outcome.
Product capability does not replace organizational policy. The role owner still defines sources, permissions, acceptance, escalation, and the Day 30 decision.
Q1How long does it take to onboard an AI employee?
There is no universal duration. Use the first 30 days as a review structure, not a promise. A narrow internal role may pass early stages quickly; a consequential role may need more time, specialist review, narrower scope, or no independent-action stage.
Q2What should an AI employee do in its first week?
It should learn the approved role and source hierarchy, reconstruct task states, rehearse refusal and escalation, and run representative work in observe or shadow mode. It should not receive broad live authority merely because configuration is complete.
Q3When should an AI employee receive tool access?
Give the minimum read access needed for observation and shadow work. Add a write or external action only after the relevant output passes, the action is necessary, approval and limits are enforceable, postconditions are observable, and recovery is tested.
Q4What metrics matter during the first 30 days?
Track accepted outcomes, first-pass quality, material corrections, reviewer minutes, escalation recall and precision, durable completion, postcondition pass, handover reconciliation, total cost per accepted outcome, unauthorized actions, and worst-error severity.
Q5Should the AI employee run independently after 30 days?
Only for allowlisted task-and-action combinations whose evidence supports bounded shifts. Keep new, sensitive, high-impact, irreversible, or hard-to-detect actions human-led or approval-gated. Day 30 is a decision point, not an autonomy deadline.
Q6What should happen if the AI employee performs poorly?
Classify the failure. Repair sources, instructions, tools, permissions, evaluation, or process when appropriate; narrow failed subtypes; roll back authority; or stop the role when value, governability, ownership, or risk cannot meet the approved threshold.
