Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentPricingBlogStoryContact

How to Choose an AI Employee Platform: A 12-Point Evaluation Framework

Napkin-style sketch of a clipboard scorecard with twelve criteria rows beside a row of locked gates, with an amber highlight on one gate labeled approvals
Fig 0A high average score must never compensate for a failed security, approval, recovery, or data-handling requirement.

The best AI employee platform is not the one with the longest feature list. It is the one that can carry your defined role to an accepted outcome, inside your authority and risk boundaries, with evidence you can inspect.

That changes the buying process.

Do not begin with vendors. Begin with one role, one recurring outcome, one representative workload, and the actions the system may take. Then score every platform on the same 12 criteria:

  1. role model and persistence;
  2. triggers and demand control;
  3. context and memory governance;
  4. tool and action capability;
  5. permission and approval granularity;
  6. task state and handovers;
  7. observability and audit evidence;
  8. evaluation and KPI support;
  9. team and delegation architecture;
  10. model capability and output breadth;
  11. pricing and total-cost transparency; and
  12. data, security, privacy, and portability.

Use proof, not promises. A marketing claim is not equal to a configured demonstration. A demonstration is not equal to a representative pilot. And a high average score must never compensate for a failed security, approval, recovery, or data-handling requirement.

Use the scorecard, evidence scale, platform-shape comparison, red flags, and decision rule to evaluate the category before ranking named vendors. Once you have scored your shortlist, use a current AI employee platform comparison to investigate the specific products that fit your winning platform shape.

On this page · 23 sectionsOpen
  1. The 12-Point AI Employee Platform Scorecard
  2. What Should You Define Before Comparing AI Employee Platforms?
  3. How Should You Score Vendor Evidence?
  4. Can the Platform Represent a Persistent Role?
  5. Can the Platform Start the Right Work Without Creating Duplicate or Unbounded Work?
  6. Is Context and Memory Useful, Correctable, and Governed?
  7. Can the Platform Perform the Role’s Real Actions?
  8. Can You Restrict Authority at the Action Level?
  9. Can Work Pause, Resume, Transfer, and End Cleanly?
  10. Can You Reconstruct What Happened?
  11. Can the Platform Measure Role Performance and Regression?
  12. Does the Platform Support the Team You Need?
  13. Is the Model Capability Relevant to the Whole Job?
  14. Can You Forecast Total Cost per Accepted Outcome?
  15. Does the Data and Security Posture Fit the Deployment?
  16. Which Type of AI Employee Platform Fits the Role?
  17. What Should You Ask a Vendor to Prove in the Demo?
  18. What Are the Red Flags When Choosing an AI Employee Platform?
  19. How Does CellCog Map to the 12-Point Framework?
  20. How Do You Make the Final Platform Decision?
  21. A Copyable AI Employee Platform Evaluation Template
  22. AI Employee Platform Selection Checklist
  23. The Short Version
Key points7 · 32 min full read
  1. Define a role before evaluating software. A vendor cannot prove role ownership against an undefined job.
  2. Use the same workload, permissions, failure cases, and acceptance rubric for every platform.
  3. Score 12 criteria with role-specific weights totaling 100, and score evidence from 0 to 5, where 0 is absent and 5 is production-like proof.
  4. Apply non-compensating gates to high-risk actions, access revocation, interruption and recovery, audit evidence, and data fit.
  5. Compare total cost per accepted outcome, not plan price, credits, seats, or successful demo run.
  6. Treat public benchmarks only as evidence for the capability measured.
  7. Shortlist a vendor only when its gates pass, weighted score clears your threshold, and remaining evidence gaps have owners and deadlines.

§ 01The 12-Point AI Employee Platform Scorecard

The weights below are an illustrative starting point for a knowledge-work role that reads internal information, creates deliverables, and takes some external actions. They are not an industry standard.

Move weight toward permissions, audit, and data controls for higher-risk roles. Move it toward tool reach and output breadth for production-heavy roles. Move it toward triggers, task state, and handovers for operational queues.

# Evaluation criterion Sample weight The question to answer
1 Role model and persistence 8 Can the platform preserve a standing responsibility across work sessions?
2 Triggers and demand control 7 Can the right events start work once, within clear limits?
3 Context and memory governance 8 Can it retain useful context without creating an ungoverned memory pool?
4 Tool and action capability 10 Can it complete the required work in the real systems involved?
5 Permissions and approvals 12 Can authority be restricted by action, resource, condition, and risk?
6 Task state and handovers 8 Can work pause, resume, transfer, and finish without losing obligations?
7 Observability and audit evidence 10 Can you reconstruct what happened, why, and with which inputs?
8 Evaluation and KPI support 10 Can you measure accepted outcomes and regressions over time?
9 Team and delegation architecture 5 Can multiple roles coordinate without multiplying ambiguity and access?
10 Model capability and output breadth 7 Can the reasoning engine produce the artifacts and judgments this role needs?
11 Pricing and total-cost transparency 7 Can you forecast cost for the workload and compare accepted outcomes?
12 Data, security, privacy, and portability 8 Does the operating and contractual data posture fit the deployment?
Total 100
Table 1The 12 evaluation criteria with sample weights

The central rule is simple:

Score the platform you can prove, not the product the vendor can describe.

§ 02What Should You Define Before Comparing AI Employee Platforms?

Define the job before the software.

An AI employee platform is only valuable in relation to a bounded responsibility. “Help with marketing” is not a testable responsibility. “Every Monday, assemble the approved channel data, produce the weekly acquisition review in the standard format, identify material changes, and route budget recommendations to the growth lead” is.

If the role itself is still unclear, use the AI employee hiring process to turn the workload into a role contract before opening a vendor spreadsheet.

Write a one-page role card

Role-card field Example: weekly acquisition analyst
Recurring outcome Accepted acquisition review by 10 a.m. each Monday
Trigger Weekly schedule after approved data refresh
Inputs Analytics, ad platforms, CRM, campaign calendar
Required actions Read sources, calculate changes, draft report, update task
Prohibited actions Change budget, edit source data, contact customers
Approval boundary Human approves any spend recommendation
Acceptance evidence Sources linked, calculations checked, template complete
Escalation Missing data, conflicting attribution, abnormal spend
Supervisor Growth lead
Stop condition Two material data-integrity failures in 30 days
Table 2An example role card for a weekly acquisition analyst

The platform must support this operating contract. It does not earn credit for an impressive capability the role does not use.

Build a representative test pack

Use 10-20 examples drawn from real work, sanitized where required. Include:

  • ordinary tasks;
  • incomplete inputs;
  • ambiguous requests;
  • conflicting sources;
  • a denied action;
  • a tool or API failure;
  • a duplicate trigger;
  • a stale instruction;
  • a high-risk action that must escalate; and
  • a task that should be stopped rather than completed.

The test pack prevents a vendor from choosing only favorable examples.

Separate must-pass gates from scored preferences

Some requirements are not tradeable.

Gate Example requirement Why an average cannot compensate
Authority gate A spend change always requires enforceable approval Strong writing does not offset unauthorized spend
Access gate Credentials and connectors can be revoked centrally A departing owner or incident needs an immediate stop
Recovery gate A run can be interrupted and a failed action contained Autonomous retries can compound damage
Evidence gate Actions and approvals can be reconstructed An untraceable outcome cannot be governed
Data gate Retention, subprocessors, storage, and deletion fit policy Business utility does not waive legal or contractual limits
Table 3Five non-compensating gates and why an average cannot save them

Reject or pause a vendor that fails a gate. Do not let a 90 in model quality erase a zero in approval enforcement.

§ 03How Should You Score Vendor Evidence?

Every score needs both a number and a receipt.

Use this six-level evidence scale:

Score Evidence level What qualifies What does not
0 Absent or contradicted Requirement is unavailable, prohibited, or disproved “On the roadmap”
1 Marketing claim Vendor states the capability Landing-page language or sales assurance
2 Documented or demonstrated Current documentation or a vendor-run happy path shows it Slides without product behavior
3 Buyer-configured proof Your team configures it and tests an ordinary case Vendor-controlled example
4 Representative pilot Your workload, tools, failures, and acceptance rubric are exercised One polished task
5 Production-like proof Repeated representative work with exported evidence and stable results A larger demo without operational controls
Table 4The 0-5 evidence scale, from absent to production-like proof

Documentation can prove that a control exists. It cannot prove that your configuration works under your conditions.

Calculate the weighted score

For each criterion:

weighted points = criterion weight × evidence score ÷ 5

Then add the 12 weighted results.

Suppose one platform receives these evidence scores:

Criterion Weight Evidence score Weighted points
Role model 8 4 6.4
Triggers 7 4 5.6
Memory 8 3 4.8
Tools 10 4 8.0
Permissions 12 3 7.2
Task state 8 4 6.4
Observability 10 3 6.0
Evaluation 10 3 6.0
Teams 5 2 2.0
Model capability 7 4 5.6
Cost 7 3 4.2
Data posture 8 4 6.4
Total 100 68.6
Table 5A worked weighted-score example

A 68.6 is not automatically good or bad. Your policy decides.

An illustrative decision policy might be:

  • 80-100: shortlist if every gate passes;
  • 65-79.9: continue only through a bounded evidence plan;
  • below 65: do not advance for this role; and
  • any failed gate: do not deploy, regardless of total.

Publish the weights, scores, evidence links, evaluator, and date. Otherwise the spreadsheet becomes a record of preference disguised as analysis.

§ 04Can the Platform Represent a Persistent Role?

Criterion 1 asks whether the product can hold a standing responsibility - not merely produce a strong response.

If your use case is one-off analysis, a capable assistant may be enough. The AI employee definition becomes relevant when work needs a durable role, continuity, initiative, action, and accountability.

What to evaluate

Look for configurable objects or behaviors covering:

  • role purpose;
  • goals and boundaries;
  • supervisor;
  • recurring responsibilities;
  • identity or communication channel;
  • schedule or availability;
  • open tasks;
  • role-specific instructions;
  • context carried between work sessions;
  • performance expectations; and
  • retirement or reassignment.

The platform does not need to copy a human HR system. It does need a durable place for the role’s obligations to live.

Ask for a restart test

Start a multi-step task. Interrupt it. End the session. Resume later from a different interface or scheduled run.

The system should recover:

  • what outcome it owes;
  • what has already happened;
  • which source is authoritative;
  • what remains blocked;
  • what action is next;
  • which approval is still valid; and
  • when the task should stop.

Conversation history alone is weak evidence. It may preserve words without preserving an explicit operating state.

Score role ownership, not anthropomorphic presentation

Weak signal Stronger signal
The agent has a human name The role has durable responsibilities and boundaries
The interface has an avatar Work state persists across sessions
The vendor says “digital worker” A trigger, outcome, supervisor, and stop condition are configurable
The model speaks proactively The role knows which obligation to advance
The product shows a persona The role can be reassigned, narrowed, paused, or retired
Table 6Weak signals versus stronger signals of role persistence

A friendly identity can help collaboration. It is not evidence of operational persistence.

§ 05Can the Platform Start the Right Work Without Creating Duplicate or Unbounded Work?

Criterion 2 covers triggers and demand control.

An AI employee may wake on a schedule, event, email, message, queue item, webhook, threshold, or direct request. The important question is not whether a trigger exists. It is whether the trigger starts the right work exactly as intended.

Evaluate trigger coverage

Trigger type Useful for Failure to test
Schedule Reports, reviews, monitoring Time-zone drift, missed runs, holiday rules
Email or message Support, coordination, intake Loops, spoofing, duplicate threads
Application event CRM, ticketing, operations Replayed or out-of-order events
Queue item Bounded operational work Lost priority or double ownership
Threshold Anomaly or exception detection Noisy repeated alerts
Human request Ad hoc escalation or assignment Ambiguous authority
Table 7Trigger types and the failure each must be tested against

Demand control is part of reliability

Ask whether the system supports:

  • deduplication;
  • idempotency;
  • rate limits;
  • concurrency limits;
  • retry caps;
  • backoff;
  • budget caps;
  • quiet hours;
  • priority rules;
  • stale-event handling;
  • cancellation; and
  • a dead-letter or exception path.

“Runs 24/7” is a capacity claim. It is not a control design.

Run a duplicate-trigger test

Send the same event twice. Delay a prerequisite. Then deliver the prerequisite after the original task has expired.

A credible platform should not:

  • send two customer messages;
  • create two records;
  • continue an obsolete task;
  • retry indefinitely; or
  • hide which event produced which action.

Score the vendor on the behavior you observe and the controls you can configure.

§ 06Is Context and Memory Useful, Correctable, and Governed?

Criterion 3 covers what the AI employee knows across work sessions.

Persistent memory can reduce repeated prompting. It can also preserve a false fact, stale policy, confidential detail, or malicious instruction. The buying question is not “Does it have memory?” It is “Can we govern what becomes memory, where it applies, when it expires, and how it is corrected?”

The full design problem is covered in the guide to AI employee memory. At platform-selection stage, test the minimum controls directly.

Distinguish five context layers

Context layer Example Required control
Role instructions “Never approve a refund” Version, owner, precedence
Authoritative knowledge Refund policy Source, freshness, access scope
Working context Current customer case Task or account isolation
Learned preference Preferred report format Correction and provenance
Episodic record Prior action and result Retention, retrieval, audit
Table 8Five context layers and the control each requires

A single undifferentiated “memory” box makes conflicts difficult to resolve.

Test correction, conflict, and expiry

Give the platform an old policy, then provide a newer authoritative version. Ask it to perform a task where the two conflict.

The system should:

  • prefer the defined authoritative source;
  • expose the conflict;
  • avoid silently blending incompatible rules;
  • show which source affected the decision;
  • let an authorized user correct or remove the bad record; and
  • prevent one customer or role’s context from leaking into another.

Ask what memory is not

Vendors often combine the following objects under one label:

  • chat history;
  • retrieved documents;
  • vector search;
  • generated summaries;
  • saved preferences;
  • task state; and
  • long-term learned records.

Ask which layer stores each object, who can see it, how long it remains, whether it is used for model training, and what deletion actually removes. A memory feature without a lifecycle is a future incident queue.

§ 07Can the Platform Perform the Role’s Real Actions?

Criterion 4 covers tools and action capability.

A platform that can draft a perfect update but cannot read the approved data, write to the task system, or send through the governed channel may still leave the human as workflow operator.

List the role’s required actions before evaluating connectors.

Build an action inventory

Action System Mode Expected proof
Read campaign data Analytics platform Read Correct account and date range
Read opportunity stage CRM Read Correct field and latest value
Create weekly review Document system Write Correct template and folder
Update task status Work tracker Write Valid transition and evidence link
Propose budget change Approval queue Propose No direct spend mutation
Notify growth lead Email or chat Send Correct recipient, thread, and summary
Table 9The action inventory: system, mode, and expected proof

“Connects to the CRM” earns no credit if the role needs an unsupported object or write operation.

Evaluate four action modes

  1. API or native connector: structured, often easier to scope and validate.
  2. Browser or computer use: broad reach, but interface changes and session state matter.
  3. Code execution: flexible, but sandboxing, secrets, dependencies, and output validation matter.
  4. Artifact generation: documents, spreadsheets, presentations, code, media, or dashboards.

One mode is not universally superior. Score the mode against the specific task, evidence needs, and failure exposure.

Test the last mile

Ask the vendor to complete the full chain:

  1. identify the correct source;
  2. retrieve the correct record;
  3. transform or reason over it;
  4. create the required artifact;
  5. place it in the correct destination;
  6. update task state;
  7. request approval if needed; and
  8. return evidence.

Many demos stop at step 4. Operational value often depends on steps 5-8.

§ 08Can You Restrict Authority at the Action Level?

Criterion 5 receives the highest sample weight because tools turn model errors into business actions.

A connected account is not a permission model. “The agent can use Gmail” does not tell you whether it can read, draft, send, delete, change settings, contact any recipient, or act only inside an approved thread.

Use the detailed AI employee permissions and approvals framework to design the role. During vendor selection, require the platform to implement that design.

Ask for an action envelope

Dimension Example restriction
Tool CRM only
Operation Read and propose; no delete
Resource Accounts tagged pilot
Recipient Approved customer domain
Amount Refund proposal below $100
Time Weekdays, 8 a.m.-6 p.m.
Volume Maximum 20 external messages per day
Data class No restricted financial data
Approval Named manager, exact action preview
Expiry Permission ends after 14 days
Table 10Ten dimensions of an enforceable action envelope

The platform should enforce the boundary. A prompt that says “please do not” is an instruction, not an authorization control.

Test denial and revocation

Run three cases:

  1. ask the role to perform an explicitly denied action;
  2. revoke a connector or permission while a task is open; and
  3. alter the target or parameters after an approval is granted.

Expected behavior:

  • the denied action does not execute;
  • revocation takes effect predictably;
  • the task moves to a visible blocked state;
  • the approval applies only to the exact reviewed action;
  • replay or parameter substitution is rejected; and
  • the event appears in the audit record.

OpenAI’s agent guide recommends pairing guardrails with authentication, authorization, access controls, and normal software security measures. It also identifies high-risk or irreversible actions and repeated failure as reasons for human intervention.

OWASP’s AI Agent Security Cheat Sheet similarly recommends least-privilege tools, explicit approval for high-impact actions, action previews, audit trails, interruption, and separate validation of irreversible execution.

Score the enforced path

Evidence Suggested score ceiling
“Enterprise-grade permissions” claim 1
Documented connector scopes 2
Buyer configures read versus write 3
Pilot proves denial, approval, expiry, and revocation 4
Repeated production-like evidence plus audit export 5
Table 11Evidence and the score ceiling it supports

Do not award points for controls the vendor says it can custom-build later.

§ 09Can Work Pause, Resume, Transfer, and End Cleanly?

Criterion 6 covers task state and handovers.

Long-running work does not move directly from started to done. It waits for data, approval, a reply, a scheduled follow-up, another role, or a changed condition.

Require explicit task states

A useful minimum is:

  • queued;
  • in progress;
  • waiting on external event;
  • waiting on human;
  • blocked;
  • completed;
  • failed;
  • canceled; and
  • reopened.

The exact labels can differ. The state transitions should not be hidden inside prose.

Inspect the handover object

A handover should carry:

  • original outcome;
  • work completed;
  • evidence created;
  • decisions made;
  • source versions used;
  • actions taken;
  • permissions consumed;
  • unresolved questions;
  • current status;
  • next action;
  • next owner; and
  • deadline or wake condition.

“Here is what I did” is a summary. “Here is the accepted state from which the next actor can continue” is a handover.

Test three interruptions

  1. Session interruption: stop midway and resume later.
  2. Human dependency: require approval, then approve with a modification.
  3. Role transfer: move the task to another agent or a person.

The platform should not repeat completed actions, lose the modified approval, or declare success before the postcondition is true.

Define completion as a postcondition

For a weekly report, document generated may be insufficient. Completion might require:

  • correct reporting period;
  • all required sources present;
  • calculations reconciled;
  • report stored in the approved folder;
  • owner notified;
  • evidence attached; and
  • no unresolved critical exception.

Task state is how a platform turns impressive outputs into dependable operations.

§ 10Can You Reconstruct What Happened?

Criterion 7 covers observability and audit evidence.

An AI employee can make decisions across many tool calls and work sessions. A supervisor needs more than a final answer and a token total.

Require three evidence layers

Layer Buyer question Useful fields
Operational Is work moving? task, owner, state, duration, retry, queue
Decision Why did it choose this path? instruction version, sources, confidence, exception
Action What changed in the world? actor, tool, operation, target, parameters, result, approval
Table 12Operational, decision, and action evidence layers

A polished natural-language recap can help humans scan. It should not be the only record.

Ask for correlation and export

You should be able to connect:

trigger → run → source → decision → approval → tool action → result → task state

Test whether the platform can:

  • search by task, customer, role, action, and time;
  • correlate a tool action to the triggering event;
  • distinguish proposed from executed action;
  • show retries and partial failures;
  • preserve approval evidence;
  • export records;
  • restrict log access;
  • redact sensitive fields; and
  • apply a retention policy.

Test a failed run, not just a successful one

Disable a tool after a run has been configured to use it. Then ask:

  • Did the run retry?
  • How many times?
  • Did any partial action succeed?
  • Was the task status correct?
  • Did the user receive the right escalation?
  • Can an evaluator reconstruct the sequence without asking the agent to remember it?

NIST’s AI Risk Management Framework Core treats documentation, ongoing monitoring, production behavior, incident response, recovery, override, and change management as lifecycle concerns - not an end-of-project checklist.

§ 11Can the Platform Measure Role Performance and Regression?

Criterion 8 separates activity from accepted work.

Runs, prompts, tokens, messages, and documents are operational telemetry. They do not show whether the role produced useful outcomes.

Define an accepted outcome

For each task type, specify:

  • required fields;
  • quality rubric;
  • source requirements;
  • allowed actions;
  • approval state;
  • completion evidence;
  • exception rule; and
  • reviewer.

Then measure:

Metric What it reveals Common distortion
Acceptance rate Share of outputs accepted against rubric Easy tasks inflate the rate
First-pass acceptance Quality before human correction Reviewer standards drift
Correction time Human effort required after generation Editing can be hidden outside platform
Cycle time Speed from valid trigger to accepted result Waiting time may be misclassified
Reopen rate Premature or fragile completion Reopened work counted as new work
Escalation precision Whether the right cases reach humans Avoiding all work can look “safe”
Action error rate Incorrect or unauthorized executions Near misses go unrecorded
Cost per accepted outcome Economic efficiency of usable work Failed and reviewed work omitted
Table 13Eight role metrics and the distortion each is vulnerable to

Require a regression set

The platform, model, tools, prompts, policies, and source data will change. Keep a versioned test set containing ordinary cases, known failures, and abuse cases.

Run it:

  • before deployment;
  • after a model or prompt change;
  • after a connector or permission change;
  • after a memory or retrieval change;
  • after an incident; and
  • on a regular schedule.

OpenAI’s practical guide advises teams to establish a performance baseline with evaluations before optimizing model cost and latency. NIST likewise calls for documented test sets and performance evidence under conditions similar to deployment.

Separate model evaluation from role evaluation

A reasoning benchmark can show that an engine performs well on a defined set of research tasks. It does not prove:

  • correct account access;
  • reliable event handling;
  • action authorization;
  • governed memory;
  • handover quality;
  • recovery from your tool failures;
  • accepted business outcomes; or
  • predictable unit economics.

Credit a benchmark only inside the criterion it actually measures. Then run the role test.

§ 12Does the Platform Support the Team You Need?

Criterion 9 covers delegation and multi-role coordination.

Do not buy a multi-agent architecture because the diagram looks organizational. A second role is useful only when specialization, independent verification, access separation, or workload routing improves the outcome enough to justify more coordination.

Anthropic’s guidance on building effective agents recommends adding complexity only when it demonstrably improves results. It also notes that autonomy can increase cost and compound errors.

Compare three coordination patterns

Pattern Best fit Evaluation question
One general role Coherent workload with shared context Can one role meet the rubric reliably?
Manager plus specialists Distinct skills or parallel work Are delegation, acceptance, and escalation explicit?
Independent maker and checker High-value verification Is the checker genuinely independent and empowered to reject?
Table 14Three coordination patterns and the question each must answer

Use the general-purpose versus specialized AI agent architecture guide to score task variance, context, tool choice, permission separation, evaluation, coordination, and accepted-system-outcome cost before awarding points for team features.

Test delegation as a contract

When one role delegates to another, inspect:

  • task identity;
  • requested outcome;
  • relevant context;
  • data and tool boundary;
  • deadline;
  • acceptance rubric;
  • authority;
  • evidence returned;
  • failure path; and
  • final accountable owner.

A message between agents is not a governed handover.

Check blast-radius control

Multi-agent systems can multiply access and error propagation. Ask whether:

  • roles have separate permissions;
  • one role can grant authority to another;
  • shared memory is scoped;
  • inter-agent messages are attributable;
  • a compromised task can spread;
  • recursion and delegation depth are capped; and
  • the whole team can be paused centrally.

If the intended role works as one bounded agent, team features should have a low weight. Buy the simplest architecture that proves the outcome.

If the first role later exposes a stable reason to add workers or a manager, design the operating model before expanding the org chart. The AI organization framework covers ownership, interfaces, shared state, escalation, supervision, controls, and coordination economics.

§ 13Is the Model Capability Relevant to the Whole Job?

Criterion 10 covers the reasoning and production engine under the employee scaffolding.

Model quality matters. A persistent role built on weak reasoning will persistently produce weak work. But “uses the latest models” is not an evaluation result.

Map capabilities to role stages

Role stage Capability to test
Interpret request Ambiguity handling and instruction following
Gather evidence Search, retrieval, source discrimination
Plan Decomposition, prioritization, stop conditions
Produce Writing, analysis, code, data, document, or media quality
Act Correct tool selection and parameters
Verify Self-check, reconciliation, test execution
Escalate Uncertainty recognition and concise handoff
Table 15Role stages and the capability each stage tests

A general-purpose platform may be valuable when one role needs research, spreadsheets, documents, code, images, or dashboards in the same workflow. A specialized platform may outperform it when one narrow task benefits from proprietary data, dedicated interfaces, and deeply tuned workflows.

Ask how models are selected

Evaluate:

  • fixed versus configurable model;
  • routing by task or risk;
  • fallback behavior;
  • latency and cost tradeoffs;
  • context limits;
  • supported modalities;
  • model-change notification;
  • regression testing after upgrades;
  • provider and regional constraints; and
  • whether model choice changes data terms.

Use benchmark evidence precisely

For every benchmark, record:

  • task set;
  • sample size;
  • date;
  • judge;
  • metric;
  • tested product or model;
  • configuration;
  • public reproducibility;
  • relationship to your workflow; and
  • what the benchmark does not measure.

“Ranked first in research” supports a research-capability claim when the result is current and reproducible. It does not support “best AI employee platform” without action, governance, reliability, and economic evidence.

§ 14Can You Forecast Total Cost per Accepted Outcome?

Criterion 11 covers pricing and total cost.

Plan price is only one term. Seats, credits, tasks, model tokens, runtime, tool calls, storage, concurrency, and successful outcomes can all be billing units.

Use the full AI employee cost framework when building a business case. During vendor selection, normalize every candidate against the same workload.

Calculate the complete operating cost

monthly operating cost = subscription + usage + integration + supervision + correction + monitoring + failure exposure

Then:

cost per accepted outcome = monthly operating cost ÷ accepted outcomes

Illustrative example:

Cost component Platform A Platform B
Subscription and usage $700 $1,100
Integration and maintenance $300 $150
Human review and correction $1,200 $500
Monitoring and exception handling $250 $200
Expected rework or failure cost $350 $150
Monthly operating cost $2,800 $2,100
Accepted outcomes 140 120
Cost per accepted outcome $20.00 $17.50
Table 16Two platforms, one workload - vendor bill versus unit economics

Platform A has the lower vendor bill. Platform B has the lower cost per accepted result.

The numbers are illustrative, not market benchmarks.

Run a fixed-workload pricing test

Give each vendor the same:

  • 20 ordinary tasks;
  • five hard tasks;
  • three failure cases;
  • output quality requirement;
  • review standard;
  • tool set;
  • concurrency;
  • context volume; and
  • time window.

Record billed units, failed attempts, retries, reviewer time, and accepted outcomes.

Ask about the edges

Verify:

  • minimum commitment;
  • included usage;
  • overage;
  • credit expiry;
  • premium model multipliers;
  • tool or connector fees;
  • storage;
  • background or scheduled-run charges;
  • failed-run charges;
  • retry charges;
  • support tier;
  • implementation services;
  • cancellation;
  • data export; and
  • price-change notice.

Pricing pages change. Check the vendor’s live terms on the decision date. CellCog, for example, currently publishes credit-based pricing and explains that more complex operations use more credits; a buyer should still measure credits against a representative role workload.

§ 15Does the Data and Security Posture Fit the Deployment?

Criterion 12 covers the complete information path, not only the base model.

An AI employee may receive data from email, documents, databases, browser sessions, connected apps, human prompts, long-term memory, logs, and other agents. Map all of it.

Build a data-flow record

Data question Evidence to request
What enters the system? Data inventory and classification
Where is it processed and stored? Regions, architecture, subprocessors
Who can access it? Roles, support access, tenant isolation
Is it used for training? Contract and provider-specific terms
How long is it retained? Object-level retention schedule
How is it deleted? Deletion scope, timing, backups, subprocessors
What appears in logs and memory? Redaction and access controls
How does data leave? Connectors, actions, exports, external messages
What happens on termination? Export format, deletion, revocation
What happens during an incident? Notification, containment, recovery
Table 17Ten data questions and the evidence to request

Evaluate the whole supply chain

Ask about:

  • hosting provider;
  • model providers;
  • connector or integration brokers;
  • email, voice, or messaging providers;
  • analytics and error monitoring;
  • support tooling;
  • data residency;
  • encryption;
  • identity and access;
  • tenant isolation;
  • backups;
  • vulnerability management;
  • incident response; and
  • contractual commitments.

Certifications can provide useful third-party evidence. They do not replace your workflow-specific review.

Test portability and exit

Before purchase, ask the vendor to show how you export:

  • role instructions;
  • task history;
  • memory or learned records;
  • generated artifacts;
  • evaluation sets and results;
  • audit evidence;
  • connector configuration; and
  • usage and cost data.

Then ask what is deleted after termination and on what timeline.

A system can be easy to start and difficult to leave. Portability belongs in selection, not offboarding.

§ 16Which Type of AI Employee Platform Fits the Role?

The category contains different product shapes. Compare the shape before comparing logos.

Platform shape Best when Common strength Common buying risk
Assistant-first product A person remains in the loop Flexible collaboration and broad capability Human still carries recurring state
Workflow builder Path and rules are mostly known Predictability and explicit orchestration Brittle around ambiguity and changing work
Specialized AI employee One function dominates Deep role-specific workflow and integrations Narrow fit, less reusable outside the role
General-purpose AI employee platform A role spans research, artifacts, actions, and follow-through One operating layer across varied work Breadth may require more buyer configuration
Internal or developer framework Control and proprietary workflow justify engineering Custom architecture and data boundary Build, evaluation, maintenance, and on-call burden
Table 18Five platform shapes, their strengths, and their buying risks

If the shortlist mixes a managed platform with an internal architecture, run the build-versus-buy AI employee lifecycle comparison before assigning a platform score. It prevents a prototype estimate from being compared with a subscription that already includes parts of the runtime, connector, memory, and operating layer.

Choose an assistant-first product when

  • work is irregular;
  • the user benefits from constant collaboration;
  • the output is advisory;
  • a person naturally owns follow-through; and
  • standing access would add more risk than value.

Choose a workflow builder when

  • the sequence is known;
  • rules can be expressed deterministically;
  • inputs and outputs are structured;
  • exceptions are limited; and
  • predictability matters more than open-ended judgment.

Choose a specialized platform when

  • one high-volume role drives the business case;
  • domain data or workflow depth matters;
  • the product has evidence on your exact channel and outcome; and
  • specialization offsets vendor concentration.

Choose a general-purpose AI employee platform when

  • the role crosses applications and artifact types;
  • work is recurring but not fully deterministic;
  • persistent task state and follow-through matter;
  • one role may later coordinate with others; and
  • the buyer wants to configure jobs rather than build agent infrastructure.

Choose an internal build when

  • the workflow is a durable competitive advantage;
  • data or execution constraints cannot be met by a vendor;
  • the engineering team can own evaluation, observability, security, and upgrades;
  • time to first value is less important; and
  • the three-year ownership cost is justified.

Platform shape narrows the field. The 12-point scorecard chooses within it.

§ 17What Should You Ask a Vendor to Prove in the Demo?

A demo should be a controlled evidence session, not a product tour.

This is a shortlist test - not the complete pilot design.

Send the vendor the same six cases

  1. Normal case: complete the ordinary workflow.
  2. Missing-input case: identify the missing prerequisite and stop correctly.
  3. Conflicting-source case: apply the defined source hierarchy.
  4. Denied-action case: refuse an action outside the role’s authority.
  5. Tool-failure case: contain partial work, cap retries, and escalate.
  6. High-risk case: preview the exact action and route approval.

Watch the operator, not only the output

Ask the vendor to show:

  • role configuration;
  • trigger setup;
  • memory source and correction;
  • tool permission;
  • approval policy;
  • task state;
  • run evidence;
  • failure record;
  • cost or usage record; and
  • export.

If the vendor hides setup behind a services team, price and score that dependency.

Capture proof in a decision log

Field Record
Requirement Exact criterion and role need
Evidence URL, screenshot, export, configuration, or test result
Evidence level 0-5
Limitation What remains unproved
Owner Buyer or vendor
Due date Before shortlist, pilot, or deployment
Decision effect Gate, score, or informational
Table 19The decision-log fields for every requirement

A later procurement process may ask dozens of legal, technical, and commercial questions. Use the 50-question AI employee platform RFP after this shortlist test to request architecture, configuration, logs, tests, assurance, pricing, and contract evidence without turning the first demo into a generic security questionnaire.

Once mandatory proof gaps are closed, use the AI employee pilot decision framework to freeze the baseline, representative sample, staged authority, KPI thresholds, budget, and go/revise/switch/stop rules before scored work begins.

§ 18What Are the Red Flags When Choosing an AI Employee Platform?

Most red flags are evidence problems, not missing buzzwords.

Red flag Why it matters What to request
“Unlimited autonomy” No operating boundary is specified Action envelope and stop policy
Permissions described only per connector One connector contains many risky operations Operation- and resource-level enforcement
Memory cannot be inspected or corrected Errors can persist across work Provenance, edit, deletion, scope
Logs are only narrative summaries Sequence and responsibility are hard to reconstruct Structured correlated action record
No explicit task states Waiting and failure disappear into chat State model and transition evidence
Scheduled work has no duplicate protection Replayed triggers can repeat actions Idempotency and cancellation test
Evaluation means one public benchmark Role reliability remains unmeasured Representative test set and KPI record
Pricing hides retries or background runs Unit cost becomes unpredictable Fixed-workload usage report
Data language is vague Retention and subprocessors remain unknown Contractual data-flow answers
No export or deletion demonstration Exit risk is undiscovered Sample export and deletion procedure
Model upgrade happens without notice or regression Performance can change silently Change policy and versioned eval
Table 20Eleven red flags, why each matters, and what to request

Beware the feature-parity trap

Two vendors may both list memory, tools, approvals, and schedules. They can still differ dramatically in:

  • granularity;
  • defaults;
  • configuration burden;
  • evidence;
  • failure behavior;
  • scale;
  • operator usability; and
  • contractual commitment.

Score the operating behavior.

Beware the demo-quality trap

A good result can come from:

  • vendor-selected data;
  • hidden preparation;
  • expert prompt engineering;
  • a manually repaired run;
  • broad permissions;
  • unlimited retries; or
  • an expensive model configuration.

Ask to reproduce the result in your environment with visible settings and billing.

Beware the average-score trap

A platform with excellent output breadth and low price can still be wrong for a role if it cannot:

  • enforce the approval boundary;
  • recover from a partial action;
  • prevent cross-account context leakage;
  • meet retention policy; or
  • show what it changed.

Gates protect the buyer from averaging away existential risk.

§ 19How Does CellCog Map to the 12-Point Framework?

CellCog positions itself as a general-purpose AI employee platform. Its current public materials provide evidence for several scorecard areas, but a buyer should still run the same role-specific proof required of every vendor.

The assessment below reflects public pages reviewed in July 2026. Product behavior, terms, and documentation can change.

What the public product material establishes

Criterion Publicly described CellCog evidence Buyer should still verify
Role and persistence Roles with goals, an inbox, schedule, memory, tasks, shifts, and wake conditions Restart behavior, role versioning, retirement
Triggers Email, Slack message, schedule, and event wake conditions Dedupe, retry, concurrency, cancellation
Memory Persistent context across shifts Provenance, correction, scope, expiry, export
Tools and actions Work on files, browser sessions, and connected apps Exact operations, connector limits, failure containment
Permissions and approvals User-defined permissions and approval routing are described Action/resource granularity, revocation, approval binding
Task state and handovers Task board states, shifts, and “dear next-me” handovers are shown Reopen behavior, partial failure, transfer evidence
Observability Task and KPI surfaces are described Structured action logs, correlation, export, retention
Evaluation and KPIs Employees can maintain KPI dashboards Regression sets, accepted-outcome definitions, evaluation export
Teams and delegation Employees can hand off work and form teams Delegation depth, permission propagation, team pause
Model capability Research, data, code, documents, media, dashboards, and apps are described Performance on the buyer’s complete role
Pricing Public credit-based pricing and additional-credit terms Fixed-workload consumption and all implementation costs
Data posture A public privacy policy describes categories, providers, storage, retention, and rights Contractual fit, identity controls, assurance, deletion evidence
Table 21CellCog’s public evidence per criterion and what buyers should still verify

CellCog’s AI Employees page describes an employee with goals, permissions, an inbox, task board, persistent memory, shifts, event-based wake conditions, KPI dashboards, tools, approvals, and handovers. These are relevant category signals. They are not a substitute for your configured denial, failure, and outcome tests.

How to interpret the benchmark

CellCog’s benchmark page reported in July 2026 that CellCog Max ranked first on the public DeepResearch Bench leaderboard, with a 55.78 overall score under the stated judge and methodology.

That is relevant evidence for research capability. It does not by itself establish:

  • reliable scheduling;
  • correct connector use;
  • permission enforcement;
  • handovers;
  • audit completeness;
  • role-level acceptance;
  • security fit; or
  • cost per accepted outcome.

Score it under model capability, verify the public leaderboard, and then test the role.

How to interpret the privacy material

CellCog’s current privacy policy describes AI employee configuration, mailbox content, memory and work artifacts, connected-account actions, service providers, US-based hosting, encryption, retention, and access, correction, deletion, and portability rights.

That transparency helps a buyer map the data flow. A procurement team should still verify how the policy and contract apply to its plan, users, data classes, regions, deletion requirements, authentication model, and assurance needs.

The fairest CellCog shortlist test

Configure one real role. Then require CellCog to:

  1. wake from a normal trigger;
  2. retrieve the correct context;
  3. complete a multi-modal work product if the role needs one;
  4. attempt and refuse a denied action;
  5. request an exact approval for a high-risk action;
  6. survive a connector failure;
  7. hand over an unfinished task;
  8. show the evidence chain;
  9. report role KPI and usage; and
  10. export the records the buyer needs.

If the role passes its gates and score threshold, move CellCog into the named-vendor comparison. If it does not, the framework should say so.

§ 20How Do You Make the Final Platform Decision?

Use a three-part decision rule:

advance = all gates pass AND weighted score clears threshold AND open risks have an accepted treatment

First, compare evidence quality

Do not compare one vendor’s pilot score with another vendor’s marketing score as if both numbers mean the same thing.

Your final table should include:

Vendor Shape Gates Weighted score Lowest evidence areas Monthly role cost Decision
Platform A General employee Pass 84.2 Data export, retry policy $2,400 Shortlist
Platform B Workflow builder Pass 79.6 Open-ended reasoning $1,900 Pilot only
Platform C Specialist Fail: approval 88.0 Approval enforcement $2,100 Do not advance
Scroll to compare all columns
Table 22An illustrative final decision table

These are illustrative results.

Second, inspect the score shape

Two vendors with 82 points may create different decisions.

  • One may be balanced across every criterion.
  • One may dominate model capability and tools but trail on evidence and control.
  • One may fit the first role but not the likely second role.
  • One may require more implementation work but create stronger long-term ownership.

Review the individual criteria, not just the total.

Third, write a reversible decision

A good decision memo states:

  • selected role;
  • selected platform and shape;
  • rejected alternatives;
  • gates and evidence;
  • score and weights;
  • known limitations;
  • allowed deployment boundary;
  • supervisor;
  • review date;
  • expansion conditions;
  • pause conditions; and
  • exit path.

The first purchase should authorize a bounded role, not declare a universal platform winner.

§ 21A Copyable AI Employee Platform Evaluation Template

Use this structure for each vendor:

Role and decision context

  • Role:
  • Recurring outcome:
  • Supervisor:
  • Systems:
  • Allowed actions:
  • Prohibited actions:
  • Approval boundary:
  • Acceptance rubric:
  • Monthly volume:
  • Risk level:

Non-compensating gates

  • High-risk actions require enforceable approval.
  • Access can be revoked centrally and promptly.
  • Runs can be interrupted; retries and spend are bounded.
  • Actions, approvals, and task state are reconstructable.
  • Data processing, retention, deletion, and contract fit policy.

Weighted score

Criterion Weight Evidence score 0-5 Weighted points Evidence link Open gap
Role model and persistence
Triggers and demand control
Context and memory
Tools and actions
Permissions and approvals
Task state and handovers
Observability and audit
Evaluations and KPIs
Teams and delegation
Model capability
Pricing and total cost
Data, security, and portability
Total 100
Scroll to compare all columns
Table 23The scoring template

Decision

  • Gates:
  • Weighted score:
  • Cost per accepted outcome:
  • Evidence gaps:
  • Risk treatments:
  • Decision:
  • Decision owner:
  • Valid until:

§ 22AI Employee Platform Selection Checklist

Role readiness

  • One recurring outcome is defined.
  • Inputs and authoritative sources are named.
  • Required and prohibited actions are listed.
  • The supervisor and escalation path are explicit.
  • Acceptance evidence and stop conditions are written.
  • A representative test pack exists.

Evidence design

  • Every vendor receives the same cases.
  • Evidence scores use one 0-5 definition.
  • Weights total 100 and reflect role risk.
  • Gate failures cannot be averaged away.
  • Scores include a link or artifact.
  • Evidence dates and evaluators are recorded.

Operating capability

  • The role resumes after interruption.
  • Duplicate triggers do not duplicate actions.
  • Memory can be inspected, corrected, and scoped.
  • Required tools support exact operations.
  • Denied actions are technically blocked.
  • Approvals bind to the reviewed action.
  • Connector revocation changes behavior.
  • Task states cover waiting, failure, cancellation, and reopen.
  • Handovers preserve obligations and evidence.

Reliability and measurement

  • Runs correlate triggers, decisions, approvals, actions, and results.
  • Failed and partial runs are visible.
  • Logs are searchable and exportable.
  • Accepted outcomes are defined.
  • Review and correction time are counted.
  • Regression tests cover ordinary, edge, and abuse cases.
  • Model or configuration changes trigger reevaluation.

Economics and governance

  • Pricing is normalized against one workload.
  • Failed runs, retries, background work, and review are included.
  • Cost per accepted outcome is calculated.
  • Data flow and subprocessors are mapped.
  • Retention, deletion, and training use are understood.
  • Identity, access, incident, and recovery controls fit the role.
  • Role data and evidence can be exported.
  • Expansion, pause, and exit conditions are documented.

§ 23The Short Version

Choose an AI employee platform in this order:

  1. define one role;
  2. write the action and risk boundary;
  3. create a representative test pack;
  4. set non-compensating gates;
  5. assign role-specific weights to the 12 criteria;
  6. score evidence from claim to production-like proof;
  7. compare platform shapes;
  8. normalize total cost per accepted outcome;
  9. document limitations and risk treatments; and
  10. authorize only a bounded first deployment.

The point is not to find the vendor with the most AI. It is to find the operating system that can own the right work, prove what happened, and return control when the boundary is reached.

If CellCog fits your winning platform shape, inspect its AI Employee operating layer and score it with the template above. Then compare it with the relevant alternatives on the AI employee platform comparison page.

Frequently asked6 questions

Q1What is an AI employee platform?

An AI employee platform is software for assigning AI a standing, bounded role. Beyond generating answers, it may provide triggers, persistent context, task state, tools, permissions, approvals, handovers, measurement, and supervision. Product labels vary, so evaluate the operating behavior rather than the name.

Q2What is the best AI employee platform?

The best platform is the one that passes your non-negotiable gates and produces the highest evidence-weighted score for a defined role at an acceptable cost per accepted outcome. There is no responsible universal winner across every role, risk level, data boundary, and buying constraint.

Q3How is an AI employee platform different from an agent builder?

An agent builder usually gives a technical or operations team components for designing agentic workflows. An AI employee platform usually packages more of the standing-role layer: identity, recurring responsibilities, work state, schedules, memory, supervision, and performance. Products can overlap. Compare what the buyer must build and operate.

Q4Should a small business use the same 12 criteria?

Yes, but change the weights and evidence burden. A low-risk internal research role may use a short configured proof and lightweight controls. A role that sends external messages, handles sensitive data, or changes business systems still needs permissions, approval, evidence, and recovery even in a small company.

Q5Are AI agent benchmarks useful when selecting a platform?

Yes, when the benchmark is current, public, relevant, and interpreted narrowly. A research benchmark can support research-quality evaluation. It cannot establish reliable triggers, permissions, tool use, handovers, security, or business outcomes. Add representative role tests.

Q6How many AI employee platforms should enter a pilot?

Usually two or three well-matched candidates are enough after platform-shape screening and documented demos. More vendors can dilute testing quality. The important constraint is that every candidate receives the same workload, failure cases, acceptance rubric, controls, and cost accounting.

Published 31 July 2026 All Choosing a platform →