Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentPricingBlogStoryContact

AI Employee Platform RFP Checklist: 50 Questions and Proof Requests

Napkin-style sketch of a questionnaire sheet with fifty numbered rows feeding into an evidence folder, with an amber stamp reading proof on the folder
Fig 0Accept a claim only at the level its evidence supports.

An AI employee platform RFP should ask more than whether a vendor has memory, tools, approvals, logs, and enterprise security. It should ask how each capability works, where its boundary sits, what the customer must configure, and which artifact proves the answer.

That last requirement changes the result.

“Yes, we support approvals” is an assertion. A current control description, admin configuration, denied-action demonstration, approval record, and contractual commitment show what the answer means.

Use the 50 questions below after you have defined one role and created a shortlist. Send every vendor the same role card, action boundary, data profile, workload, and response format. Require each response to identify the supported product tier, default behavior, configuration dependency, customer responsibility, evidence, known limitation, and contractual status.

This is a practical first-pass RFP, not a universal compliance questionnaire or legal opinion. A regulated, public-sector, critical-infrastructure, or high-impact deployment may need hundreds of additional controls. The Cloud Security Alliance’s current AI Controls Matrix v1.1, for example, contains 247 control objectives across 18 domains and includes a mapped AI vendor questionnaire.

On this page · 24 sectionsOpen
  1. What Is an AI Employee Platform RFP?
  2. What Should You Send Vendors Before the 50 Questions?
  3. How Should Vendors Format Their RFP Responses?
  4. Who Should Review the RFP?
  5. The 50 AI Employee Platform RFP Questions at a Glance
  6. Domain 1: Role and Operating Architecture - Questions 1-5
  7. Domain 2: Tools, Integrations, and Execution - Questions 6-10
  8. Domain 3: Identity, Permissions, and Approvals - Questions 11-15
  9. Domain 4: Context, Knowledge, and Memory - Questions 16-20
  10. Domain 5: Data Privacy and Lifecycle - Questions 21-25
  11. Domain 6: Security Assurance and Supply Chain - Questions 26-30
  12. Domain 7: Evaluation, Reliability, and Safety - Questions 31-35
  13. Domain 8: Observability and Human Governance - Questions 36-40
  14. Domain 9: Pricing, Implementation, and Support - Questions 41-45
  15. Domain 10: Contract, Continuity, and Exit - Questions 46-50
  16. Which Evidence Pack Should Every Vendor Return?
  17. Which Six Tasks Should the Vendor Demonstrate Live?
  18. How Should You Evaluate the 50 RFP Answers?
  19. Which RFP Answers Should Become Contract Terms?
  20. What Are the Most Common AI Vendor RFP Red Flags?
  21. Which 15 Questions Should a Small Team Ask First?
  22. How Does CellCog Map to This RFP?
  23. Copyable AI Employee Platform RFP Response Template
  24. The Short Version
Key points7 · 32 min full read
  1. Send the RFP only after one role, workload, data boundary, and action envelope are defined.
  2. Require Yes, Partial, No, or Not applicable - then require evidence, tier, configuration, owner, limitation, and contract status.
  3. Use all 50 questions for consequential or enterprise deployments; use the 15-question core for a low-risk self-serve role.
  4. Treat permissions, revocation, recovery, audit evidence, incident response, data use, deletion, and exit as possible knockout requirements.
  5. Make vendors demonstrate a normal task, denied action, duplicate trigger, poisoned context, tool failure, and complete evidence export.
  6. Do not treat a certification, public benchmark, polished demo, or roadmap statement as proof of the complete role.
  7. Score the resulting evidence with the 12-point platform framework, then move only qualified vendors into a bounded pilot.

§ 01What Is an AI Employee Platform RFP?

An AI employee platform request for proposal is a structured request for product, security, privacy, reliability, commercial, and contractual evidence from vendors being considered for a standing AI role.

It is narrower than a complete enterprise risk program and broader than a feature checklist.

Document Primary job Typical owner What it cannot prove alone
AI employee platform RFP Compare role fit, controls, evidence, cost, and terms Business owner + procurement Repeated performance in your environment
Security questionnaire Assess standard security controls and supplier risk Security/GRC Business outcome quality
Data processing addendum Bind privacy roles and processing terms Privacy/legal Product behavior
Architecture review Inspect data flows, trust boundaries, and integrations Security/IT/engineering Commercial fit
Live proof session Observe selected controls and failure paths Technical evaluator Stability across many cases
Pilot Measure representative role performance against a baseline Business + evaluation owner Unwritten contractual commitments
Table 1How the RFP relates to the other diligence documents

Use the RFP to decide what is documented, demonstrable, independently assured, contractually committed, still unverified, or unavailable. The bounded AI employee pilot should then test the remaining performance uncertainty against a baseline, representative cases, staged authority, and precommitted exit rules.

When a 50-question RFP is justified

Use the full checklist when the AI employee will:

  • read confidential or personal data;
  • access authenticated applications;
  • send external messages;
  • create, update, approve, delete, or publish records;
  • run on schedules or events without a person present;
  • retain memory across tasks;
  • delegate to other agents;
  • affect customers, money, rights, security, or regulated work;
  • become operationally important; or
  • require an enterprise agreement, security review, or data processing addendum.

A read-only research role using public information may not need the same process as a role that updates a CRM, emails customers, or runs commands on a workstation.

What this RFP should not become

Do not ask 50 generic questions before defining the job. Vendors will answer against their broadest product capability while your team evaluates a narrower deployment.

Do not use the RFP to:

  • select a role;
  • replace security or legal review;
  • reproduce an industry framework without scoping;
  • award points for irrelevant enterprise features;
  • demand sensitive vendor internals that can be verified another way;
  • force every answer into Yes;
  • turn roadmap promises into current capability; or
  • declare a pilot successful before it runs.

The RFP converts a defined role into proof requests.

§ 02What Should You Send Vendors Before the 50 Questions?

Every vendor should receive the same 1-2 page scope sheet. Without it, “supported” can mean different things in every response.

Role scope sheet

Field Buyer-provided example
Role Weekly acquisition analyst
Recurring outcome Accepted channel review by 10 a.m. Monday
Valid triggers Weekly schedule after approved data refresh
Systems Analytics, ad accounts, CRM, document store, task tracker
Data classes Internal business data; no payment credentials
Allowed actions Read, calculate, draft, save, update task
Prohibited actions Change budgets, edit source data, contact customers
Approval boundary Any external distribution or spend recommendation
Expected volume 4 standard runs plus 2 exception cases per month
Acceptance evidence Reconciled figures, linked sources, complete template
Escalation Missing source, conflicting attribution, abnormal spend
Supervisor Growth lead
Retention need 12 months of reports; shorter operational logs if policy allows
Deployment region Buyer-specified requirement
Exit requirement Export role config, tasks, evidence, artifacts, and usage
Table 2The buyer-provided scope sheet every vendor receives

The scope sheet should use real workflow objects: account, record, file, recipient, tool operation, task state, source, approver, and accepted outcome.

Action-risk tier

Tier Example RFP depth
0 Public-web research with no authenticated tools Condensed 15-question core
1 Internal read and draft with human publication 50 questions, lighter contract review
2 Reversible writes or external action after approval Full RFP and live control demonstrations
3 Consequential, irreversible, rights-affecting, regulated, or privileged action Full RFP plus specialist security, privacy, legal, and domain review
Table 3Four action-risk tiers and the RFP depth each demands

The tier is an internal scoping device, not an industry standard. Increase rigor when the action’s blast radius grows.

Buyer constraints

Tell vendors which answers are knockout requirements before they respond.

Examples:

  • no inference-data training;
  • named processing regions;
  • central access revocation;
  • enforceable approval before defined actions;
  • exportable action evidence;
  • incident notification within a contractual period;
  • deletion within an accepted period;
  • supported identity controls;
  • a minimum service commitment; or
  • an acceptable exit format.

A vendor can give an excellent answer and still be wrong for the deployment.

§ 03How Should Vendors Format Their RFP Responses?

Reject free-form sales essays. Require one structured row per question.

Response field Required entry
Status Yes, Partial, No, or Not applicable
Scope Product, plan, region, interface, and workload covered
Default/configured Default behavior or configuration required
Customer responsibility Control or operation the buyer must own
Evidence ID One or more evidence artifacts
Limitation Known boundary, exception, or unsupported condition
Roadmap Target date and dependency, kept separate from current status
Contract status Standard term, negotiable term, or non-contractual statement
Evidence date Version or last-reviewed date
Vendor owner Person accountable for follow-up
Table 4The required response format per question

Use evidence codes

Code Evidence artifact Best used to prove
DOC Current product, security, or policy documentation Defined behavior and scope
ARC Architecture or data-flow diagram Components, trust boundaries, and processors
CFG Redacted admin configuration or export Available control and granularity
LOG Redacted structured log or audit export What is recorded and reconstructable
DEMO Live vendor demonstration using buyer case Executable product behavior
TEST Buyer-run or jointly run test result Behavior under buyer conditions
ASSURE Independent audit, certification, or assessment report Assessed control scope and period
CONT Contract, DPA, SLA, security addendum, or order form Enforceable commitment
COST Quote, rate card, usage export, or invoice example Commercial mechanics
Table 5Nine evidence codes and what each best proves

No artifact proves everything. A SOC report may support control assurance but not your role’s denial behavior. A demo may prove behavior but not create a contractual service level.

Protect sensitive evidence

Allow vendors to:

  • redact customer identifiers;
  • provide audit reports under NDA;
  • use a secure data room;
  • give an auditor’s report instead of exploit details;
  • demonstrate a control without disclosing secrets; and
  • describe an architecture boundary without publishing attack-enabling configuration.

The objective is sufficient assurance, not forced disclosure.

§ 04Who Should Review the RFP?

No single reviewer owns all 50 answers.

Reviewer Primary domains Decision owned
Business/role owner Role, workflow, acceptance, escalation Does the product support useful work?
IT/integration owner Connectors, APIs, identity, environments Can it fit the operating stack?
Security/GRC Access, logs, assurance, incident, supply chain Is residual security risk acceptable?
Privacy/legal Data flow, retention, terms, IP, liability Are processing and contractual terms acceptable?
Procurement/finance Pricing, support, renewal, exit Is the commercial commitment controllable?
Evaluation owner Demo cases, evidence ledger, open gaps Has the vendor proved the scored claims?
Executive sponsor Risk acceptance and deployment boundary Should the organization proceed?
Table 6Seven reviewers and the decision each owns

CISA’s Secure by Demand Guide recommends considering product security before procurement, putting appropriate requirements into the contract during procurement, and continuing assessment after purchase. That lifecycle matters for AI employees because model, prompt, tool, memory, and provider changes can alter behavior after the contract is signed.

§ 05The 50 AI Employee Platform RFP Questions at a Glance

Domain Questions What the domain establishes
1. Role and operating architecture 1-5 Whether the product can represent standing work
2. Tools, integrations, and execution 6-10 Whether it can act in the required systems safely
3. Identity, permissions, and approvals 11-15 Whether authority can be constrained and revoked
4. Context, knowledge, and memory 16-20 Whether persistent context is governed
5. Data privacy and lifecycle 21-25 Where data goes, why, and for how long
6. Security assurance and supply chain 26-30 How the vendor builds, tests, and responds
7. Evaluation, reliability, and safety 31-35 Whether behavior is measured under relevant conditions
8. Observability and human governance 36-40 Whether decisions and actions can be supervised
9. Pricing, implementation, and support 41-45 What the role will cost and require operationally
10. Contract, continuity, and exit 46-50 What happens when terms, providers, or needs change
Table 7Ten domains, fifty questions

The 10 domains mirror the operating path from role definition through termination. Five Yes answers in one domain cannot compensate for a failed knockout requirement in another.

§ 06Domain 1: Role and Operating Architecture - Questions 1-5

These questions establish whether the product carries standing work or only exposes a capable model, prompt, or workflow surface.

# Exact RFP question Minimum evidence request Red flag
1 How does the platform represent a persistent role, including purpose, goals, supervisor, boundaries, recurring responsibilities, and retirement? DOC + CFG showing the role object and lifecycle Persona or prompt presented as complete role persistence
2 Which state persists across sessions, shifts, interfaces, model changes, and operator changes, and which state does not? ARC + restart DEMO + data dictionary “The conversation remembers” without explicit state objects
3 Which schedules, messages, events, queues, webhooks, or human requests can start work, and how are duplicate, stale, or replayed triggers handled? Trigger DOC + duplicate-event TEST “Runs 24/7” with no idempotency or cancellation behavior
4 Which task states, dependencies, deadlines, postconditions, reopen paths, and handover objects are supported? CFG + interrupted-task DEMO + sample handover Only running and done, or state hidden in narrative chat
5 Which models, orchestration layers, retrieval systems, memory services, tool brokers, and execution environments participate in the service? Versioned ARC with provider responsibilities “Proprietary AI” without component or responsibility boundaries
Table 8Questions 1-5 with minimum evidence and red flags

What a strong response shows

A strong response connects:

role → trigger → task state → context → plan → action → evidence → postcondition → handover

It also separates native product capability from professional services, custom code, third-party tools, and buyer-operated components.

Question 5 does not require source code. It requires enough architecture to understand the value chain, data path, control ownership, and failure dependencies.

Knockout candidates

Consider a knockout when:

  • the role cannot be paused or retired;
  • duplicate triggers can repeat consequential actions;
  • unfinished work cannot be distinguished from completed work; or
  • critical provider dependencies are undisclosed.

§ 07Domain 2: Tools, Integrations, and Execution - Questions 6-10

An AI employee creates operational risk when reasoning becomes action. Ask about exact operations, not connector logos.

# Exact RFP question Minimum evidence request Red flag
6 For every required system, which objects and operations are supported through native API, connector, browser/computer use, code, or custom integration? Operation-level matrix + DEMO of the buyer’s last mile “Connects to CRM” without object, read/write, or account scope
7 How are tool inputs, parameters, outputs, schemas, and postconditions validated before and after execution? DOC + invalid-parameter TEST + sample structured result Model-generated parameters sent directly to consequential tools
8 How are credentials stored, rotated, scoped, isolated, audited, and prevented from entering prompts, memory, output, or logs? ARC + CFG + relevant ASSURE evidence Shared static credentials or secrets visible in agent context
9 What happens on timeout, partial success, rate limit, unavailable dependency, invalid response, or ambiguous tool result? Failure-state DEMO + retry/idempotency DOC + LOG Unlimited retries, silent partial success, or duplicate write
10 How are custom integrations developed, tested, versioned, approved, sandboxed, monitored, and maintained after an API changes? Integration lifecycle DOC + test environment evidence Production-first testing or undocumented customer-owned code
Table 9Questions 6-10 with minimum evidence and red flags

What a strong response shows

The vendor should distinguish:

  • connection from authorization;
  • tool availability from supported operation;
  • model choice from execution policy;
  • retry from safe recovery;
  • browser reach from durable integration; and
  • artifact generation from confirmed delivery.

The question is not whether the platform can reach thousands of tools. It is whether the selected role can use five required operations within a bounded, observable action path.

Agent-specific security evidence

OWASP’s AI Agent Security Cheat Sheet identifies tool abuse, privilege escalation, data exfiltration, memory poisoning, excessive autonomy, cascading failures, and unbounded cost as agent-specific risks. Its recommendations include least-privilege tool scopes, high-impact approval, structured monitoring, cost and retry limits, and adversarial testing.

Ask the vendor to map its evidence to the threats that apply to your role.

§ 08Domain 3: Identity, Permissions, and Approvals - Questions 11-15

The full AI employee permissions and approvals framework explains how to build an action envelope. These five RFP questions ask whether a platform can enforce it.

# Exact RFP question Minimum evidence request Red flag
11 Can access be restricted by tool, operation, resource, account, data class, recipient, value, time, volume, environment, and role? CFG of buyer-specified envelope + denial TEST Permission exists only at connector or prompt level
12 Which identity controls are available for human users, service identities, AI roles, administrators, and support personnel? Identity/control matrix + ASSURE scope Shared administrator accounts or unclear service identity
13 Can approval bind to the exact actor, action, tool, target, parameters, amount, expiry, and one-time execution? Approval CFG + modified-parameter and replay TEST + LOG Approval to a vague plan rather than the executed operation
14 How quickly can an administrator revoke a connector, credential, permission, trigger, role, or entire workspace, including work already in progress? Live revocation DEMO + documented propagation behavior Revocation applies only after a run completes or support intervenes
15 Which administrative controls support separation of duties, least privilege, multifactor authentication, role-based access, single sign-on, provisioning, and periodic access review? Tier-specific DOC + CFG + contract/roadmap separation “Enterprise controls” without named availability or product tier
Table 10Questions 11-15 with minimum evidence and red flags

What a strong response shows

The response should identify:

  • which controls the platform enforces;
  • which controls come from the connected application;
  • which controls require middleware;
  • which controls are only instructions;
  • which plan includes the control; and
  • how a denied, expired, revoked, or altered action appears in evidence.

Do not mark a vendor down merely because a buyer-desired control sits in the connected system. Mark it down when the combined path cannot enforce or prove the boundary.

Knockout candidates

For a consequential role, common knockout failures include:

  • no central stop;
  • no enforceable approval;
  • approval parameters can change after review;
  • no operation-level restriction;
  • no administrator attribution; or
  • no workable identity model for the deployment.

§ 09Domain 4: Context, Knowledge, and Memory - Questions 16-20

Persistent memory improves continuity and expands the attack and privacy surface. The AI employee memory guide provides the lifecycle design; the RFP should require proof of the platform controls.

# Exact RFP question Minimum evidence request Red flag
16 Which context types are stored - role instructions, authoritative knowledge, working context, preferences, task state, summaries, and episodic records - and where? Context data dictionary + ARC One undefined memory store for unrelated objects
17 How are source provenance, freshness, authority, instruction precedence, conflict, and uncertainty represented during retrieval and decision-making? Conflicting-source TEST + source evidence Newest or most similar content silently overrides policy
18 How is context isolated by tenant, workspace, role, user, account, customer, project, task, and environment? Isolation ARC + negative-access TEST + ASSURE Shared context without documented boundaries
19 Who can inspect, correct, suppress, expire, export, and delete each memory type, and how do changes propagate to derived summaries or indexes? Correction/deletion DEMO + lifecycle DOC Memory is persistent but cannot be located or corrected
20 How does the platform detect and contain prompt injection, memory poisoning, malicious retrieval, poisoned tool output, and sensitive-data persistence? Threat model + adversarial TEST + incident/eval evidence “The model is trained to ignore attacks” as the only control
Table 11Questions 16-20 with minimum evidence and red flags

What a strong response shows

Strong memory answers are object-specific. They distinguish among:

  • chat history;
  • uploaded documents;
  • retrieved knowledge;
  • task state;
  • generated summaries;
  • learned preferences; and
  • action evidence.

The vendor should explain which layer is authoritative, which is advisory, and what happens when two layers conflict.

Ask for propagation evidence

Delete or correct a source record, then test:

  1. direct retrieval;
  2. generated summary;
  3. open task;
  4. long-term memory;
  5. audit record; and
  6. export.

Deletion of a visible document does not prove removal from derived context, backups, or required audit evidence.

§ 10Domain 5: Data Privacy and Lifecycle - Questions 21-25

Data can flow through the platform, model providers, connector brokers, email services, cloud infrastructure, support tools, analytics, logs, memory, and external actions.

# Exact RFP question Minimum evidence request Red flag
21 Provide a data inventory and flow diagram showing data categories, purpose, component, processor, region, transfer, storage, and recipient. ARC + subprocessor list + processing register Privacy policy without service-specific data flow
22 Is customer input, output, feedback, memory, tool data, or metadata used to train or improve any vendor or third-party model, and what opt-out or contract applies? DOC + DPA/CONT + provider terms “We do not train” without covering subprocessors or metadata
23 How is data encrypted in transit and at rest, how are keys and secrets managed, and which data may remain unencrypted for operation or support? Security DOC + relevant ASSURE scope Blanket “encrypted” statement without exceptions or boundaries
24 What retention schedule applies to prompts, outputs, memory, mail, files, logs, telemetry, backups, deleted objects, and support records? Object-level schedule + deletion DEMO + CONT One retention statement for every data type
25 How can the customer access, correct, export, restrict, and delete data during service and after termination, including derived and subprocessor-held data? Sample export + deletion procedure + DPA terms Manual support promise with no scope, format, or timeline
Table 12Questions 21-25 with minimum evidence and red flags

What a strong response shows

The vendor distinguishes:

  • controller and processor roles where applicable;
  • customer content from service telemetry;
  • active storage from backup retention;
  • account deletion from object deletion;
  • inference processing from training use;
  • primary vendor from subprocessor; and
  • legal retention from product convenience.

Do not accept “zero retention” without identifying the provider, data type, endpoint, plan, and exception.

Match the RFP to applicable law and policy

Privacy requirements depend on data, people, location, sector, and use. Ask qualified privacy and legal reviewers to tailor this domain.

The RFP should record the organization’s requirement. It should not ask the vendor to decide which law applies to the buyer.

§ 11Domain 6: Security Assurance and Supply Chain - Questions 26-30

AI employee platforms combine normal SaaS risk with models, prompts, retrieval, memory, tools, code, browsers, and third-party services.

# Exact RFP question Minimum evidence request Red flag
26 Which independent audits, certifications, penetration tests, or assessments cover the exact service, region, period, and controls proposed? ASSURE evidence under NDA if needed Certification name with no scope, entity, period, or exceptions
27 Describe secure development, threat modeling, code review, dependency management, vulnerability disclosure, patching, secrets scanning, and release controls. SDLC DOC + vulnerability policy + assurance excerpt “Industry standard SDLC” with no practices or ownership
28 How are tenants isolated across application, storage, retrieval, memory, logs, execution environments, browser/computer sessions, caches, and support access? Isolation ARC + test/assurance evidence Application-level tenant ID treated as the only isolation control
29 How are model providers, open-source components, connectors, plugins, tool servers, datasets, and other subprocessors inventoried, assessed, monitored, and changed? Supplier inventory + change policy + CONT notification Material provider can change with no notice or review
30 Provide the incident-response process for AI misuse, data exposure, unauthorized action, cross-tenant access, prompt injection, model/provider event, and service compromise. Incident plan summary + notification terms + exercise evidence Cyber incident process excludes AI action or model events
Table 13Questions 26-30 with minimum evidence and red flags

What a strong response shows

A certificate is useful when the scope matches the purchased service. It is not a universal pass.

Ask:

  • Which legal entity was assessed?
  • Which service and infrastructure were included?
  • Which period was covered?
  • Were important controls customer-operated?
  • Were exceptions found?
  • Is the report current?
  • Does the evidence cover AI-specific components?

Use broader frameworks when the risk requires them

The 50 questions are intentionally shorter than a complete control catalog. CSA’s July 2026 AICM v1.1 overview says its AI-CAIQ contains 320 questions aligned to 247 controls. Enterprise security teams can use that resource for deeper self- and third-party assessment rather than expanding an improvised spreadsheet indefinitely.

§ 12Domain 7: Evaluation, Reliability, and Safety - Questions 31-35

The vendor must show how it knows the system works - not only that the model can produce a good answer.

# Exact RFP question Minimum evidence request Red flag
31 Which evaluation sets test model output, tool choice, action correctness, memory, escalation, refusal, recovery, and end-to-end role outcomes? Eval inventory + sample cases + versioned results One benchmark used as proof for the whole platform
32 Which metrics track accepted outcomes, first-pass quality, correction time, false completion, escalation, unauthorized action, severity, latency, and cost? Metric definitions + sample report + buyer mapping Tokens, runs, or user ratings presented as role success
33 Which model, prompt, policy, tool, retrieval, memory, or provider changes trigger regression tests, approval, customer notice, or rollback? Change policy + release evidence + rollback DEMO Silent updates with no version or regression evidence
34 Which adversarial tests cover direct and indirect prompt injection, tool abuse, privilege escalation, data exfiltration, memory poisoning, approval manipulation, and runaway loops? Threat-linked test matrix + redacted results Generic red team with no agent action path
35 What limits and recovery mechanisms govern run duration, retries, tokens, cost, concurrency, delegation depth, tool chaining, timeout, fallback, and safe failure? Limit CFG + failure TEST + service DOC Unlimited autonomy or retry described as a benefit
Table 14Questions 31-35 with minimum evidence and red flags

What a strong response shows

The evaluation set should resemble deployment:

  • same role;
  • same tool classes;
  • same data types;
  • same action risk;
  • realistic failures;
  • explicit acceptance;
  • severity-aware errors;
  • known limitations; and
  • versioned results.

NIST’s AI RMF Core calls for documented test sets, evaluation under conditions similar to deployment, production monitoring, and response and recovery processes. Ask the vendor which artifacts support each lifecycle stage.

Do not confuse benchmark and role proof

A benchmark can support the capability it measures. It does not automatically prove:

  • connector correctness;
  • permission enforcement;
  • memory isolation;
  • reliable triggers;
  • task handover;
  • incident containment; or
  • cost per accepted outcome.

Record the benchmark’s task set, date, configuration, judge, and scope - then keep the role evaluation separate.

§ 13Domain 8: Observability and Human Governance - Questions 36-40

The buyer should be able to reconstruct work without asking the same model to narrate what it remembers.

# Exact RFP question Minimum evidence request Red flag
36 Which structured fields connect trigger, run, role, task, source, instruction, model, decision, approval, tool action, result, cost, and final state? Redacted LOG + schema + correlation DEMO Final chat transcript is the only evidence
37 Can authorized users search, filter, retain, redact, export, and restrict logs by role, customer, account, action, severity, and time? Admin CFG + sample export + retention controls Export requires bespoke engineering or excludes action detail
38 Which monitors and alerts detect failed runs, repeated retries, abnormal cost, unauthorized requests, stale work, performance regression, data exposure, and unusual tool use? Monitoring catalog + alert TEST + routing evidence Dashboard exists but nobody or nothing receives alerts
39 How do humans review, approve, interrupt, override, correct, roll back, reopen, escalate, and assume ownership of work? End-to-end DEMO + policy DOC + LOG Human-in-the-loop means only reading the final answer
40 Which administrative changes are recorded, who can investigate them, and how are incident evidence, legal holds, and audit retention handled? Admin-history LOG + access matrix + retention terms Prompt, policy, or permission changes are unattributed
Table 15Questions 36-40 with minimum evidence and red flags

What a strong response shows

Strong evidence separates:

  • user instruction;
  • system or role instruction;
  • retrieved content;
  • model decision;
  • policy decision;
  • human approval;
  • execution request;
  • external-system result; and
  • final task state.

The buyer needs structured evidence that shows inputs, control decisions, actions, results, and responsibility without depending on a model-generated explanation of its own behavior.

Test the investigation path

Give the evaluator one suspicious external message and ask them to find:

  1. which event started the task;
  2. which role handled it;
  3. which source influenced the decision;
  4. which instruction and model version applied;
  5. which approval was requested;
  6. which tool call executed;
  7. what the external system returned;
  8. what the platform retried;
  9. what the human changed; and
  10. how the task ended.

If the evidence exists but cannot be found during an incident, observability is incomplete.

§ 14Domain 9: Pricing, Implementation, and Support - Questions 41-45

The AI employee cost framework separates vendor price from setup, integration, review, correction, monitoring, and failure exposure. The RFP should make every vendor quote the same work.

When “build internally” remains on the shortlist, require the internal sponsor to answer the same ownership questions through a three-year build-versus-buy AI employee analysis. Framework access and prototype labor are not substitutes for a costed evaluation, security, memory, monitoring, incident, change, and exit plan.

# Exact RFP question Minimum evidence request Red flag
41 Which units are billed for subscription, seats, credits, tokens, runs, time, tools, storage, models, background work, retries, failures, support, and overage? COST rate card + billing definitions + invoice example “Simple usage pricing” without failed-run or premium-model rules
42 Price the buyer’s fixed workload, including ordinary cases, hard cases, failures, review, expected correction, and forecast range. Written same-workload quote + assumptions + sensitivity Vendor quotes unrelated plan minimum rather than role consumption
43 Which implementation tasks, integrations, data preparation, configuration, evaluation, training, change management, and services are required, by whom, and when? Implementation plan + RACI + dependencies + fees “Deploy in minutes” excludes production configuration
44 Which support channels, hours, severity levels, response targets, resolution targets, escalation paths, service levels, status pages, and remedies apply? Support policy + SLA/CONT + sample escalation Dedicated support is marketing language, not a named service
45 Which required capabilities depend on roadmap, beta, custom work, partner products, buyer code, or a higher tier, and which will be committed in the order form? Gap list + commercial proposal + CONT status Roadmap capability scored as currently available
Table 16Questions 41-45 with minimum evidence and red flags

What a strong response shows

The quote should normalize:

  • task mix;
  • context volume;
  • tool calls;
  • model tier;
  • artifact type;
  • retries;
  • failures;
  • concurrency;
  • review;
  • correction;
  • support;
  • implementation; and
  • term.

The lowest subscription may create the highest operating cost. The RFP must preserve both.

Keep commercials current

Require a quote-valid-until date, renewal rules, overage treatment, usage-export format, price-change notice, termination right, and assumptions.

Do not freeze a pricing page into a multi-year decision without a contract.

§ 15Domain 10: Contract, Continuity, and Exit - Questions 46-50

The last five questions determine whether the evidence survives a dispute, outage, provider change, acquisition, or termination.

# Exact RFP question Minimum evidence request Red flag
46 Who owns or may use customer inputs, outputs, configurations, evaluations, feedback, memory, fine-tuned assets, logs, and generated artifacts? Terms + DPA + IP clauses + usage-right matrix Broad secondary use hidden outside the response
47 How do the agreement and DPA allocate responsibility, liability, indemnity, warranty, acceptable use, customer supervision, unauthorized action, data event, and third-party failure? Standard and proposed CONT language Product promise conflicts with liability or user-responsibility term
48 What continuity and fallback exist for service, model, cloud, connector, region, key personnel, or subprocessor outage or termination? Business continuity summary + dependency/fallback ARC + SLA Critical role depends on one undisclosed service with no fallback
49 In which documented formats can the buyer export role configuration, prompts/instructions, tasks, memory, artifacts, logs, evaluations, usage, and evidence during service and at exit? Complete sample export + schema + transition support terms PDF summary is the only exit artifact
50 What is the termination and deletion process, including access cut-off, credential revocation, export window, active deletion, backup expiry, subprocessor deletion, certification, and exceptions? Exit runbook + deletion CONT + sample certificate or record “Data deleted on request” with no object scope or timeline
Table 17Questions 46-50 with minimum evidence and red flags

What a strong response shows

The contract should not say “autonomous work” in the product description and leave every consequence undefined in the terms.

Review the following materials for contradictions:

  • product documentation;
  • security response;
  • privacy policy;
  • DPA;
  • terms;
  • SLA;
  • support policy;
  • order form;
  • statement of work; and
  • insurance or assurance evidence.

NIST’s third-party guidance

The NIST Generative AI Profile recommends use-case-based supplier assessment, third-party inventory, ongoing monitoring, incident planning, fallbacks, and contract terms addressing ownership, rights, security, serious-incident notification, response times, and critical support.

That is the procurement lesson: a current answer must become an owned control, continuing monitor, or enforceable term where the risk requires it.

§ 16Which Evidence Pack Should Every Vendor Return?

The questionnaire is incomplete without a deliverable list.

Core evidence pack

Request:

  1. product and architecture overview;
  2. role, task-state, trigger, and handover documentation;
  3. operation-level integration matrix;
  4. permission and approval configuration evidence;
  5. identity and administrator-control matrix;
  6. data inventory and flow diagram;
  7. current subprocessor list;
  8. object-level retention and deletion schedule;
  9. sample data and evidence export;
  10. redacted structured run and action logs;
  11. evaluation methodology and sample result;
  12. change and regression policy;
  13. incident-response summary and notification term;
  14. security assurance package;
  15. same-workload commercial quote;
  16. implementation plan and RACI;
  17. support policy and proposed SLA;
  18. standard terms and DPA;
  19. business continuity and exit summary; and
  20. current gaps, betas, custom dependencies, and roadmap items.

The request should specify an acceptable secure delivery method.

Evidence freshness rules

Artifact Freshness question
Documentation Which product version and date does it cover?
Architecture Has the provider, region, or data path changed?
Configuration Is it available in the proposed tier today?
Log/export Was it generated from the proposed product surface?
Test result Which model, prompt, tools, policies, and date were used?
Assurance report Which entity, service, period, and exceptions apply?
Contract Is this the operative standard or negotiated language?
Quote How long is it valid and which assumptions can change?
Table 18Freshness questions per artifact type

An undated screenshot is not durable evidence.

Evidence ownership

For every accepted artifact, record:

  • buyer owner;
  • source vendor;
  • secure location;
  • reviewed date;
  • approved use;
  • expiration or refresh trigger;
  • linked RFP question;
  • linked risk or requirement; and
  • final disposition.

Procurement evidence becomes operational evidence only when someone maintains it.

§ 17Which Six Tasks Should the Vendor Demonstrate Live?

The RFP asks what the product can do. The proof session checks a small number of claims before a full pilot.

Demo 1: Normal role path

The role should:

  1. wake from the approved trigger;
  2. load the correct context;
  3. complete the requested work;
  4. use the permitted tool;
  5. satisfy the postcondition;
  6. attach evidence; and
  7. enter the correct final state.

Demo 2: Denied action

Ask the role to perform an operation outside its action envelope.

Require:

  • technical denial;
  • clear reason;
  • no external side effect;
  • correct task state;
  • escalation if appropriate; and
  • structured log evidence.

Demo 3: Duplicate or stale trigger

Send the same event twice, then replay it after the task has expired.

Require:

  • one valid task;
  • no duplicate external action;
  • visible deduplication or rejection;
  • bounded retry; and
  • correlation to the original event.

Demo 4: Conflicting or poisoned context

Place a malicious or stale instruction inside a retrieved document and provide a newer authoritative policy.

Require:

  • source hierarchy;
  • conflict detection;
  • no unauthorized persistence;
  • safe refusal or escalation; and
  • evidence of which source controlled the decision.

Demo 5: Tool failure after partial progress

Make the destination unavailable after one step succeeds.

Require:

  • correct partial-state record;
  • no repeated side effect;
  • bounded retry;
  • recovery or compensation path;
  • human escalation; and
  • final reconciliation.

Demo 6: Investigation and export

Give a reviewer the task identifier and ask them to reconstruct:

  • trigger;
  • role;
  • context;
  • instruction;
  • model;
  • decision;
  • approval;
  • action;
  • result;
  • retry;
  • cost; and
  • final state.

Then export the relevant record.

These six tasks validate shortlist claims. They do not replace a representative pilot with a baseline, sample design, duration, KPI thresholds, and go/no-go criteria.

§ 18How Should You Evaluate the 50 RFP Answers?

Use the evidence scale and weights already defined in the AI employee platform selection guide. Keep this RFP focused on producing the evidence that guide needs.

Classify every answer

Classification Meaning Procurement action
Verified Evidence directly supports the scoped answer Score at the demonstrated evidence level
Conditional Works only with a plan, configuration, partner, service, or buyer control Price and assign the dependency
Contract pending Behavior is shown but commitment is unresolved Do not treat as enforceable
Unverified Assertion or incomplete artifact Request proof or lower the evidence score
Gap Required capability is unavailable Apply knockout or risk treatment
Not applicable Requirement does not apply to the scoped role Record the rationale
Table 19Six answer classifications and the procurement action for each

Keep gates outside the average

Potential gates include:

  • high-risk approval;
  • central revocation;
  • safe failure and interruption;
  • action evidence;
  • acceptable training/data-use terms;
  • required identity control;
  • acceptable incident term;
  • retention and deletion fit; and
  • workable exit.

One failed gate can stop the procurement even if the vendor has 49 strong answers.

Price every “Partial”

Partial can mean:

  • higher subscription tier;
  • professional services;
  • custom integration;
  • buyer-built middleware;
  • connected-system control;
  • manual review;
  • additional vendor;
  • policy process;
  • delayed roadmap; or
  • accepted risk.

Translate it into owner, cost, delivery date, operational burden, and residual risk.

§ 19Which RFP Answers Should Become Contract Terms?

Contractualization should follow impact.

Claim When to bind it Possible document
Data use and training Whenever customer data terms matter DPA, security addendum, agreement
Processing region When residency or transfer is material DPA/order form
Retention and deletion When policy or law requires a period DPA/security addendum
Incident notification When delayed notice changes response DPA/security addendum
Availability and support When the role becomes operationally important SLA/order form
Model or subprocessor change notice When change alters risk or compliance DPA/service terms
Required controls or product tier When procurement depends on them Order form/SOW
Pricing and overage Always for material spend Order form
Export and transition When switching cost matters Agreement/SOW
Remediation or termination right When a control failure is unacceptable Agreement/addendum
Table 20Which claims to bind, when, and where

Do not attempt to put every marketing sentence into a contract. Bind the facts whose failure would change the purchase or create unacceptable exposure.

Resolve contradictions

If a sales answer, security questionnaire, privacy policy, terms, and order form disagree:

  1. identify the conflict;
  2. ask the vendor to resolve it in writing;
  3. identify which document controls;
  4. update the RFP disposition; and
  5. do not rely on the most favorable non-binding statement.

The procurement record should show what the buyer actually accepted.

§ 20What Are the Most Common AI Vendor RFP Red Flags?

Red flag answer What it usually hides Required follow-up
“Yes” with no artifact Claim scope is unknown Evidence ID, tier, date, limitation
“Enterprise-grade” No named control Exact control and enforcement
“Industry standard” No standard, scope, or assessor Named framework and evidence
“We use trusted models” Application controls are omitted Full value-chain responsibility
“We do not train on your data” Metadata or subprocessors are excluded Data-type and provider-specific term
“Human in the loop” Human sees only final output Exact approval and interruption path
“Full audit trail” Narrative transcript only Structured fields and export
“Encrypted” Key, state, or exception unclear Transit/at-rest scope and key ownership
“Delete on request” Backup or derivative data omitted Object schedule and certification
“Unlimited autonomy” Limits and stop conditions absent Retry, cost, action, and time caps
“Roadmap” in current-status column Future capability inflates score Separate No/Partial from roadmap
“Customizable” Services or buyer engineering required Scope, owner, price, and maintenance
“Compliant with…” Buyer treats compliance as inherited Scope, role, evidence, and shared responsibility
“Proprietary security” Control cannot be assessed Alternative assurance under NDA
“Best-in-class model” Role controls remain unproved Role-specific test and acceptance evidence
Table 21Fifteen red-flag answers and the required follow-up

A vague answer is not always evidence of a bad product. It is evidence of incomplete diligence.

§ 21Which 15 Questions Should a Small Team Ask First?

For a low-risk role, begin with these question IDs:

  1. Q1: How is the persistent role represented?
  2. Q3: Which triggers start work, and how are duplicates handled?
  3. Q6: Which exact tool operations are supported?
  4. Q9: What happens on timeout or partial tool failure?
  5. Q11: How granular are permissions?
  6. Q13: Does approval bind to the exact action?
  7. Q14: How is access or the role revoked?
  8. Q16: Which context and memory types are stored?
  9. Q18: How is context isolated?
  10. Q21: Where does data flow?
  11. Q22: Is any data used for training or improvement?
  12. Q24: What are retention and deletion periods?
  13. Q30: How are AI-specific incidents handled?
  14. Q36: Can the action path be reconstructed and exported?
  15. Q41: Which work consumes billable units?

Add all 50 questions before the role gains external action, sensitive data, broad access, operational criticality, or enterprise commitment.

The small-team version should reduce administrative weight, not remove the controls that contain the role.

§ 22How Does CellCog Map to This RFP?

CellCog should receive the same 50 questions as every shortlisted vendor.

Its current public pages answer part of the RFP and identify the gaps a buyer should take to sales, security, privacy, and legal review. The observations below reflect public materials reviewed in July 2026.

Evidence available publicly

RFP area Current public CellCog material What it supports
Role architecture AI Employees product and support material Goals, permissions, inbox, schedule, memory, task board, shifts, wake conditions, KPIs, approvals, and handovers are described
Tools and action Product and connector material Connected apps, computer/file work, browser work, email, and external action are described
Team architecture AI Employees and AI organization material Delegation, handoffs, and AI employees managing AI employees are described
Data flow CellCog privacy policy Data categories, providers, US storage, encryption, retention, and rights are described
Responsibility and service terms CellCog terms of service User responsibility, autonomous action, connected services, credits, restrictions, and service terms are described
Commercial model Pricing and role pages Credit-based self-serve pricing and usage estimates are described
Model capability Benchmark and product pages Research benchmark plus multi-format product capabilities are described
Table 22Public CellCog material mapped to RFP areas

Public evidence helps a buyer prepare. It does not answer every operation-, tenant-, plan-, assurance-, SLA-, incident-, export-, or contract-specific question.

Questions to send CellCog directly

Ask CellCog to identify and demonstrate:

  • operation- and resource-level permission granularity;
  • approval binding, expiry, replay protection, and revocation behavior;
  • structured action-log fields and export;
  • memory provenance, correction, isolation, and deletion propagation;
  • identity and administrator controls available by plan;
  • current independent security assurance and scope;
  • incident notification and support commitments;
  • service levels for the intended deployment;
  • fixed-workload credit consumption;
  • complete role, memory, task, evaluation, and evidence export; and
  • deletion and transition behavior at termination.

These are evidence requests, not claims that the capability is absent. Public materials do not establish the complete answer.

What CellCog should be asked to prove live

CellCog AI Employees publicly describes scheduled and event-triggered shifts, persistent memory, task boards, approvals, tools, KPI dashboards, and handovers. Use one scoped role to demonstrate:

  1. wake;
  2. context retrieval;
  3. tool action;
  4. denied action;
  5. exact approval;
  6. connector failure;
  7. handover;
  8. action evidence;
  9. usage record; and
  10. export.

Then send the same script to every shortlisted general-purpose platform.

Move qualified vendors from comparison to evidence review

First, compare candidate shapes and current vendor claims on the AI employee platform comparison. Then send the 50-question response template to each shortlisted vendor. If CellCog remains qualified, use its contact path to request the plan-specific evidence and commercial terms that public pages cannot establish.

§ 23Copyable AI Employee Platform RFP Response Template

Procurement header

  • Buyer:
  • Role:
  • Business outcome:
  • Action-risk tier:
  • Systems:
  • Data classes:
  • Regions:
  • Volume:
  • Required identity controls:
  • Required service level:
  • Retention requirement:
  • Contract term:
  • Knockout requirements:
  • Response due:

Vendor header

  • Vendor:
  • Product:
  • Proposed plan:
  • Proposed region:
  • Primary respondent:
  • Security respondent:
  • Privacy/legal respondent:
  • Commercial respondent:
  • Response date:
  • Evidence data-room location:

Question response row

Field Vendor response
Question ID
Status: Yes / Partial / No / N/A
Scoped answer
Supported product/plan/region
Default or configuration required
Customer responsibility
Evidence IDs
Known limitation
Roadmap item and date
Contract status
Vendor owner
Buyer disposition
Table 23The per-question response row

Evidence register

Evidence ID Type Title/version Date Questions supported Restrictions Buyer owner Refresh trigger
Scroll to compare all columns
Table 24The evidence-register columns

Gap register

Question Gap Knockout? Compensating control Owner Cost Due date Contract treatment Final decision
Scroll to compare all columns
Table 25The gap-register columns

§ 24The Short Version

An AI employee platform RFP works when every answer has:

  1. a scoped role;
  2. a clear status;
  3. a named product tier;
  4. an enforcement or operating explanation;
  5. customer responsibility;
  6. current evidence;
  7. a known limitation;
  8. a roadmap field separate from present capability;
  9. contractual status; and
  10. an accountable owner.

The 50 questions cover the entire chain:

role → trigger → context → tool → permission → action → evidence → outcome → incident → exit

Send the same chain to every vendor. Accept a claim only at the level its evidence supports. Put material promises into the appropriate agreement. Then let the pilot answer the remaining question: can this platform perform the role reliably enough to buy?

Frequently asked6 questions

Q1How many questions should an AI platform RFP contain?

Use enough questions to cover the scoped role’s architecture, actions, data, controls, evidence, cost, and exit. This template uses 50. A low-risk role can begin with 15. A high-risk enterprise assessment may require the CSA AI-CAIQ’s broader control set plus sector, legal, and organizational requirements.

Q2Is an AI RFP different from a SaaS security questionnaire?

Yes. A SaaS questionnaire covers important cloud, identity, development, incident, and privacy controls. An AI employee RFP adds role persistence, triggers, tools, action authorization, model and prompt change, memory, prompt injection, evaluation, human override, cost limits, and end-to-end action evidence.

Q3Should vendors answer every question with yes or no?

No. Require Yes, Partial, No, or Not applicable, followed by scope and proof. A precise No can be safer than a vague Yes. Keep roadmap capability separate from current status and identify configuration or customer-owned controls.

Q4Does SOC 2 or ISO certification replace the RFP?

No. Independent assurance may support specific controls over a defined service and period. It does not prove your role’s trigger, tool operation, denied action, approval binding, memory behavior, accepted outcome, price, or exit format. Review scope and then test the role-specific path.

Q5Should a vendor disclose prompts, source code, or security secrets?

Not necessarily. Ask for sufficient assurance through redacted architecture, control documentation, configuration, logs, demonstrations, independent reports, and contract terms. Use an NDA or secure data room when appropriate. Do not create security risk by demanding unnecessary sensitive details.

Q6What happens after vendors return the RFP?

Validate artifacts, classify each answer, apply knockout gates, score the evidence against the platform-selection framework, resolve contract-critical gaps, and invite only qualified vendors into a bounded proof session or pilot. Preserve the response, evidence date, owner, limitation, and refresh trigger.

Published 31 July 2026 All Choosing a platform →