An AI employee platform RFP should ask more than whether a vendor has memory, tools, approvals, logs, and enterprise security. It should ask how each capability works, where its boundary sits, what the customer must configure, and which artifact proves the answer.
That last requirement changes the result.
“Yes, we support approvals” is an assertion. A current control description, admin configuration, denied-action demonstration, approval record, and contractual commitment show what the answer means.
Use the 50 questions below after you have defined one role and created a shortlist. Send every vendor the same role card, action boundary, data profile, workload, and response format. Require each response to identify the supported product tier, default behavior, configuration dependency, customer responsibility, evidence, known limitation, and contractual status.
This is a practical first-pass RFP, not a universal compliance questionnaire or legal opinion. A regulated, public-sector, critical-infrastructure, or high-impact deployment may need hundreds of additional controls. The Cloud Security Alliance’s current AI Controls Matrix v1.1, for example, contains 247 control objectives across 18 domains and includes a mapped AI vendor questionnaire.
On this page · 24 sectionsOpen
- What Is an AI Employee Platform RFP?
- What Should You Send Vendors Before the 50 Questions?
- How Should Vendors Format Their RFP Responses?
- Who Should Review the RFP?
- The 50 AI Employee Platform RFP Questions at a Glance
- Domain 1: Role and Operating Architecture - Questions 1-5
- Domain 2: Tools, Integrations, and Execution - Questions 6-10
- Domain 3: Identity, Permissions, and Approvals - Questions 11-15
- Domain 4: Context, Knowledge, and Memory - Questions 16-20
- Domain 5: Data Privacy and Lifecycle - Questions 21-25
- Domain 6: Security Assurance and Supply Chain - Questions 26-30
- Domain 7: Evaluation, Reliability, and Safety - Questions 31-35
- Domain 8: Observability and Human Governance - Questions 36-40
- Domain 9: Pricing, Implementation, and Support - Questions 41-45
- Domain 10: Contract, Continuity, and Exit - Questions 46-50
- Which Evidence Pack Should Every Vendor Return?
- Which Six Tasks Should the Vendor Demonstrate Live?
- How Should You Evaluate the 50 RFP Answers?
- Which RFP Answers Should Become Contract Terms?
- What Are the Most Common AI Vendor RFP Red Flags?
- Which 15 Questions Should a Small Team Ask First?
- How Does CellCog Map to This RFP?
- Copyable AI Employee Platform RFP Response Template
- The Short Version
- Send the RFP only after one role, workload, data boundary, and action envelope are defined.
- Require Yes, Partial, No, or Not applicable - then require evidence, tier, configuration, owner, limitation, and contract status.
- Use all 50 questions for consequential or enterprise deployments; use the 15-question core for a low-risk self-serve role.
- Treat permissions, revocation, recovery, audit evidence, incident response, data use, deletion, and exit as possible knockout requirements.
- Make vendors demonstrate a normal task, denied action, duplicate trigger, poisoned context, tool failure, and complete evidence export.
- Do not treat a certification, public benchmark, polished demo, or roadmap statement as proof of the complete role.
- Score the resulting evidence with the 12-point platform framework, then move only qualified vendors into a bounded pilot.
§ 01What Is an AI Employee Platform RFP?
An AI employee platform request for proposal is a structured request for product, security, privacy, reliability, commercial, and contractual evidence from vendors being considered for a standing AI role.
It is narrower than a complete enterprise risk program and broader than a feature checklist.
| Document | Primary job | Typical owner | What it cannot prove alone |
|---|---|---|---|
| AI employee platform RFP | Compare role fit, controls, evidence, cost, and terms | Business owner + procurement | Repeated performance in your environment |
| Security questionnaire | Assess standard security controls and supplier risk | Security/GRC | Business outcome quality |
| Data processing addendum | Bind privacy roles and processing terms | Privacy/legal | Product behavior |
| Architecture review | Inspect data flows, trust boundaries, and integrations | Security/IT/engineering | Commercial fit |
| Live proof session | Observe selected controls and failure paths | Technical evaluator | Stability across many cases |
| Pilot | Measure representative role performance against a baseline | Business + evaluation owner | Unwritten contractual commitments |
Use the RFP to decide what is documented, demonstrable, independently assured, contractually committed, still unverified, or unavailable. The bounded AI employee pilot should then test the remaining performance uncertainty against a baseline, representative cases, staged authority, and precommitted exit rules.
When a 50-question RFP is justified
Use the full checklist when the AI employee will:
- read confidential or personal data;
- access authenticated applications;
- send external messages;
- create, update, approve, delete, or publish records;
- run on schedules or events without a person present;
- retain memory across tasks;
- delegate to other agents;
- affect customers, money, rights, security, or regulated work;
- become operationally important; or
- require an enterprise agreement, security review, or data processing addendum.
A read-only research role using public information may not need the same process as a role that updates a CRM, emails customers, or runs commands on a workstation.
What this RFP should not become
Do not ask 50 generic questions before defining the job. Vendors will answer against their broadest product capability while your team evaluates a narrower deployment.
Do not use the RFP to:
- select a role;
- replace security or legal review;
- reproduce an industry framework without scoping;
- award points for irrelevant enterprise features;
- demand sensitive vendor internals that can be verified another way;
- force every answer into Yes;
- turn roadmap promises into current capability; or
- declare a pilot successful before it runs.
The RFP converts a defined role into proof requests.
§ 02What Should You Send Vendors Before the 50 Questions?
Every vendor should receive the same 1-2 page scope sheet. Without it, “supported” can mean different things in every response.
Role scope sheet
| Field | Buyer-provided example |
|---|---|
| Role | Weekly acquisition analyst |
| Recurring outcome | Accepted channel review by 10 a.m. Monday |
| Valid triggers | Weekly schedule after approved data refresh |
| Systems | Analytics, ad accounts, CRM, document store, task tracker |
| Data classes | Internal business data; no payment credentials |
| Allowed actions | Read, calculate, draft, save, update task |
| Prohibited actions | Change budgets, edit source data, contact customers |
| Approval boundary | Any external distribution or spend recommendation |
| Expected volume | 4 standard runs plus 2 exception cases per month |
| Acceptance evidence | Reconciled figures, linked sources, complete template |
| Escalation | Missing source, conflicting attribution, abnormal spend |
| Supervisor | Growth lead |
| Retention need | 12 months of reports; shorter operational logs if policy allows |
| Deployment region | Buyer-specified requirement |
| Exit requirement | Export role config, tasks, evidence, artifacts, and usage |
The scope sheet should use real workflow objects: account, record, file, recipient, tool operation, task state, source, approver, and accepted outcome.
Action-risk tier
| Tier | Example | RFP depth |
|---|---|---|
| 0 | Public-web research with no authenticated tools | Condensed 15-question core |
| 1 | Internal read and draft with human publication | 50 questions, lighter contract review |
| 2 | Reversible writes or external action after approval | Full RFP and live control demonstrations |
| 3 | Consequential, irreversible, rights-affecting, regulated, or privileged action | Full RFP plus specialist security, privacy, legal, and domain review |
The tier is an internal scoping device, not an industry standard. Increase rigor when the action’s blast radius grows.
Buyer constraints
Tell vendors which answers are knockout requirements before they respond.
Examples:
- no inference-data training;
- named processing regions;
- central access revocation;
- enforceable approval before defined actions;
- exportable action evidence;
- incident notification within a contractual period;
- deletion within an accepted period;
- supported identity controls;
- a minimum service commitment; or
- an acceptable exit format.
A vendor can give an excellent answer and still be wrong for the deployment.
§ 03How Should Vendors Format Their RFP Responses?
Reject free-form sales essays. Require one structured row per question.
| Response field | Required entry |
|---|---|
| Status | Yes, Partial, No, or Not applicable |
| Scope | Product, plan, region, interface, and workload covered |
| Default/configured | Default behavior or configuration required |
| Customer responsibility | Control or operation the buyer must own |
| Evidence ID | One or more evidence artifacts |
| Limitation | Known boundary, exception, or unsupported condition |
| Roadmap | Target date and dependency, kept separate from current status |
| Contract status | Standard term, negotiable term, or non-contractual statement |
| Evidence date | Version or last-reviewed date |
| Vendor owner | Person accountable for follow-up |
Use evidence codes
| Code | Evidence artifact | Best used to prove |
|---|---|---|
| DOC | Current product, security, or policy documentation | Defined behavior and scope |
| ARC | Architecture or data-flow diagram | Components, trust boundaries, and processors |
| CFG | Redacted admin configuration or export | Available control and granularity |
| LOG | Redacted structured log or audit export | What is recorded and reconstructable |
| DEMO | Live vendor demonstration using buyer case | Executable product behavior |
| TEST | Buyer-run or jointly run test result | Behavior under buyer conditions |
| ASSURE | Independent audit, certification, or assessment report | Assessed control scope and period |
| CONT | Contract, DPA, SLA, security addendum, or order form | Enforceable commitment |
| COST | Quote, rate card, usage export, or invoice example | Commercial mechanics |
No artifact proves everything. A SOC report may support control assurance but not your role’s denial behavior. A demo may prove behavior but not create a contractual service level.
Protect sensitive evidence
Allow vendors to:
- redact customer identifiers;
- provide audit reports under NDA;
- use a secure data room;
- give an auditor’s report instead of exploit details;
- demonstrate a control without disclosing secrets; and
- describe an architecture boundary without publishing attack-enabling configuration.
The objective is sufficient assurance, not forced disclosure.
§ 04Who Should Review the RFP?
No single reviewer owns all 50 answers.
| Reviewer | Primary domains | Decision owned |
|---|---|---|
| Business/role owner | Role, workflow, acceptance, escalation | Does the product support useful work? |
| IT/integration owner | Connectors, APIs, identity, environments | Can it fit the operating stack? |
| Security/GRC | Access, logs, assurance, incident, supply chain | Is residual security risk acceptable? |
| Privacy/legal | Data flow, retention, terms, IP, liability | Are processing and contractual terms acceptable? |
| Procurement/finance | Pricing, support, renewal, exit | Is the commercial commitment controllable? |
| Evaluation owner | Demo cases, evidence ledger, open gaps | Has the vendor proved the scored claims? |
| Executive sponsor | Risk acceptance and deployment boundary | Should the organization proceed? |
CISA’s Secure by Demand Guide recommends considering product security before procurement, putting appropriate requirements into the contract during procurement, and continuing assessment after purchase. That lifecycle matters for AI employees because model, prompt, tool, memory, and provider changes can alter behavior after the contract is signed.
§ 05The 50 AI Employee Platform RFP Questions at a Glance
| Domain | Questions | What the domain establishes |
|---|---|---|
| 1. Role and operating architecture | 1-5 | Whether the product can represent standing work |
| 2. Tools, integrations, and execution | 6-10 | Whether it can act in the required systems safely |
| 3. Identity, permissions, and approvals | 11-15 | Whether authority can be constrained and revoked |
| 4. Context, knowledge, and memory | 16-20 | Whether persistent context is governed |
| 5. Data privacy and lifecycle | 21-25 | Where data goes, why, and for how long |
| 6. Security assurance and supply chain | 26-30 | How the vendor builds, tests, and responds |
| 7. Evaluation, reliability, and safety | 31-35 | Whether behavior is measured under relevant conditions |
| 8. Observability and human governance | 36-40 | Whether decisions and actions can be supervised |
| 9. Pricing, implementation, and support | 41-45 | What the role will cost and require operationally |
| 10. Contract, continuity, and exit | 46-50 | What happens when terms, providers, or needs change |
The 10 domains mirror the operating path from role definition through termination. Five Yes answers in one domain cannot compensate for a failed knockout requirement in another.
§ 06Domain 1: Role and Operating Architecture - Questions 1-5
These questions establish whether the product carries standing work or only exposes a capable model, prompt, or workflow surface.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 1 | How does the platform represent a persistent role, including purpose, goals, supervisor, boundaries, recurring responsibilities, and retirement? | DOC + CFG showing the role object and lifecycle | Persona or prompt presented as complete role persistence |
| 2 | Which state persists across sessions, shifts, interfaces, model changes, and operator changes, and which state does not? | ARC + restart DEMO + data dictionary | “The conversation remembers” without explicit state objects |
| 3 | Which schedules, messages, events, queues, webhooks, or human requests can start work, and how are duplicate, stale, or replayed triggers handled? | Trigger DOC + duplicate-event TEST | “Runs 24/7” with no idempotency or cancellation behavior |
| 4 | Which task states, dependencies, deadlines, postconditions, reopen paths, and handover objects are supported? | CFG + interrupted-task DEMO + sample handover | Only running and done, or state hidden in narrative chat |
| 5 | Which models, orchestration layers, retrieval systems, memory services, tool brokers, and execution environments participate in the service? | Versioned ARC with provider responsibilities | “Proprietary AI” without component or responsibility boundaries |
What a strong response shows
A strong response connects:
role → trigger → task state → context → plan → action → evidence → postcondition → handover
It also separates native product capability from professional services, custom code, third-party tools, and buyer-operated components.
Question 5 does not require source code. It requires enough architecture to understand the value chain, data path, control ownership, and failure dependencies.
Knockout candidates
Consider a knockout when:
- the role cannot be paused or retired;
- duplicate triggers can repeat consequential actions;
- unfinished work cannot be distinguished from completed work; or
- critical provider dependencies are undisclosed.
§ 07Domain 2: Tools, Integrations, and Execution - Questions 6-10
An AI employee creates operational risk when reasoning becomes action. Ask about exact operations, not connector logos.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 6 | For every required system, which objects and operations are supported through native API, connector, browser/computer use, code, or custom integration? | Operation-level matrix + DEMO of the buyer’s last mile | “Connects to CRM” without object, read/write, or account scope |
| 7 | How are tool inputs, parameters, outputs, schemas, and postconditions validated before and after execution? | DOC + invalid-parameter TEST + sample structured result | Model-generated parameters sent directly to consequential tools |
| 8 | How are credentials stored, rotated, scoped, isolated, audited, and prevented from entering prompts, memory, output, or logs? | ARC + CFG + relevant ASSURE evidence | Shared static credentials or secrets visible in agent context |
| 9 | What happens on timeout, partial success, rate limit, unavailable dependency, invalid response, or ambiguous tool result? | Failure-state DEMO + retry/idempotency DOC + LOG | Unlimited retries, silent partial success, or duplicate write |
| 10 | How are custom integrations developed, tested, versioned, approved, sandboxed, monitored, and maintained after an API changes? | Integration lifecycle DOC + test environment evidence | Production-first testing or undocumented customer-owned code |
What a strong response shows
The vendor should distinguish:
- connection from authorization;
- tool availability from supported operation;
- model choice from execution policy;
- retry from safe recovery;
- browser reach from durable integration; and
- artifact generation from confirmed delivery.
The question is not whether the platform can reach thousands of tools. It is whether the selected role can use five required operations within a bounded, observable action path.
Agent-specific security evidence
OWASP’s AI Agent Security Cheat Sheet identifies tool abuse, privilege escalation, data exfiltration, memory poisoning, excessive autonomy, cascading failures, and unbounded cost as agent-specific risks. Its recommendations include least-privilege tool scopes, high-impact approval, structured monitoring, cost and retry limits, and adversarial testing.
Ask the vendor to map its evidence to the threats that apply to your role.
§ 08Domain 3: Identity, Permissions, and Approvals - Questions 11-15
The full AI employee permissions and approvals framework explains how to build an action envelope. These five RFP questions ask whether a platform can enforce it.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 11 | Can access be restricted by tool, operation, resource, account, data class, recipient, value, time, volume, environment, and role? | CFG of buyer-specified envelope + denial TEST | Permission exists only at connector or prompt level |
| 12 | Which identity controls are available for human users, service identities, AI roles, administrators, and support personnel? | Identity/control matrix + ASSURE scope | Shared administrator accounts or unclear service identity |
| 13 | Can approval bind to the exact actor, action, tool, target, parameters, amount, expiry, and one-time execution? | Approval CFG + modified-parameter and replay TEST + LOG | Approval to a vague plan rather than the executed operation |
| 14 | How quickly can an administrator revoke a connector, credential, permission, trigger, role, or entire workspace, including work already in progress? | Live revocation DEMO + documented propagation behavior | Revocation applies only after a run completes or support intervenes |
| 15 | Which administrative controls support separation of duties, least privilege, multifactor authentication, role-based access, single sign-on, provisioning, and periodic access review? | Tier-specific DOC + CFG + contract/roadmap separation | “Enterprise controls” without named availability or product tier |
What a strong response shows
The response should identify:
- which controls the platform enforces;
- which controls come from the connected application;
- which controls require middleware;
- which controls are only instructions;
- which plan includes the control; and
- how a denied, expired, revoked, or altered action appears in evidence.
Do not mark a vendor down merely because a buyer-desired control sits in the connected system. Mark it down when the combined path cannot enforce or prove the boundary.
Knockout candidates
For a consequential role, common knockout failures include:
- no central stop;
- no enforceable approval;
- approval parameters can change after review;
- no operation-level restriction;
- no administrator attribution; or
- no workable identity model for the deployment.
§ 09Domain 4: Context, Knowledge, and Memory - Questions 16-20
Persistent memory improves continuity and expands the attack and privacy surface. The AI employee memory guide provides the lifecycle design; the RFP should require proof of the platform controls.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 16 | Which context types are stored - role instructions, authoritative knowledge, working context, preferences, task state, summaries, and episodic records - and where? | Context data dictionary + ARC | One undefined memory store for unrelated objects |
| 17 | How are source provenance, freshness, authority, instruction precedence, conflict, and uncertainty represented during retrieval and decision-making? | Conflicting-source TEST + source evidence | Newest or most similar content silently overrides policy |
| 18 | How is context isolated by tenant, workspace, role, user, account, customer, project, task, and environment? | Isolation ARC + negative-access TEST + ASSURE | Shared context without documented boundaries |
| 19 | Who can inspect, correct, suppress, expire, export, and delete each memory type, and how do changes propagate to derived summaries or indexes? | Correction/deletion DEMO + lifecycle DOC | Memory is persistent but cannot be located or corrected |
| 20 | How does the platform detect and contain prompt injection, memory poisoning, malicious retrieval, poisoned tool output, and sensitive-data persistence? | Threat model + adversarial TEST + incident/eval evidence | “The model is trained to ignore attacks” as the only control |
What a strong response shows
Strong memory answers are object-specific. They distinguish among:
- chat history;
- uploaded documents;
- retrieved knowledge;
- task state;
- generated summaries;
- learned preferences; and
- action evidence.
The vendor should explain which layer is authoritative, which is advisory, and what happens when two layers conflict.
Ask for propagation evidence
Delete or correct a source record, then test:
- direct retrieval;
- generated summary;
- open task;
- long-term memory;
- audit record; and
- export.
Deletion of a visible document does not prove removal from derived context, backups, or required audit evidence.
§ 10Domain 5: Data Privacy and Lifecycle - Questions 21-25
Data can flow through the platform, model providers, connector brokers, email services, cloud infrastructure, support tools, analytics, logs, memory, and external actions.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 21 | Provide a data inventory and flow diagram showing data categories, purpose, component, processor, region, transfer, storage, and recipient. | ARC + subprocessor list + processing register | Privacy policy without service-specific data flow |
| 22 | Is customer input, output, feedback, memory, tool data, or metadata used to train or improve any vendor or third-party model, and what opt-out or contract applies? | DOC + DPA/CONT + provider terms | “We do not train” without covering subprocessors or metadata |
| 23 | How is data encrypted in transit and at rest, how are keys and secrets managed, and which data may remain unencrypted for operation or support? | Security DOC + relevant ASSURE scope | Blanket “encrypted” statement without exceptions or boundaries |
| 24 | What retention schedule applies to prompts, outputs, memory, mail, files, logs, telemetry, backups, deleted objects, and support records? | Object-level schedule + deletion DEMO + CONT | One retention statement for every data type |
| 25 | How can the customer access, correct, export, restrict, and delete data during service and after termination, including derived and subprocessor-held data? | Sample export + deletion procedure + DPA terms | Manual support promise with no scope, format, or timeline |
What a strong response shows
The vendor distinguishes:
- controller and processor roles where applicable;
- customer content from service telemetry;
- active storage from backup retention;
- account deletion from object deletion;
- inference processing from training use;
- primary vendor from subprocessor; and
- legal retention from product convenience.
Do not accept “zero retention” without identifying the provider, data type, endpoint, plan, and exception.
Match the RFP to applicable law and policy
Privacy requirements depend on data, people, location, sector, and use. Ask qualified privacy and legal reviewers to tailor this domain.
The RFP should record the organization’s requirement. It should not ask the vendor to decide which law applies to the buyer.
§ 11Domain 6: Security Assurance and Supply Chain - Questions 26-30
AI employee platforms combine normal SaaS risk with models, prompts, retrieval, memory, tools, code, browsers, and third-party services.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 26 | Which independent audits, certifications, penetration tests, or assessments cover the exact service, region, period, and controls proposed? | ASSURE evidence under NDA if needed | Certification name with no scope, entity, period, or exceptions |
| 27 | Describe secure development, threat modeling, code review, dependency management, vulnerability disclosure, patching, secrets scanning, and release controls. | SDLC DOC + vulnerability policy + assurance excerpt | “Industry standard SDLC” with no practices or ownership |
| 28 | How are tenants isolated across application, storage, retrieval, memory, logs, execution environments, browser/computer sessions, caches, and support access? | Isolation ARC + test/assurance evidence | Application-level tenant ID treated as the only isolation control |
| 29 | How are model providers, open-source components, connectors, plugins, tool servers, datasets, and other subprocessors inventoried, assessed, monitored, and changed? | Supplier inventory + change policy + CONT notification | Material provider can change with no notice or review |
| 30 | Provide the incident-response process for AI misuse, data exposure, unauthorized action, cross-tenant access, prompt injection, model/provider event, and service compromise. | Incident plan summary + notification terms + exercise evidence | Cyber incident process excludes AI action or model events |
What a strong response shows
A certificate is useful when the scope matches the purchased service. It is not a universal pass.
Ask:
- Which legal entity was assessed?
- Which service and infrastructure were included?
- Which period was covered?
- Were important controls customer-operated?
- Were exceptions found?
- Is the report current?
- Does the evidence cover AI-specific components?
Use broader frameworks when the risk requires them
The 50 questions are intentionally shorter than a complete control catalog. CSA’s July 2026 AICM v1.1 overview says its AI-CAIQ contains 320 questions aligned to 247 controls. Enterprise security teams can use that resource for deeper self- and third-party assessment rather than expanding an improvised spreadsheet indefinitely.
§ 12Domain 7: Evaluation, Reliability, and Safety - Questions 31-35
The vendor must show how it knows the system works - not only that the model can produce a good answer.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 31 | Which evaluation sets test model output, tool choice, action correctness, memory, escalation, refusal, recovery, and end-to-end role outcomes? | Eval inventory + sample cases + versioned results | One benchmark used as proof for the whole platform |
| 32 | Which metrics track accepted outcomes, first-pass quality, correction time, false completion, escalation, unauthorized action, severity, latency, and cost? | Metric definitions + sample report + buyer mapping | Tokens, runs, or user ratings presented as role success |
| 33 | Which model, prompt, policy, tool, retrieval, memory, or provider changes trigger regression tests, approval, customer notice, or rollback? | Change policy + release evidence + rollback DEMO | Silent updates with no version or regression evidence |
| 34 | Which adversarial tests cover direct and indirect prompt injection, tool abuse, privilege escalation, data exfiltration, memory poisoning, approval manipulation, and runaway loops? | Threat-linked test matrix + redacted results | Generic red team with no agent action path |
| 35 | What limits and recovery mechanisms govern run duration, retries, tokens, cost, concurrency, delegation depth, tool chaining, timeout, fallback, and safe failure? | Limit CFG + failure TEST + service DOC | Unlimited autonomy or retry described as a benefit |
What a strong response shows
The evaluation set should resemble deployment:
- same role;
- same tool classes;
- same data types;
- same action risk;
- realistic failures;
- explicit acceptance;
- severity-aware errors;
- known limitations; and
- versioned results.
NIST’s AI RMF Core calls for documented test sets, evaluation under conditions similar to deployment, production monitoring, and response and recovery processes. Ask the vendor which artifacts support each lifecycle stage.
Do not confuse benchmark and role proof
A benchmark can support the capability it measures. It does not automatically prove:
- connector correctness;
- permission enforcement;
- memory isolation;
- reliable triggers;
- task handover;
- incident containment; or
- cost per accepted outcome.
Record the benchmark’s task set, date, configuration, judge, and scope - then keep the role evaluation separate.
§ 13Domain 8: Observability and Human Governance - Questions 36-40
The buyer should be able to reconstruct work without asking the same model to narrate what it remembers.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 36 | Which structured fields connect trigger, run, role, task, source, instruction, model, decision, approval, tool action, result, cost, and final state? | Redacted LOG + schema + correlation DEMO | Final chat transcript is the only evidence |
| 37 | Can authorized users search, filter, retain, redact, export, and restrict logs by role, customer, account, action, severity, and time? | Admin CFG + sample export + retention controls | Export requires bespoke engineering or excludes action detail |
| 38 | Which monitors and alerts detect failed runs, repeated retries, abnormal cost, unauthorized requests, stale work, performance regression, data exposure, and unusual tool use? | Monitoring catalog + alert TEST + routing evidence | Dashboard exists but nobody or nothing receives alerts |
| 39 | How do humans review, approve, interrupt, override, correct, roll back, reopen, escalate, and assume ownership of work? | End-to-end DEMO + policy DOC + LOG | Human-in-the-loop means only reading the final answer |
| 40 | Which administrative changes are recorded, who can investigate them, and how are incident evidence, legal holds, and audit retention handled? | Admin-history LOG + access matrix + retention terms | Prompt, policy, or permission changes are unattributed |
What a strong response shows
Strong evidence separates:
- user instruction;
- system or role instruction;
- retrieved content;
- model decision;
- policy decision;
- human approval;
- execution request;
- external-system result; and
- final task state.
The buyer needs structured evidence that shows inputs, control decisions, actions, results, and responsibility without depending on a model-generated explanation of its own behavior.
Test the investigation path
Give the evaluator one suspicious external message and ask them to find:
- which event started the task;
- which role handled it;
- which source influenced the decision;
- which instruction and model version applied;
- which approval was requested;
- which tool call executed;
- what the external system returned;
- what the platform retried;
- what the human changed; and
- how the task ended.
If the evidence exists but cannot be found during an incident, observability is incomplete.
§ 14Domain 9: Pricing, Implementation, and Support - Questions 41-45
The AI employee cost framework separates vendor price from setup, integration, review, correction, monitoring, and failure exposure. The RFP should make every vendor quote the same work.
When “build internally” remains on the shortlist, require the internal sponsor to answer the same ownership questions through a three-year build-versus-buy AI employee analysis. Framework access and prototype labor are not substitutes for a costed evaluation, security, memory, monitoring, incident, change, and exit plan.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 41 | Which units are billed for subscription, seats, credits, tokens, runs, time, tools, storage, models, background work, retries, failures, support, and overage? | COST rate card + billing definitions + invoice example | “Simple usage pricing” without failed-run or premium-model rules |
| 42 | Price the buyer’s fixed workload, including ordinary cases, hard cases, failures, review, expected correction, and forecast range. | Written same-workload quote + assumptions + sensitivity | Vendor quotes unrelated plan minimum rather than role consumption |
| 43 | Which implementation tasks, integrations, data preparation, configuration, evaluation, training, change management, and services are required, by whom, and when? | Implementation plan + RACI + dependencies + fees | “Deploy in minutes” excludes production configuration |
| 44 | Which support channels, hours, severity levels, response targets, resolution targets, escalation paths, service levels, status pages, and remedies apply? | Support policy + SLA/CONT + sample escalation | Dedicated support is marketing language, not a named service |
| 45 | Which required capabilities depend on roadmap, beta, custom work, partner products, buyer code, or a higher tier, and which will be committed in the order form? | Gap list + commercial proposal + CONT status | Roadmap capability scored as currently available |
What a strong response shows
The quote should normalize:
- task mix;
- context volume;
- tool calls;
- model tier;
- artifact type;
- retries;
- failures;
- concurrency;
- review;
- correction;
- support;
- implementation; and
- term.
The lowest subscription may create the highest operating cost. The RFP must preserve both.
Keep commercials current
Require a quote-valid-until date, renewal rules, overage treatment, usage-export format, price-change notice, termination right, and assumptions.
Do not freeze a pricing page into a multi-year decision without a contract.
§ 15Domain 10: Contract, Continuity, and Exit - Questions 46-50
The last five questions determine whether the evidence survives a dispute, outage, provider change, acquisition, or termination.
| # | Exact RFP question | Minimum evidence request | Red flag |
|---|---|---|---|
| 46 | Who owns or may use customer inputs, outputs, configurations, evaluations, feedback, memory, fine-tuned assets, logs, and generated artifacts? | Terms + DPA + IP clauses + usage-right matrix | Broad secondary use hidden outside the response |
| 47 | How do the agreement and DPA allocate responsibility, liability, indemnity, warranty, acceptable use, customer supervision, unauthorized action, data event, and third-party failure? | Standard and proposed CONT language | Product promise conflicts with liability or user-responsibility term |
| 48 | What continuity and fallback exist for service, model, cloud, connector, region, key personnel, or subprocessor outage or termination? | Business continuity summary + dependency/fallback ARC + SLA | Critical role depends on one undisclosed service with no fallback |
| 49 | In which documented formats can the buyer export role configuration, prompts/instructions, tasks, memory, artifacts, logs, evaluations, usage, and evidence during service and at exit? | Complete sample export + schema + transition support terms | PDF summary is the only exit artifact |
| 50 | What is the termination and deletion process, including access cut-off, credential revocation, export window, active deletion, backup expiry, subprocessor deletion, certification, and exceptions? | Exit runbook + deletion CONT + sample certificate or record | “Data deleted on request” with no object scope or timeline |
What a strong response shows
The contract should not say “autonomous work” in the product description and leave every consequence undefined in the terms.
Review the following materials for contradictions:
- product documentation;
- security response;
- privacy policy;
- DPA;
- terms;
- SLA;
- support policy;
- order form;
- statement of work; and
- insurance or assurance evidence.
NIST’s third-party guidance
The NIST Generative AI Profile recommends use-case-based supplier assessment, third-party inventory, ongoing monitoring, incident planning, fallbacks, and contract terms addressing ownership, rights, security, serious-incident notification, response times, and critical support.
That is the procurement lesson: a current answer must become an owned control, continuing monitor, or enforceable term where the risk requires it.
§ 16Which Evidence Pack Should Every Vendor Return?
The questionnaire is incomplete without a deliverable list.
Core evidence pack
Request:
- product and architecture overview;
- role, task-state, trigger, and handover documentation;
- operation-level integration matrix;
- permission and approval configuration evidence;
- identity and administrator-control matrix;
- data inventory and flow diagram;
- current subprocessor list;
- object-level retention and deletion schedule;
- sample data and evidence export;
- redacted structured run and action logs;
- evaluation methodology and sample result;
- change and regression policy;
- incident-response summary and notification term;
- security assurance package;
- same-workload commercial quote;
- implementation plan and RACI;
- support policy and proposed SLA;
- standard terms and DPA;
- business continuity and exit summary; and
- current gaps, betas, custom dependencies, and roadmap items.
The request should specify an acceptable secure delivery method.
Evidence freshness rules
| Artifact | Freshness question |
|---|---|
| Documentation | Which product version and date does it cover? |
| Architecture | Has the provider, region, or data path changed? |
| Configuration | Is it available in the proposed tier today? |
| Log/export | Was it generated from the proposed product surface? |
| Test result | Which model, prompt, tools, policies, and date were used? |
| Assurance report | Which entity, service, period, and exceptions apply? |
| Contract | Is this the operative standard or negotiated language? |
| Quote | How long is it valid and which assumptions can change? |
An undated screenshot is not durable evidence.
Evidence ownership
For every accepted artifact, record:
- buyer owner;
- source vendor;
- secure location;
- reviewed date;
- approved use;
- expiration or refresh trigger;
- linked RFP question;
- linked risk or requirement; and
- final disposition.
Procurement evidence becomes operational evidence only when someone maintains it.
§ 17Which Six Tasks Should the Vendor Demonstrate Live?
The RFP asks what the product can do. The proof session checks a small number of claims before a full pilot.
Demo 1: Normal role path
The role should:
- wake from the approved trigger;
- load the correct context;
- complete the requested work;
- use the permitted tool;
- satisfy the postcondition;
- attach evidence; and
- enter the correct final state.
Demo 2: Denied action
Ask the role to perform an operation outside its action envelope.
Require:
- technical denial;
- clear reason;
- no external side effect;
- correct task state;
- escalation if appropriate; and
- structured log evidence.
Demo 3: Duplicate or stale trigger
Send the same event twice, then replay it after the task has expired.
Require:
- one valid task;
- no duplicate external action;
- visible deduplication or rejection;
- bounded retry; and
- correlation to the original event.
Demo 4: Conflicting or poisoned context
Place a malicious or stale instruction inside a retrieved document and provide a newer authoritative policy.
Require:
- source hierarchy;
- conflict detection;
- no unauthorized persistence;
- safe refusal or escalation; and
- evidence of which source controlled the decision.
Demo 5: Tool failure after partial progress
Make the destination unavailable after one step succeeds.
Require:
- correct partial-state record;
- no repeated side effect;
- bounded retry;
- recovery or compensation path;
- human escalation; and
- final reconciliation.
Demo 6: Investigation and export
Give a reviewer the task identifier and ask them to reconstruct:
- trigger;
- role;
- context;
- instruction;
- model;
- decision;
- approval;
- action;
- result;
- retry;
- cost; and
- final state.
Then export the relevant record.
These six tasks validate shortlist claims. They do not replace a representative pilot with a baseline, sample design, duration, KPI thresholds, and go/no-go criteria.
§ 18How Should You Evaluate the 50 RFP Answers?
Use the evidence scale and weights already defined in the AI employee platform selection guide. Keep this RFP focused on producing the evidence that guide needs.
Classify every answer
| Classification | Meaning | Procurement action |
|---|---|---|
| Verified | Evidence directly supports the scoped answer | Score at the demonstrated evidence level |
| Conditional | Works only with a plan, configuration, partner, service, or buyer control | Price and assign the dependency |
| Contract pending | Behavior is shown but commitment is unresolved | Do not treat as enforceable |
| Unverified | Assertion or incomplete artifact | Request proof or lower the evidence score |
| Gap | Required capability is unavailable | Apply knockout or risk treatment |
| Not applicable | Requirement does not apply to the scoped role | Record the rationale |
Keep gates outside the average
Potential gates include:
- high-risk approval;
- central revocation;
- safe failure and interruption;
- action evidence;
- acceptable training/data-use terms;
- required identity control;
- acceptable incident term;
- retention and deletion fit; and
- workable exit.
One failed gate can stop the procurement even if the vendor has 49 strong answers.
Price every “Partial”
Partial can mean:
- higher subscription tier;
- professional services;
- custom integration;
- buyer-built middleware;
- connected-system control;
- manual review;
- additional vendor;
- policy process;
- delayed roadmap; or
- accepted risk.
Translate it into owner, cost, delivery date, operational burden, and residual risk.
§ 19Which RFP Answers Should Become Contract Terms?
Contractualization should follow impact.
| Claim | When to bind it | Possible document |
|---|---|---|
| Data use and training | Whenever customer data terms matter | DPA, security addendum, agreement |
| Processing region | When residency or transfer is material | DPA/order form |
| Retention and deletion | When policy or law requires a period | DPA/security addendum |
| Incident notification | When delayed notice changes response | DPA/security addendum |
| Availability and support | When the role becomes operationally important | SLA/order form |
| Model or subprocessor change notice | When change alters risk or compliance | DPA/service terms |
| Required controls or product tier | When procurement depends on them | Order form/SOW |
| Pricing and overage | Always for material spend | Order form |
| Export and transition | When switching cost matters | Agreement/SOW |
| Remediation or termination right | When a control failure is unacceptable | Agreement/addendum |
Do not attempt to put every marketing sentence into a contract. Bind the facts whose failure would change the purchase or create unacceptable exposure.
Resolve contradictions
If a sales answer, security questionnaire, privacy policy, terms, and order form disagree:
- identify the conflict;
- ask the vendor to resolve it in writing;
- identify which document controls;
- update the RFP disposition; and
- do not rely on the most favorable non-binding statement.
The procurement record should show what the buyer actually accepted.
§ 20What Are the Most Common AI Vendor RFP Red Flags?
| Red flag answer | What it usually hides | Required follow-up |
|---|---|---|
| “Yes” with no artifact | Claim scope is unknown | Evidence ID, tier, date, limitation |
| “Enterprise-grade” | No named control | Exact control and enforcement |
| “Industry standard” | No standard, scope, or assessor | Named framework and evidence |
| “We use trusted models” | Application controls are omitted | Full value-chain responsibility |
| “We do not train on your data” | Metadata or subprocessors are excluded | Data-type and provider-specific term |
| “Human in the loop” | Human sees only final output | Exact approval and interruption path |
| “Full audit trail” | Narrative transcript only | Structured fields and export |
| “Encrypted” | Key, state, or exception unclear | Transit/at-rest scope and key ownership |
| “Delete on request” | Backup or derivative data omitted | Object schedule and certification |
| “Unlimited autonomy” | Limits and stop conditions absent | Retry, cost, action, and time caps |
| “Roadmap” in current-status column | Future capability inflates score | Separate No/Partial from roadmap |
| “Customizable” | Services or buyer engineering required | Scope, owner, price, and maintenance |
| “Compliant with…” | Buyer treats compliance as inherited | Scope, role, evidence, and shared responsibility |
| “Proprietary security” | Control cannot be assessed | Alternative assurance under NDA |
| “Best-in-class model” | Role controls remain unproved | Role-specific test and acceptance evidence |
A vague answer is not always evidence of a bad product. It is evidence of incomplete diligence.
§ 21Which 15 Questions Should a Small Team Ask First?
For a low-risk role, begin with these question IDs:
- Q1: How is the persistent role represented?
- Q3: Which triggers start work, and how are duplicates handled?
- Q6: Which exact tool operations are supported?
- Q9: What happens on timeout or partial tool failure?
- Q11: How granular are permissions?
- Q13: Does approval bind to the exact action?
- Q14: How is access or the role revoked?
- Q16: Which context and memory types are stored?
- Q18: How is context isolated?
- Q21: Where does data flow?
- Q22: Is any data used for training or improvement?
- Q24: What are retention and deletion periods?
- Q30: How are AI-specific incidents handled?
- Q36: Can the action path be reconstructed and exported?
- Q41: Which work consumes billable units?
Add all 50 questions before the role gains external action, sensitive data, broad access, operational criticality, or enterprise commitment.
The small-team version should reduce administrative weight, not remove the controls that contain the role.
§ 22How Does CellCog Map to This RFP?
CellCog should receive the same 50 questions as every shortlisted vendor.
Its current public pages answer part of the RFP and identify the gaps a buyer should take to sales, security, privacy, and legal review. The observations below reflect public materials reviewed in July 2026.
Evidence available publicly
| RFP area | Current public CellCog material | What it supports |
|---|---|---|
| Role architecture | AI Employees product and support material | Goals, permissions, inbox, schedule, memory, task board, shifts, wake conditions, KPIs, approvals, and handovers are described |
| Tools and action | Product and connector material | Connected apps, computer/file work, browser work, email, and external action are described |
| Team architecture | AI Employees and AI organization material | Delegation, handoffs, and AI employees managing AI employees are described |
| Data flow | CellCog privacy policy | Data categories, providers, US storage, encryption, retention, and rights are described |
| Responsibility and service terms | CellCog terms of service | User responsibility, autonomous action, connected services, credits, restrictions, and service terms are described |
| Commercial model | Pricing and role pages | Credit-based self-serve pricing and usage estimates are described |
| Model capability | Benchmark and product pages | Research benchmark plus multi-format product capabilities are described |
Public evidence helps a buyer prepare. It does not answer every operation-, tenant-, plan-, assurance-, SLA-, incident-, export-, or contract-specific question.
Questions to send CellCog directly
Ask CellCog to identify and demonstrate:
- operation- and resource-level permission granularity;
- approval binding, expiry, replay protection, and revocation behavior;
- structured action-log fields and export;
- memory provenance, correction, isolation, and deletion propagation;
- identity and administrator controls available by plan;
- current independent security assurance and scope;
- incident notification and support commitments;
- service levels for the intended deployment;
- fixed-workload credit consumption;
- complete role, memory, task, evaluation, and evidence export; and
- deletion and transition behavior at termination.
These are evidence requests, not claims that the capability is absent. Public materials do not establish the complete answer.
What CellCog should be asked to prove live
CellCog AI Employees publicly describes scheduled and event-triggered shifts, persistent memory, task boards, approvals, tools, KPI dashboards, and handovers. Use one scoped role to demonstrate:
- wake;
- context retrieval;
- tool action;
- denied action;
- exact approval;
- connector failure;
- handover;
- action evidence;
- usage record; and
- export.
Then send the same script to every shortlisted general-purpose platform.
Move qualified vendors from comparison to evidence review
First, compare candidate shapes and current vendor claims on the AI employee platform comparison. Then send the 50-question response template to each shortlisted vendor. If CellCog remains qualified, use its contact path to request the plan-specific evidence and commercial terms that public pages cannot establish.
§ 23Copyable AI Employee Platform RFP Response Template
Procurement header
- Buyer:
- Role:
- Business outcome:
- Action-risk tier:
- Systems:
- Data classes:
- Regions:
- Volume:
- Required identity controls:
- Required service level:
- Retention requirement:
- Contract term:
- Knockout requirements:
- Response due:
Vendor header
- Vendor:
- Product:
- Proposed plan:
- Proposed region:
- Primary respondent:
- Security respondent:
- Privacy/legal respondent:
- Commercial respondent:
- Response date:
- Evidence data-room location:
Question response row
| Field | Vendor response |
|---|---|
| Question ID | |
| Status: Yes / Partial / No / N/A | |
| Scoped answer | |
| Supported product/plan/region | |
| Default or configuration required | |
| Customer responsibility | |
| Evidence IDs | |
| Known limitation | |
| Roadmap item and date | |
| Contract status | |
| Vendor owner | |
| Buyer disposition |
Evidence register
| Evidence ID | Type | Title/version | Date | Questions supported | Restrictions | Buyer owner | Refresh trigger |
|---|---|---|---|---|---|---|---|
Gap register
| Question | Gap | Knockout? | Compensating control | Owner | Cost | Due date | Contract treatment | Final decision |
|---|---|---|---|---|---|---|---|---|
§ 24The Short Version
An AI employee platform RFP works when every answer has:
- a scoped role;
- a clear status;
- a named product tier;
- an enforcement or operating explanation;
- customer responsibility;
- current evidence;
- a known limitation;
- a roadmap field separate from present capability;
- contractual status; and
- an accountable owner.
The 50 questions cover the entire chain:
role → trigger → context → tool → permission → action → evidence → outcome → incident → exit
Send the same chain to every vendor. Accept a claim only at the level its evidence supports. Put material promises into the appropriate agreement. Then let the pilot answer the remaining question: can this platform perform the role reliably enough to buy?
Q1How many questions should an AI platform RFP contain?
Use enough questions to cover the scoped role’s architecture, actions, data, controls, evidence, cost, and exit. This template uses 50. A low-risk role can begin with 15. A high-risk enterprise assessment may require the CSA AI-CAIQ’s broader control set plus sector, legal, and organizational requirements.
Q2Is an AI RFP different from a SaaS security questionnaire?
Yes. A SaaS questionnaire covers important cloud, identity, development, incident, and privacy controls. An AI employee RFP adds role persistence, triggers, tools, action authorization, model and prompt change, memory, prompt injection, evaluation, human override, cost limits, and end-to-end action evidence.
Q3Should vendors answer every question with yes or no?
No. Require Yes, Partial, No, or Not applicable, followed by scope and proof. A precise No can be safer than a vague Yes. Keep roadmap capability separate from current status and identify configuration or customer-owned controls.
Q4Does SOC 2 or ISO certification replace the RFP?
No. Independent assurance may support specific controls over a defined service and period. It does not prove your role’s trigger, tool operation, denied action, approval binding, memory behavior, accepted outcome, price, or exit format. Review scope and then test the role-specific path.
Q5Should a vendor disclose prompts, source code, or security secrets?
Not necessarily. Ask for sufficient assurance through redacted architecture, control documentation, configuration, logs, demonstrations, independent reports, and contract terms. Use an NDA or secure data room when appropriate. Do not create security risk by demanding unnecessary sensitive details.
Q6What happens after vendors return the RFP?
Validate artifacts, classify each answer, apply knockout gates, score the evidence against the platform-selection framework, resolve contract-critical gaps, and invite only qualified vendors into a bounded proof session or pilot. Preserve the response, evidence date, owner, limitation, and refresh trigger.
