The build-vs-buy AI employee decision is not “engineers or subscription.” It is a choice about which team will own the agent’s production lifecycle: role design, model changes, tools, credentials, memory, evaluations, observability, approvals, incidents, support, and eventual exit.
Building can be the right answer when the agent is strategically differentiating, the workflow requires unusual infrastructure, and your team already operates secure production software. Buying usually wins when the role is operational rather than proprietary, speed matters, and a platform can meet your non-negotiable control and evidence requirements.
Do not compare a 3-week prototype with a monthly plan. Compare 3 years of the same accepted workload. Use the AI employee platform evaluation framework for the broader shortlist, then make the architecture and ownership decision separately.
On this page · 14 sectionsOpen
- What Does “Build vs Buy an AI Employee” Actually Mean?
- When Should You Build an Internal AI Employee?
- When Should You Buy an AI Employee Platform?
- What Must a Production AI Employee Include Beyond the Prototype?
- How Do Build and Buy Compare Across the 12 Ownership Layers?
- What Does the Three-Year Build-vs-Buy Cost Model Include?
- Which Hidden Maintenance Costs Change the Decision?
- How Should Security, Privacy, and Supplier Risk Affect the Choice?
- How Do You Decide Between Build, Buy, and Hybrid?
- Where Does CellCog Fit in the Build-vs-Buy Decision?
- Which Build-vs-Buy Assumptions Fail Most Often?
- Who Owns the Decision After the Contract or Build Approval?
- What 30-Day Process Produces a Defensible Decision?
- Make the Decision on Lifecycle Ownership
- Build when the agent itself creates durable competitive advantage, the integration or policy boundary is genuinely unusual, and you can fund a permanent product, platform, security, and evaluation function.
- Buy when the role uses common business systems, time to useful work matters, and a vendor can prove the required permissions, memory, logs, recovery, data, and exit controls.
- A framework lowers code required for a prototype; it does not remove production ownership for evaluations, identity, secrets, monitoring, incidents, model changes, or support.
- Compare 3-year total cost of ownership, including internal labor, infrastructure, model and tool usage, review, correction, downtime, security, compliance, and expected failure exposure.
- Use a hybrid architecture when your proprietary policy, data, or orchestration should remain internal but commodity execution, research, media, or document capabilities can be delegated.
- Treat failed permission, audit, recovery, data, or exit requirements as knockout gates. A cheaper option does not compensate for an unsafe authority boundary.
§ 01What Does “Build vs Buy an AI Employee” Actually Mean?
“Build” and “buy” are endpoints on a control spectrum. Most production systems sit between them.
| Architecture | You own | Vendor owns | Best fit | Main risk |
|---|---|---|---|---|
| Build from model APIs | Agent runtime, tools, memory, orchestration, UI, evals, security, operations | Foundation model service | Differentiating agent product or unusual control plane | Permanent engineering and reliability burden |
| Build on agent framework | Application, production controls, most integrations, evaluation, operations | Framework primitives; model service | Strong engineering team that wants reusable primitives | Prototype speed mistaken for production readiness |
| Buy configurable platform | Role configuration, policies, acceptance, supervision, vendor management | Runtime, common tools, product surfaces, platform operations | Business role using common workflows | Platform limits or vendor dependency |
| Hybrid | Proprietary data, policy, orchestration, system of record | Selected execution or specialist capabilities | Differentiated control with commodity execution | Ambiguous boundary and duplicated monitoring |
| Outsource implementation | Requirements, acceptance, governance, ongoing ownership | Initial build or configuration services | Temporary capability gap | Knowledge leaves with the implementer |
The table’s practical point is that “buy” never means “vendor owns the outcome.” Your business still owns the role, permitted actions, approval policy, source quality, acceptance criteria, and consequences.
Build is an operating model
A built agent needs a product owner, technical owner, security owner, domain reviewer, and incident owner. Those responsibilities remain after the first workflow passes.
The operating question is not whether your engineers can call a model and tool. It is whether the organization can keep the complete system useful when the model, prompt, API, data, policy, and business process change.
Buy is a supplier relationship
A purchased platform concentrates more of the runtime and product surface in a supplier. You still need to verify the supported plan, configuration dependency, customer responsibility, data path, support route, and exit mechanism.
Use the 50-question AI employee platform RFP to convert vendor claims into documents, configurations, logs, demonstrations, tests, assurance material, contract terms, and pricing evidence.
Hybrid is a boundary, not a compromise slogan
A useful hybrid pattern keeps proprietary policy and source-of-truth decisions inside your environment while delegating bounded execution to a platform or external agent.
For example, an internal service can decide which accounts qualify for research, strip prohibited fields, send an approved research task, validate the returned artifact, and require a human before updating the CRM. CellCog’s Agent-to-Agent surface is relevant to this model because it exposes capabilities to other agents through API, SDK, plugins, and skills; the buyer still has to validate the live interface and control boundary.
§ 02When Should You Build an Internal AI Employee?
Build when control produces more durable value than the operating burden consumes.
The agent is part of the product
If customers buy your agent’s behavior, latency, reliability, workflow, or proprietary intelligence, the agent is not merely an internal efficiency tool. Its architecture may be core intellectual property.
Strong build signals include:
- the agent is a paid customer-facing product;
- proprietary orchestration materially changes the outcome;
- response time or infrastructure placement is product-critical;
- your evaluation data creates a defensible learning loop;
- product economics require control of model routing and caching;
- customers require deployment conditions a platform cannot support; or
- the agent’s failure and recovery experience is part of your product promise.
The stronger the differentiation, the more reasonable permanent ownership becomes.
The workflow has a nonstandard trust boundary
Build may be justified when the role must operate inside unusual networks, use specialized hardware, meet strict residency requirements, or call proprietary systems through controls that no shortlisted platform can enforce.
“We prefer control” is too weak. Document the exact control:
- credential must never leave a named environment;
- inference must occur in an approved region;
- every write must pass a proprietary policy engine;
- artifacts must use a customer-owned encryption boundary;
- tool execution must happen inside an existing sandbox;
- approval must bind to exact parameters and expire; or
- a regulated record must be reconstructed from internal logs.
If a vendor can prove the same control, the requirement no longer automatically favors build.
You already operate the necessary platform disciplines
A strong software team is necessary but insufficient. An agent adds nondeterministic behavior, tool selection, long-running state, model dependency, and natural-language attack surfaces.
Build becomes more credible when your organization already has:
- production identity and access management;
- secrets management and scoped service accounts;
- versioned deployment and rollback;
- structured application and security logging;
- on-call ownership and incident response;
- privacy and supplier-risk operations;
- test-data governance;
- model and prompt evaluation infrastructure;
- cost and rate controls; and
- domain experts who can grade the actual work.
If 7 of these 10 capabilities do not exist, the internal build includes an organizational platform project, not only an agent project.
§ 03When Should You Buy an AI Employee Platform?
Buy when the workflow is important but the infrastructure beneath it is not your differentiator.
The role uses common business systems
Inbox triage, research, recurring reporting, content production, CRM preparation, document work, spreadsheet analysis, and task coordination often use common interfaces. A managed platform can spread the cost of connectors, schedules, workspaces, memory, and monitoring across many customers.
The buyer still needs one action-level boundary. A platform advertising an email connector has not proved that it can separate read, draft, send, delete, and administrative authority.
Useful deployment speed has economic value
Time to first demo is a weak metric. Time to the first accepted recurring outcome is better.
Buying can avoid months of work on:
- account and workspace administration;
- agent role configuration;
- inbox and trigger infrastructure;
- task-state interfaces;
- common connectors;
- artifact storage;
- usage metering;
- approval surfaces;
- operator dashboards; and
- support workflows.
That avoided work has value only if the vendor surface actually fits the role. A fast start followed by connector rewrites and manual review is not fast time to value.
The vendor can carry commodity maintenance
Model providers change APIs, capabilities, limits, safety behavior, and prices. Connected applications change authentication, schemas, rate limits, and permissions. Production agent systems also require retries, queues, timeouts, trace storage, and recovery behavior.
A platform earns margin when it absorbs a meaningful share of that change while preserving your required controls. It does not earn margin merely by wrapping a model in a polished chat interface.
Your internal opportunity cost is high
An engineering team building an internal marketing or operations agent is not building customer features, data infrastructure, reliability work, or security improvements.
Calculate the contribution of the displaced roadmap, not just payroll. A 4-person team diverted for 6 months consumes 24 person-months before the first year of maintenance begins.
§ 04What Must a Production AI Employee Include Beyond the Prototype?
A prototype proves that one path can work. Production needs to prove that bounded work keeps working under normal variation and known failure.
Role and state
The system needs a durable definition of:
- purpose and recurring outcome;
- supervisor and accountable owner;
- goals and key performance indicators;
- valid triggers;
- task states;
- permitted and prohibited actions;
- approval conditions;
- source and memory rules;
- completion evidence; and
- retirement or suspension.
Without state, the “employee” is a repeated prompt with implied continuity.
Tools and authority
Tools need schemas, authentication, authorization, input validation, error handling, idempotency, and audit events. The model’s decision to request a tool call is not itself authorization.
The current OWASP AI Agent Security Cheat Sheet recommends least privilege, explicit approval for high-impact actions, action previews, audit trails, interrupt and rollback paths, structured monitoring, and limits on retries, cost, and tool chains.
Memory and context
Persistent memory adds continuity and a new data system. It needs provenance, scope, read/write permissions, freshness, correction, deletion, isolation, and poisoning defenses.
A governed AI employee memory lifecycle should separate source records, working state, approved knowledge, and retained evidence so a useful note does not silently become permanent truth.
Evaluation and regression
An agent can complete the task through different trajectories. Evaluation therefore needs outcome checks, trace review, tool assertions, policy checks, and repeated trials where variability matters.
Anthropic’s current guide to evaluating AI agents separates tasks, trials, graders, and transcripts. It also distinguishes capability evaluations - what the agent can do - from regression evaluations - whether it still performs previously reliable work.
Monitoring and response
Production operation needs:
- version and configuration identifiers;
- trigger, tool, approval, state, and artifact events;
- latency, usage, retries, and cost;
- acceptance and correction results;
- anomalies and policy denials;
- owner alerts;
- suspension and rollback;
- incident classification; and
- recovery evidence.
The NIST AI RMF Core calls for testing before deployment and during operation, performance measurement under conditions similar to deployment, production monitoring, incident response, recovery, override, and decommissioning.
§ 05How Do Build and Buy Compare Across the 12 Ownership Layers?
Score the same role, not generic architecture preferences.
| Ownership layer | Build advantage | Buy advantage | Evidence to compare |
|---|---|---|---|
| Role model | Unlimited custom schema | Ready configuration and operator UI | Role export and lifecycle demo |
| Models | Routing, fine control, direct contracts | Managed model selection and upgrades | Model policy and change notice |
| Tools | Exact internal interfaces | Existing connector catalog | Live tool, scope, and failure test |
| Identity | Native enterprise control | Faster packaged administration | Service identity and revocation proof |
| Permissions | Custom policy engine | Productized approval surface | Denied-action and approval-binding test |
| Memory | Custom storage and residency | Managed persistence and retrieval | Data flow, correction, deletion, export |
| Evals | Exact domain suite | Reusable evaluation features | Representative task and regression result |
| Observability | Native telemetry integration | Packaged operator view | Complete trace and export |
| Reliability | Custom SLO and topology | Shared managed operations | Support, status, recovery, commitment |
| Security/privacy | Direct architecture control | Vendor assurance and shared investment | Architecture, policy, assurance, contract |
| Economics | Avoid platform margin at scale | Avoid fixed team and platform cost | 3-year same-workload model |
| Exit | Source and runtime ownership | Faster initial adoption | Config, data, artifact, log export test |
The build advantage is strongest where customization matters. The buy advantage is strongest where repeated platform work can be shared. Neither column wins by default.
Use non-compensating gates
Before weighted scoring, define requirements that cannot be averaged away:
- forbidden data use;
- missing action-level permission;
- no reliable revocation;
- no required audit evidence;
- unacceptable recovery path;
- unsupported deployment boundary;
- no workable export or deletion; or
- no accountable operating owner.
A 15% lower estimate should not rescue a failed critical control.
Weight control only where it creates value
Teams often award “maximum control” a high score without naming what the control changes. That biases the spreadsheet toward build.
Replace generic control with testable attributes: policy precision, deployment location, model routing, data retention, change timing, interface behavior, observability, or exit. Score only the attributes the role needs.
Weight speed as accepted work
Do not score “implementation speed” from contract signature to first demo. Score time to:
- configured role;
- working integration;
- representative shadow run;
- passed negative tests;
- accepted output;
- approved bounded action; and
- stable recurring operation.
This prevents configuration theater from winning the decision.
§ 06What Does the Three-Year Build-vs-Buy Cost Model Include?
Use total lifecycle cost, not platform price or initial developer time.
Build-cost formula
For one role over 36 months:
Build TCO = discovery + engineering + platform + security/compliance + model/tool usage + domain review + operations + maintenance/change + expected failure + exit
Include loaded internal labor rather than salary alone. Include shared platform cost only in proportion to realistic use, but do not assume future agents will absorb today’s investment unless those agents are actually funded.
Buy-cost formula
Buy TCO = subscription + usage + implementation + integration + internal owner + review/correction + security/procurement + change management + expected failure + exit
Buying removes some build lines and retains most outcome-ownership lines. The vendor cannot grade your business result, decide your risk tolerance, or accept accountability for your process unless the contract explicitly says so.
Illustrative 3-year comparison
The example below is an illustrative scenario, not a market benchmark or CellCog quote.
| Cost line over 36 months | Build scenario | Buy scenario | Hybrid scenario |
|---|---|---|---|
| Discovery and role design | $30,000 | $20,000 | $25,000 |
| Initial implementation | $240,000 | $45,000 | $120,000 |
| Security, privacy, and review setup | $60,000 | $35,000 | $50,000 |
| Platform/subscription | $0 | $108,000 | $54,000 |
| Model, tool, and infrastructure usage | $72,000 | Included/variable | $54,000 |
| Maintenance and integration change | $180,000 | $30,000 | $90,000 |
| Monitoring and incident operations | $120,000 | $36,000 | $72,000 |
| Domain review and correction | $90,000 | $90,000 | $90,000 |
| Expected failure exposure | $45,000 | $45,000 | $45,000 |
| Exit or migration provision | $30,000 | $25,000 | $30,000 |
| Illustrative 3-year total | $867,000 | $434,000 | $630,000 |
In this scenario, buy costs 50% of build. That does not prove buying is generally cheaper. It shows that fixed platform labor dominates one-role economics under the stated assumptions.
Run the sensitivity test
Change 6 variables:
- number of production roles;
- initial build team;
- ongoing maintenance allocation;
- vendor subscription and usage;
- integration change rate;
- review and correction effort.
If build serves 20 genuinely reusable roles, shared infrastructure may change the result. If each role needs unique tools, policy, and evaluation, the assumed reuse may never arrive.
Calculate accepted-outcome cost
Total cost alone can reward a cheaper system that creates more rejected work.
Cost per accepted outcome = total lifecycle cost for period ÷ accepted outcomes
The 7-layer AI employee cost framework adds setup, usage, review, correction, monitoring, and expected failure exposure to the vendor price. Use it for the fixed workload in both options.
§ 07Which Hidden Maintenance Costs Change the Decision?
Maintenance is where prototype comparisons fail.
Model and prompt change
A model update can change tool choice, formatting, latency, refusals, or instruction following. A prompt change can fix one case and break another.
Budget for:
- version inventory;
- staging evaluation;
- regression suite execution;
- trace comparison;
- controlled release;
- rollback; and
- incident communication when behavior changes.
A managed platform may carry the model integration, but the buyer still needs evidence that the role’s accepted outcome did not regress.
Connector and schema drift
APIs change. OAuth scopes change. CRM fields are renamed. A document template gains a required section. A browser flow moves a button.
Connector maintenance includes authentication renewal, schema mapping, rate-limit behavior, retries, duplicate prevention, sandbox tests, and downstream reconciliation.
Evaluation-set maintenance
The eval suite must change when the work changes. Old examples can become easy, irrelevant, or contaminated.
Maintain:
- representative normal cases;
- boundary and exception cases;
- prohibited-action tests;
- past incident regressions;
- current source and policy versions;
- multiple trials for variable behavior; and
- periodic human calibration of model graders.
Human operating load
Review does not disappear because an agent is internally built or externally bought. Someone must resolve unclear inputs, correct outputs, approve actions, inspect exceptions, and revise the role.
Track reviewer minutes per accepted outcome, escalation precision and recall, oldest waiting task, and correction recurrence. Otherwise human work becomes invisible in both cost models.
§ 08How Should Security, Privacy, and Supplier Risk Affect the Choice?
Build changes where risk sits; it does not make risk disappear.
Building increases component ownership
An internal system may still depend on model APIs, cloud services, vector stores, observability vendors, browser automation, and application APIs. The supply chain may be broader than the “build” label suggests.
NIST’s Generative AI Profile recommends inventorying third-party components, defining incident ownership, rehearsing third-party response, managing fallback, and including relevant requirements in contracts.
Buying increases vendor concentration
A platform can centralize sensitive context, tools, memory, and action. Review:
- data categories and purposes;
- storage and inference locations;
- model and subprocessors;
- training or improvement use;
- encryption;
- retention and deletion;
- tenant isolation;
- employee access;
- breach and incident process;
- service continuity;
- model or subprocessor change;
- export; and
- termination.
Public policy is useful evidence, not the complete deployment contract.
CellCog-specific diligence
CellCog publicly describes US storage on Google Cloud Platform, encryption in transit and at rest, Firebase authentication, Google Cloud Storage and MongoDB services, named LLM providers, and zero-data-retention API terms for inference content. Its public material does not currently establish SOC 2, ISO 27001, HIPAA compliance, a formal uptime SLA, or a dedicated trust center.
That makes CellCog a candidate to evaluate, not an automatic pass. Buyers should confirm current architecture, plan-specific controls, legal terms, support, and role behavior directly against their deployment.
§ 09How Do You Decide Between Build, Buy, and Hybrid?
Use a two-stage decision: knockout gates first, weighted value second.
Stage 1: prove feasibility
| Gate | Build question | Buy question | Hybrid question |
|---|---|---|---|
| Outcome | Can we define and grade the role? | Can the platform represent it? | Is the responsibility boundary explicit? |
| Authority | Can policy be enforced outside model judgment? | Can permissions and approvals be demonstrated? | Which side authorizes each action? |
| Data | Can required residency and lifecycle be operated? | Are data terms and controls acceptable? | Is data minimized at the boundary? |
| Evidence | Can a run be reconstructed? | Are logs and exports sufficient? | Can traces be correlated end to end? |
| Recovery | Can we stop, roll back, and recover? | What can buyer and vendor each do? | Which owner acts during failure? |
| Economics | Is 3-year funding credible? | Is fixed-workload pricing controllable? | Does duplicated control erase savings? |
| Exit | Can another team operate the code? | Can role data and artifacts leave? | Can either component be replaced? |
Any failed mandatory gate removes the option or narrows the use case.
Stage 2: score value
Use a 100-point score only after feasibility:
- outcome fit: 20;
- control fit: 15;
- evidence and evaluation: 15;
- time to accepted work: 10;
- 3-year cost: 15;
- reliability and support: 10;
- adaptability: 5;
- data and supplier posture above the gate: 5;
- exit and portability above the gate: 5.
Document the evidence behind every score. A point without an artifact is an opinion.
Decision rules
Choose build when:
- strategic differentiation is high;
- no vendor passes a critical boundary;
- reusable internal platform demand is funded;
- engineering and operating ownership are named; and
- the 3-year case survives realistic maintenance.
Choose buy when:
- the role is operational and bounded;
- a platform passes mandatory controls;
- speed to accepted work matters;
- internal platform work has high opportunity cost; and
- vendor dependence has a workable exit.
Choose hybrid when:
- proprietary policy or data must remain internal;
- specialist execution is available as a service;
- the interface can be typed, minimized, logged, and tested; and
- end-to-end ownership is explicit.
§ 10Where Does CellCog Fit in the Build-vs-Buy Decision?
CellCog fits the configurable-platform and hybrid sides of the decision.
Configurable standing roles
CellCog AI Employees are publicly described as standing agents configured with a role, goals, KPIs, permissions, an inbox, schedules or wake conditions, task state, memory, handovers, approvals, and employee-to-employee delegation.
Those product objects can reduce the amount of role-runtime infrastructure a buyer builds. They do not remove the buyer’s work of defining the job, sources, authority, acceptance, and supervision.
General-purpose execution
The same employee runs on CellCog’s multimodal Super-Agent base. This may reduce the need to assemble separate research, code, document, spreadsheet, image, video, audio, and application-production tools.
Breadth creates its own evaluation question: does the selected role need that range, and can the team test the tools it will actually use? The general-purpose versus specialized AI agent comparison shows when shared context earns the breadth and when instruction, permission, evaluation, or capacity boundaries justify specialists. Evaluate the bounded workload, not the catalog.
Programmatic delegation
CellCog also exposes its capabilities to external agents. In a hybrid system, an internal agent could retain the proprietary decision and delegate a scoped artifact task to CellCog.
Before committing, test authentication, data minimization, task state, artifact delivery, timeout, retry, cost, evidence, and failure handling. A programmatic connection is an interface; it is not an operating contract by itself.
Current pricing is an input, not the answer
CellCog’s public pricing spans self-serve credit plans and custom organization arrangements. Because pricing and credit consumption can change, use the current CellCog pricing page and request a fixed-workload quote instead of freezing a generic monthly estimate into the decision.
§ 11Which Build-vs-Buy Assumptions Fail Most Often?
Most bad decisions are not arithmetic errors. They are ownership assumptions that the team never writes down.
| Assumption | Why it sounds reasonable | What usually remains | Test before deciding |
|---|---|---|---|
| “The prototype is 80% of the build” | The visible happy path works | Identity, negative tests, logs, recovery, support, change control | List every production layer and name its owner |
| “The vendor handles security” | The platform operates the runtime | Buyer data classification, account scope, approvals, review, contract, incident coordination | Run a shared-responsibility workshop |
| “Our data makes build unique” | Internal data is proprietary | A vendor may connect to the same governed sources without owning them | Separate data ownership from runtime ownership |
| “One platform will create lock-in” | Role state accumulates in the product | Internal code can also depend deeply on one model, framework, engineer, or cloud service | Perform a replacement exercise for both options |
| “We will reuse the platform across 20 roles” | Shared services should create scale | Role-specific tools, evals, policies, and reviewers may dominate | Fund the next 3 roles before crediting all 20 |
| “Buying is predictable” | Subscription prices are visible | Usage, services, review, correction, upgrades, and renewal can vary | Request 3 workload bands and change terms |
| “Building gives complete control” | Source code is internal | External models, APIs, packages, and specialists remain dependencies | Produce the complete component and supplier inventory |
The test column is more useful than debating the assumption. If the team cannot produce the named artifact, discount the claim in the decision model.
Prototype completion is not production completion
A prototype usually demonstrates the most favorable path with a small data sample, one operator, broad development credentials, and manual recovery. The remaining work is not cosmetic.
Before calling an internal prototype production-ready, require:
- separate development and production environments;
- scoped non-human identities;
- centrally revocable credentials;
- schemas and validation for every tool;
- idempotency or duplicate-action handling;
- retry, timeout, and spending ceilings;
- source and policy versioning;
- persistent task-state rules;
- memory write and deletion controls;
- approval binding to exact actions;
- outcome and trace evaluations;
- security abuse cases;
- operational dashboards and alerts;
- suspension, rollback, and recovery;
- on-call or business-hours response ownership;
- runbook and decision record;
- data retention and export; and
- a tested retirement path.
If 12 of these 18 items are still manual or undefined, the system may be a promising prototype, but its remaining lifecycle cost is material.
Vendor responsibility is not buyer accountability
A managed platform can operate infrastructure and product controls. It cannot infer which customer promise, legal requirement, financial boundary, or reputational risk your role must honor.
Create a shared-responsibility map with at least 8 rows:
- role and outcome definition;
- data classification and lawful use;
- account connection and credential scope;
- platform security and availability;
- action permission and approval configuration;
- model, tool, and product change;
- output review and business acceptance;
- incident detection, communication, containment, and recovery.
For each row, identify buyer owner, vendor owner, evidence, notification path, and fallback. “Shared” without named actions is unassigned work.
Proprietary data does not automatically require proprietary runtime
Data can remain buyer-owned and governed while a platform receives the minimum context required for one task. Conversely, an internal runtime may still send sensitive prompts, tool results, or files to external model and observability services.
Decide at the field and action level:
- which source remains authoritative;
- which fields the agent may read;
- which fields may enter model context;
- which fields may be stored in memory;
- which fields may appear in logs;
- which artifacts may leave the environment;
- which processor handles each copy; and
- when every copy is deleted.
This data-flow comparison often reveals that “build versus buy” was the wrong abstraction. The real choice is between two different processing chains.
Lock-in exists on both sides
Vendor lock-in can include proprietary role configuration, memory, task state, connector behavior, evaluation data, and usage economics. Internal lock-in can include undocumented code, one framework’s abstractions, custom infrastructure, a single model’s behavior, and knowledge held by 1-2 engineers.
Test portability with a 1-day exercise:
- export the role and instruction set;
- export approved source material;
- export open and completed task state;
- export representative traces and evaluation cases;
- reproduce one integration contract;
- run 5 known tasks in a replacement environment; and
- estimate gaps, labor, and elapsed time.
Run the same exercise for build and buy. Portability is demonstrated replacement effort, not an API checkbox.
Reuse should be credited only when funded
Internal-platform proposals often spread fixed cost across a future portfolio. This can make a 1-role build appear economical before another role has an owner or budget.
Use 3 cost views:
- committed case: only funded roles;
- probable case: funded roles plus approved next roles, discounted for schedule risk;
- aspirational case: the broad portfolio, shown separately and never used as the base case.
Also separate reusable services from role-specific work. Identity, logging, queues, secrets, and deployment may be reusable. Domain sources, tool permissions, acceptance tests, exception handling, and reviewer calibration often are not.
§ 12Who Owns the Decision After the Contract or Build Approval?
Architecture approval is the start of operating accountability.
| Owner | Build responsibility | Buy responsibility | Evidence at monthly review |
|---|---|---|---|
| Business owner | Outcome, scope, acceptance, escalation | Same | Accepted outcomes, exceptions, changed requirements |
| Product/platform owner | Roadmap, runtime, interface | Configuration, vendor roadmap, adoption | Version, backlog, unresolved gaps |
| Engineering/IT | System, integrations, reliability | Connections, identity, buyer-side failures | Error rate, latency, connector changes |
| Security/privacy | Architecture, testing, incidents | Supplier review, configuration, shared incidents | Denials, access, tests, open risks |
| Evaluation owner | Test harness and releases | Buyer cases and vendor-change checks | Capability and regression results |
| Finance/procurement | Internal cost governance | Usage, invoice, renewal, contract | Forecast, variance, commitments |
| Supervisor | Review, approval, correction | Same | Review time, overrides, recurring corrections |
| Executive risk owner | Residual-risk acceptance | Residual-risk acceptance | Gate failures, incidents, continue/stop decision |
The business owner and supervisor remain in both columns because neither architecture can outsource the definition of useful and acceptable work.
Define the review cadence
Use a weekly operating review during the pilot, monthly service review during stable operation, and event-driven review after material changes.
A material change includes:
- new model or model version;
- new tool, connector, or permission;
- changed memory source or retention rule;
- new external audience;
- changed evaluation threshold;
- pricing or plan migration;
- material incident;
- new subprocessor; or
- change in role outcome.
The correct cadence follows change and consequence, not an arbitrary annual procurement calendar.
Predefine switch triggers
A decision becomes safer when exit conditions are written while everyone still favors the selected option.
Possible triggers include:
- critical permission or data gate failure;
- accepted-output rate below threshold for 2 review periods;
- review effort above the business-case ceiling;
- repeated uncontained incident;
- unsupported mandatory integration;
- vendor change that breaks a contractual condition;
- internal staffing loss that leaves no operator;
- cost above the high-case budget for 2 periods; or
- failed export or recovery test.
The trigger does not always require immediate termination. It requires a recorded decision to remediate, narrow, switch, or stop.
§ 13What 30-Day Process Produces a Defensible Decision?
The decision should end with evidence, not architecture preference.
Days 1-5: define the same role
- write the recurring outcome;
- choose representative tasks;
- identify systems and data;
- map allowed and prohibited actions;
- define approval and escalation;
- set acceptance tests;
- name the supervisor; and
- document the 36-month workload.
Days 6-12: design both options
For build, create a component diagram, staffing plan, security boundary, evaluation plan, delivery milestones, and operations RACI.
For buy, send the RFP, map product capabilities to the same role, identify gaps, request evidence, and obtain same-workload pricing.
For hybrid, draw the interface and assign authorization, state, logging, retry, and incident ownership to one side.
Days 13-20: test decisive uncertainty
Do not build or configure the whole role. Test the 3-5 assumptions that can change the decision:
- hardest required integration;
- most important accepted outcome;
- prohibited-action denial;
- evidence reconstruction;
- recovery from tool failure; and
- export or replacement path.
Days 21-26: model three years
Build low, base, and high cases. Vary staffing, roles, usage, correction, integration change, incidents, vendor price, and migration.
Record assumptions separately from sourced facts. Use ranges where uncertainty is real.
Days 27-30: decide and preserve conditions
Produce a 1-page decision record:
- selected architecture;
- alternatives rejected and why;
- mandatory conditions;
- unresolved risks;
- accepted residual risk;
- cost range;
- owner;
- next test;
- review date; and
- exit trigger.
The next step is a bounded AI employee pilot with a baseline, representative cases, approval gates, and explicit exit criteria - not organization-wide deployment.
§ 14Make the Decision on Lifecycle Ownership
Build when owning the agent creates strategic value and you are willing to own the platform after launch. Buy when the role matters more than the infrastructure and a supplier can prove the required boundary. Use hybrid when proprietary judgment can be separated cleanly from commodity execution.
For a buy path, evaluate CellCog AI Employees against the same role, evidence requests, and 3-year cost model as every other option. The next decision is not “Can it demo?” It is “Who can operate this role safely, measurably, and economically for 36 months?”
Q1Is it cheaper to build an AI employee?
Not automatically. Building can avoid platform margin but adds fixed engineering, security, evaluation, integration, monitoring, incident, and maintenance costs. It becomes economically stronger when those capabilities already exist, the work is strategically differentiating, or the platform can be reused across enough funded roles.
Q2How many engineers does an internal AI agent need?
There is no universal number. Count capabilities rather than titles: product ownership, agent/runtime engineering, integrations, security, reliability, evaluation, and domain review. One engineer can prototype several layers; production accountability still needs every layer covered.
Q3Does an agent framework make buying unnecessary?
No. Anthropic’s building effective agents guidance recommends starting with the simplest architecture and warns that frameworks can add abstraction that obscures underlying prompts and responses. A framework can accelerate implementation, but your team still owns production controls, evaluation, security, operations, and support.
Q4Can we start by buying and build later?
Yes, if exit is designed before adoption. Preserve role definitions, source data, evaluation cases, artifacts, action evidence, and integration contracts in exportable forms. Avoid letting proprietary knowledge exist only in opaque vendor memory or configuration.
Q5What is the best first action after choosing build or buy?
Write the decision, mandatory controls, cost assumptions, and disconfirming evidence. Then run a representative, bounded pilot with a baseline, negative tests, approval gates, and explicit go, revise, switch, or stop criteria.
