AI market research is trustworthy only when a reviewer can trace how a question became a conclusion.
A polished report is not enough. The workflow needs to preserve:
- the business question;
- definitions and scope;
- the source plan;
- search and retrieval activity;
- inclusion and exclusion decisions;
- source identity and date;
- claim-level evidence;
- calculations;
- assumptions;
- contradictory findings;
- uncertainty;
- reviewer decisions; and
- the conditions that trigger a refresh.
An AI employee can accelerate collection, normalization, comparison, evidence extraction, and drafting. It should not turn search rank into source quality, treat a snippet as evidence, invent a missing market number, generalize a benchmark beyond its scope, or silently choose the conclusion the sponsor prefers.
The operating pattern is:
Decision question > research protocol > source plan > retrieval > screening > extraction > contradiction check > synthesis > review > accepted artifact > monitored refresh
On this page · 11 sectionsOpen
- What Is an AI Market Research Workflow?
- How Do You Turn a Business Need into a Research Question?
- How Do You Create a Source Plan?
- How Should Search, Retrieval, and Screening Work?
- How Do You Verify a Source and Extract Evidence?
- How Do You Build Claim and Contradiction Ledgers?
- How Do You Synthesize Without Overclaiming?
- What Should Human Review Check?
- How Do You Manage Updates, Memory, and Reuse?
- What Does a Complete Workflow Example Look Like?
- How Should You Pilot and Measure the Workflow?
- Start with the decision, not a broad topic. Define the population, geography, time window, category, comparison dimensions, and required confidence.
- Build a source hierarchy before searching. Prefer primary official, company, regulatory, statistical, and research sources for load-bearing claims.
- Preserve queries, source URLs, retrieval dates, document versions, inclusion/exclusion reasons, and snapshots where policy allows.
- Extract claims into a ledger. Distinguish direct fact, calculation, estimate, assumption, forecast, opinion, and recommendation.
- Search for contradiction deliberately. A credible synthesis explains conflicting definitions, periods, populations, methods, and incentives.
- Require a human reviewer to approve definitions, source sufficiency, calculations, uncertainty, recommendation, and decision use.
- Date the report and define refresh triggers. Market evidence can change even when the document still looks current.
§ 01What Is an AI Market Research Workflow?
An AI market research workflow is a controlled process for answering a bounded market question with traceable evidence.
It can support questions such as:
- Which customer segment has the strongest observable problem?
- How do 8 named competitors position and package a defined capability?
- What changed in a category during the last quarter?
- Which regulations or standards affect a proposed launch?
- What public evidence supports or contradicts an assumed market size?
- How does a buyer’s current workflow differ across 3 target roles?
- Which product claims are common, differentiated, or unsupported?
It should not promise objective truth from unlimited web browsing.
Research artifact types
| Artifact | Primary question | Typical output |
|---|---|---|
| Market landscape | Who participates and how are they categorized? | Entity map and inclusion rule |
| Competitor comparison | How do named alternatives differ? | Dimension-by-evidence matrix |
| Customer/problem research | What problem and workflow evidence exists? | Segment and problem synthesis |
| Market sizing | What population and value assumptions support a range? | Top-down and bottom-up model |
| Trend/change monitor | What materially changed since the last accepted period? | Source-linked change brief |
| Regulatory scan | Which current requirements may affect the decision? | Issue map for qualified review |
| Technology assessment | What capability, maturity, and constraint evidence exists? | Evidence-led capability brief |
| Vendor due diligence | What public and supplied evidence supports a vendor decision? | Risk, gap, and verification register |
Each artifact needs a different protocol.
The AI employee examples guide describes market research as a bounded role pattern. This article owns the research operations needed to turn that pattern into a reviewable artifact.
Accepted outcome
One accepted research artifact:
- answers the agreed question;
- uses the current protocol;
- includes required sources and segments;
- links every material claim;
- reconciles calculations;
- exposes contradiction and uncertainty;
- passes domain review;
- stays inside use and distribution rules; and
- carries an as-of date and refresh condition.
“The agent finished researching” is not an acceptance rule.
§ 02How Do You Turn a Business Need into a Research Question?
Begin with the decision that the research will support.
Research charter
| Field | Example |
|---|---|
| Decision | Choose the first industry segment for a 90-day pilot |
| Primary question | Which of 3 candidate segments has the strongest fit with the defined workflow? |
| Geography | United States |
| Time window | Evidence current through July 29, 2026 |
| Segment definition | Organization and role criteria written in protocol |
| Unit of analysis | Company, buyer role, workflow, and observable signal |
| Dimensions | Problem frequency, urgency, data readiness, budget evidence, risk, competition |
| Required sources | Official statistics, company primary sources, interviews supplied by buyer |
| Exclusions | Unsupported estimates, anonymous claims, legal conclusions |
| Deliverable | 12-page decision brief plus source and calculation ledgers |
| Reviewer | Market strategy lead and domain specialist |
| Decision deadline | August 15, 2026 |
Define terms before searching
| Term | Required definition |
|---|---|
| Market | Which products, services, workflows, and substitutes count? |
| Customer | Organization, account, user, buyer, or beneficiary? |
| Size | Revenue, spending, volume, organizations, users, or transactions? |
| Growth | Nominal, real, annual, compound, or period-over-period? |
| Competitor | Direct product, service, internal labor, automation, or no-change alternative? |
| Adoption | Trial, purchase, active use, workflow coverage, or retained use? |
| Demand | Search, stated interest, budget, purchase, usage, or renewal evidence? |
Without definitions, the workflow can combine incompatible figures into one confident paragraph.
Use a question tree
For the segment decision:
- What workflow is being improved?
- Which organizations perform it?
- Which roles own and review it?
- How frequently does it occur?
- What evidence shows pain, delay, error, or cost?
- What digital inputs and systems exist?
- What constraints or approvals apply?
- Which alternatives are used now?
- What budget or purchase evidence exists?
- What would disconfirm segment fit?
The best-task framework for AI employees can supply task-suitability dimensions. Market research should not use category enthusiasm as proof that a workflow is suitable.
§ 03How Do You Create a Source Plan?
Define source classes, authority, and expected limitations before retrieval.
Source hierarchy
| Tier | Source class | Strong use | Common limitation |
|---|---|---|---|
| 1 | Law, regulator, official statistics, standards body | Requirements and official measures | Scope, lag, jurisdiction |
| 1 | Company filing, product docs, pricing, terms | First-party company facts | Selective presentation and change |
| 1 | Original research paper or dataset | Method-specific finding | Population and method boundaries |
| 2 | Reputable industry research with disclosed method | Category context and estimates | Sponsor, sample, paywall, definition |
| 2 | Direct interview or supplied operational data | Buyer workflow and lived evidence | Sample and respondent bias |
| 3 | High-quality reporting with named sources | Event and market context | Secondary interpretation |
| 3 | Expert analysis with transparent evidence | Hypothesis and synthesis | Author perspective |
| 4 | Aggregator, directory, snippet, anonymous post | Discovery only | Weak provenance and duplication |
Tier does not decide truth automatically. A primary company page is authoritative for what the company currently says, not for the objective superiority of its product. An official statistic may be authoritative for its defined series but irrelevant to a different population.
Source-to-question map
| Research subquestion | Preferred evidence | Fallback | Exclude |
|---|---|---|---|
| Company revenue | Regulatory filing or audited report | Company investor release | Unsourced directory estimate |
| Product capability | Current official documentation | Recorded verified demonstration | Affiliate comparison |
| Current price | Official pricing/contract quote | Reseller where authoritative | Old review |
| Employment measure | Official labor statistics | Disclosed industry survey | Search snippet |
| Buyer pain | Buyer interviews and operational records | Transparent research survey | Generic opinion |
| Academic finding | Original paper and method | Systematic review | Citation without source inspection |
| Regulation | Official text and regulator guidance | Qualified legal analysis | AI-generated legal conclusion |
Official data systems can be especially useful when the protocol matches their scope. The U.S. Securities and Exchange Commission’s EDGAR data APIs provide company submissions and extracted XBRL financial-statement data. The U.S. Bureau of Labor Statistics Public Data API provides published time-series data. Neither source answers a market question by itself; each supplies defined evidence.
Record source requirements
For every required class, state:
- minimum number or coverage;
- date/freshness;
- geography;
- language;
- primary versus secondary;
- acceptable method;
- access and license;
- sensitive-data rule;
- snapshot policy;
- citation format; and
- reviewer.
§ 04How Should Search, Retrieval, and Screening Work?
Search is a reproducible activity, not a hidden prelude to writing.
Query log
| Field | Example |
|---|---|
| Query ID | Q-2026-0729-014 |
| Research question | Segment A workflow frequency |
| Source/system | BLS |
| Query/search string | Approved series IDs and years |
| Filters | United States, 2021-2025 |
| Run time | July 29, 2026, 9:14 a.m. PT |
| Results returned | 5 series |
| Included | 2 |
| Excluded | 3, wrong occupation definition |
| Operator/version | Research workflow v1.3 |
| Reviewer | Market analyst |
Retrieval record
| Field | Required value |
|---|---|
| Source ID | Stable internal ID |
| Canonical URL or DOI | Direct destination |
| Title | Exact title |
| Author/publisher | Named entity |
| Publication/update date | Source date |
| Retrieval date | When accessed |
| Version | Filing, release, model, page, or report version |
| Source type/tier | Protocol classification |
| Scope | Population, geography, period, method |
| Rights/access | Public, licensed, supplied, restricted |
| Snapshot/hash | Where policy permits |
| Inclusion reason | Subquestion supported |
| Limitations | Known boundary |
Crossref’s REST API documentation describes retrieval of scholarly metadata such as DOI records, authors, publication information, licenses, funders, updates, and retraction-related metadata. Metadata helps identify and reconcile a work; it does not replace reading the source or evaluating its method.
Screen in 2 passes
Pass 1: identity and relevance.
- Is the source what the result claims it is?
- Does the title/abstract/summary address the question?
- Is the date and geography eligible?
- Is it duplicate or superseded?
- Is full evidence available?
Pass 2: evidence and method.
- Does the source contain the needed claim or data?
- Is the measure defined?
- Is the population appropriate?
- Is the method inspectable?
- Are conflicts, sponsorship, or limitations disclosed?
- Can it be cited and used under policy?
Use exclusion codes
| Code | Reason |
|---|---|
| E01 | Wrong population |
| E02 | Wrong geography |
| E03 | Outside time window |
| E04 | Duplicate or superseded |
| E05 | No primary evidence available |
| E06 | Definition incompatible |
| E07 | Method insufficient for claim |
| E08 | Access or rights restriction |
| E09 | Unverifiable identity |
| E10 | Outside decision scope |
Do not delete excluded records. Preserve enough information to show what was considered and why it was not used.
The PRISMA 2020 statement in The BMJ is designed for transparent reporting of systematic reviews, not ordinary commercial market research. Its emphasis on reporting why a review was done, how sources were identified and selected, and what was found is a useful discipline. Do not label a market brief “systematic” unless the method actually meets the relevant standard.
§ 05How Do You Verify a Source and Extract Evidence?
Verification tests identity, authority, scope, method, and currency.
Source verification
| Check | Question |
|---|---|
| Identity | Is this the canonical publisher, filing, dataset, or paper? |
| Author | Who created it, and in what capacity? |
| Date | Publication, event, reporting, and update dates? |
| Version | Is it current, corrected, retracted, or superseded? |
| Scope | Which population, period, geography, and category? |
| Method | How was evidence collected, measured, and analyzed? |
| Incentive | Funding, sponsor, commercial interest, or advocacy? |
| Citation | Does the cited passage support the specific claim? |
| Rights | Can it be stored, quoted, shared, or transformed? |
| Stability | Will the URL or data change without versioning? |
Extract evidence, not convenient wording
For each claim, store:
- source ID;
- source location or field;
- exact variable and unit;
- population;
- period;
- geography;
- relevant excerpt within usage limits;
- paraphrase;
- calculation;
- source limitation;
- confidence;
- researcher note; and
- reviewer decision.
Do not copy a publisher’s conclusion into the report as if it were a raw observation.
Classify the claim
| Claim type | Example | Evidence requirement |
|---|---|---|
| Direct fact | Company filed Form 10-K on date X | Primary record |
| Measure | Series value was 114.2 in period Y | Dataset, series, unit, period |
| Calculation | 8.1% change | Inputs and formula |
| Estimate | Addressable organizations range 18k-24k | Model and sensitivity |
| Forecast | Adoption may reach range by year | Method, assumptions, scenario |
| Assumption | 35% of firms fit workflow rule | Explicit owner-approved input |
| Judgment | Segment has moderate readiness | Rubric and evidence |
| Recommendation | Run a pilot in segment B | Alternatives, tradeoff, owner |
The multimodal-agent guide is relevant when evidence spans tables, charts, PDFs, images, audio, video, and interfaces. Every modality still needs a source identity, extraction method, and acceptance check.
Handle interviews and surveys as bounded evidence
Interviews can reveal language, process, exceptions, incentives, and hypotheses that public sources cannot. They do not automatically estimate how common a behavior is in the wider market.
For every interview, preserve:
- recruitment source;
- inclusion criteria;
- respondent role;
- organization segment;
- date;
- consent and permitted use;
- interview guide version;
- interviewer;
- recording/transcript status;
- redactions;
- coded observations;
- direct quotation permissions;
- researcher interpretation;
- contradictions;
- reviewer; and
- retention or deletion date.
Separate 4 layers:
| Layer | Example | Permitted claim |
|---|---|---|
| Respondent statement | “We reconcile this manually each Friday.” | This respondent reported the workflow |
| Coded observation | 9 of 12 interviewees described weekly manual reconciliation | Pattern in the supplied sample |
| Interpretation | The task may have recurring automation value | Hypothesis supported by bounded evidence |
| Market claim | Most U.S. operations teams do this weekly | Not supported without representative evidence |
Do not convert 9 of 12 into 75% of the market.
Surveys need:
- target population;
- sampling frame;
- recruitment method;
- response rate;
- questionnaire wording and order;
- field dates;
- weighting;
- exclusions;
- missing-data treatment;
- uncertainty;
- sponsor; and
- complete result tables where available.
One percentage without the question wording and denominator is weak evidence.
Protect supplied research data
Interview notes, customer lists, recordings, and operational data can be confidential or personal.
| Control | Required decision |
|---|---|
| Purpose | Which research question permits use? |
| Access | Which role can view raw versus redacted data? |
| Model/tool | May the data be processed by this service and configuration? |
| Storage | Where can raw and transformed data be stored? |
| Citation | Can the respondent or company be named? |
| Reuse | Can evidence support later projects? |
| Retention | When is raw material deleted or reviewed? |
| Withdrawal | How is consent withdrawal handled? |
| Export | Can data leave the approved workspace? |
| Incident | Who responds to unexpected exposure? |
Use respondent codes in the analysis ledger where identity is unnecessary. Do not place sensitive raw transcripts into a shared briefing merely because the final quote was approved.
Score source sufficiency, not truth
A rubric can help reviewers compare evidence consistency. It cannot prove that a source is correct.
Example 0-2 fields:
| Field | 0 | 1 | 2 |
|---|---|---|---|
| Authority | Unknown/indirect | Relevant secondary | Primary/canonical for claim |
| Identity | Unverifiable | Partially verified | Canonical and verified |
| Currency | Outside decision window | Usable with qualification | Current for purpose |
| Scope fit | Wrong population/definition | Partial fit | Direct fit |
| Method | Not disclosed | Partly disclosed | Inspectable and appropriate |
| Claim fit | Does not support claim | Supports narrower claim | Direct support |
| Independence | Duplicates same origin | Unclear dependency | Independent origin |
| Limitation | Missing | Partial | Explicit and reflected |
Possible rules:
- 0-6: discovery only;
- 7-10: supporting context with qualification;
- 11-13: usable for a bounded claim with review;
- 14-16: strong evidence for the defined claim, still subject to contradiction and domain review.
The score applies to a source-claim pair. The same source can score 15 for one claim and 5 for another.
Define sufficiency by claim consequence
| Claim consequence | Minimum evidence pattern |
|---|---|
| Descriptive context | 1 inspected suitable source with limitations |
| Material comparison | Primary evidence for each compared entity |
| Market estimate | Transparent model plus at least 2 input paths |
| Strategic recommendation | Multiple evidence types and alternative analysis |
| Legal/regulatory conclusion | Official text plus qualified professional review |
| Safety/security decision | Authoritative evidence, domain owner, and control validation |
| Public product claim | Current source, exact claim fit, legal/brand review as required |
A weak source does not become strong because 20 pages repeat it. Trace apparent corroboration back to origin; several articles may all copy one estimate.
§ 06How Do You Build Claim and Contradiction Ledgers?
The claim ledger connects every material statement to evidence.
Claim ledger
| Claim ID | Claim | Type | Evidence | Support | Status |
|---|---|---|---|---|---|
| C-014 | Segment A contains 21,400 eligible entities | Estimate | S-021, S-033, model M-04 | Mixed direct and modeled | Review |
| C-015 | Workflow occurs at least monthly for most interviewees | Observation | 9 of 12 interviews | Direct, small supplied sample | Qualified |
| C-016 | Competitor X lists feature Y | Direct fact | Official docs, retrieved date | Direct | Accepted |
| C-017 | Feature Y is fully equivalent | Judgment | No sufficient evidence | Unsupported | Reject |
Status options:
Proposed > supported > qualified > contradicted > rejected > accepted > superseded
Contradiction ledger
| Issue | Source A | Source B | Possible cause | Resolution |
|---|---|---|---|---|
| Market size | $2.1B | $7.8B | Different category boundaries | Report both definitions; do not combine |
| Company count | 24,200 | 31,600 | Different years and size thresholds | Normalize or select matching series |
| Product capability | Docs say supported | Customer report says unreliable | Availability versus production performance | Separate capability and observed quality |
| Growth | 14% | 6% | Nominal vs inflation-adjusted | Use matched basis |
| Adoption | 48% | 17% | Trial versus active production | Define stages separately |
Search for disagreement deliberately:
- alternative category definitions;
- prior and current versions;
- primary source versus secondary summary;
- positive and null findings;
- sponsor and independent research;
- different geographies;
- different customer sizes;
- different time windows;
- observed behavior versus stated intent; and
- capability availability versus accepted performance.
Resolve, qualify, or preserve
Use one of 4 outcomes:
- Resolve: one source is wrong, outdated, or out of scope.
- Normalize: convert compatible measures to a common definition.
- Qualify: evidence supports a narrower claim.
- Preserve: credible sources disagree; explain why and keep uncertainty.
Do not average incompatible estimates merely to produce one number.
The ODNI ICD 203 Analytic Standards require source-quality discussion, clear uncertainty, separation of information from assumptions and judgments, and analysis of alternatives. Those practices are useful for market synthesis even though the workflow and subject matter are different.
§ 07How Do You Synthesize Without Overclaiming?
Synthesis should answer the research question while preserving the shape of the evidence.
Synthesis structure
| Section | Required content |
|---|---|
| Decision | The choice or action the research supports |
| Scope | Category, segment, geography, time, and definitions |
| Method | Sources, search, screening, extraction, limitations |
| Findings | Evidence-led answers by subquestion |
| Contradictions | Material disagreements and treatment |
| Estimates | Models, assumptions, ranges, sensitivity |
| Alternatives | Plausible competing interpretations |
| Recommendation | Named owner’s proposal and tradeoffs |
| Uncertainty | What is unknown and what would change the view |
| Appendix | Source, query, exclusion, claim, and calculation ledgers |
Use calibrated wording
| Evidence state | Appropriate wording |
|---|---|
| Direct current primary evidence | “The filing reports…” |
| Multiple aligned sources | “The reviewed sources consistently indicate…” |
| Small supplied sample | “In 9 of 12 supplied interviews…” |
| Modeled range | “Under the stated assumptions, the range is…” |
| Conflicting sources | “Estimates differ because…” |
| Incomplete evidence | “Available evidence is insufficient to determine…” |
| Unsupported hypothesis | “This remains a hypothesis to test…” |
Do not use “proves,” “guarantees,” “the market,” or “customers want” when the evidence supports only a bounded observation.
Build each finding from an evidence chain
A strong finding has 6 parts:
- Bounded answer: the narrow conclusion supported by the reviewed evidence.
- Evidence: the primary facts, observations, or measures.
- Scope: population, geography, period, and definition.
- Contradiction: credible evidence that differs or limits the conclusion.
- Uncertainty: missing information, method limits, and unstable assumptions.
- Decision implication: what the finding changes, without making the decision.
Example:
In the 12 supplied interviews with U.S. operations leaders at 50-500-person firms, 9 respondents described at least weekly manual reconciliation across 2 or more systems. That supports recurring workflow pain in this small, recruited sample. It does not estimate prevalence across the full segment. Two respondents said the work was already handled satisfactorily by deterministic automation, and 1 performed it monthly. Before treating the pattern as a market-level finding, the buyer should test a broader sampling frame and separate integration gaps from interpretation-heavy work.
The paragraph is longer than “75% of operations teams have this problem,” but it preserves what was actually observed.
Keep evidence close to the sentence it supports
Avoid one citation at the end of a paragraph containing 5 different claims. Link or footnote at the smallest practical claim unit.
Check:
- does the source support the subject of the sentence?
- does it support the verb?
- does it support the quantity or degree?
- does it cover the stated period and geography?
- is a calculation distinguishable from the source value?
- is the wording stronger than the source?
- is the cited version the one inspected?
A source showing that a feature exists does not support “customers find the feature reliable.” A survey showing purchase intent does not support adoption. A current price does not support the price last year unless an archived version exists.
Separate absence of evidence from evidence of absence
“No public documentation found” can mean:
- the capability does not exist;
- it exists but is undocumented;
- it is limited to a plan or customer;
- the search protocol missed it;
- the page is inaccessible;
- the terminology differs; or
- the feature changed after the search window.
Report the correct state:
- documented;
- verified through another approved method;
- vendor-stated but not documented;
- not found under the protocol;
- contradicted;
- not applicable; or
- unknown.
Do not turn “not found” into “does not have” without sufficient evidence.
Keep market size as a model
Use both top-down and bottom-up paths where feasible.
Top-down value = published category value x eligible share adjustments
Bottom-up value = eligible entities x qualifying units/entity x annual value/unit
Example:
| Input | Low | Base | High |
|---|---|---|---|
| Eligible entities | 18,000 | 21,000 | 24,000 |
| Qualifying workflows/entity | 1.2 | 1.6 | 2.0 |
| Annual addressable value/workflow | $4,000 | $6,000 | $8,000 |
| Implied annual range | $86.4M | $201.6M | $384.0M |
These figures are fictional.
Show sensitivity:
- a 10% change in eligible entities changes the total 10%;
- a 25% change in workflow frequency changes the total 25%;
- a 20% change in annual value changes the total 20%; and
- correlated assumptions can widen the range more than one-at-a-time tests show.
The purpose is not to manufacture precision. It is to expose which assumptions drive the decision.
§ 08What Should Human Review Check?
Research review needs domain, method, and decision perspectives.
Review roles
| Reviewer | Primary responsibility |
|---|---|
| Research owner | Protocol, completeness, source and claim ledgers |
| Domain expert | Definitions, causal logic, missing evidence, practical context |
| Data/finance reviewer | Calculation, unit, period, model, sensitivity |
| Legal/privacy reviewer | Restricted topics, rights, privacy, legal-use boundary |
| Decision owner | Relevance, alternatives, recommendation, action |
| Editor | Clarity, claim-to-source fit, unsupported language |
Review checklist
- Does the report answer the original decision question?
- Are category and segment definitions stable throughout?
- Are required source classes represented?
- Is every material claim linked to inspected evidence?
- Does each citation support the exact sentence?
- Are dates, populations, geographies, units, and methods visible?
- Are estimates labeled and reproducible?
- Are contradictory sources included fairly?
- Are assumptions and judgments separated from observations?
- Are alternative explanations considered?
- Are unavailable or restricted sources disclosed?
- Is copyrighted material used within allowed limits?
- Are sensitive inputs protected?
- Is the recommendation attributed to a human owner?
- Does the as-of date match the evidence?
The AI executive briefing workflow can consume an accepted research artifact and surface the decision. It should not skip the research review merely because the final briefing item is short.
Sample review
Give 2 reviewers the same 20 claims and ask them to rate:
- supported;
- supported with qualification;
- contradicted;
- unsupported;
- source quality;
- confidence;
- decision relevance; and
- required correction.
Use disagreement to improve the rubric and example set. Do not treat one reviewer’s acceptance as proof that the process is calibrated.
§ 09How Do You Manage Updates, Memory, and Reuse?
Research becomes stale at different speeds.
Refresh policy
| Evidence type | Possible trigger |
|---|---|
| Product documentation | Page or version change |
| Pricing | Before every purchase decision |
| Regulatory requirement | Official update or effective date |
| Company filing | New filing or amendment |
| Labor/economic series | New release or revision |
| Customer interviews | New segment, role, or material counterevidence |
| Market-size model | Key input changes beyond tolerance |
| Research paper | Correction, retraction, or relevant new study |
| Recommendation | Decision assumptions or constraints change |
Every accepted artifact should state:
- as-of date;
- monitored sources;
- next scheduled review;
- event-based triggers;
- owner;
- supersession rule; and
- downstream consumers.
Separate durable knowledge from time-bound evidence
Durable:
- research protocol;
- definitions;
- source hierarchy;
- rubrics;
- calculation method;
- approved templates.
Time-bound:
- current prices;
- company features;
- market figures;
- current filings;
- active regulations;
- current sentiment;
- current recommendation.
The memory privacy and retention guide helps decide which source content, interview data, extracts, judgments, and artifacts may be retained, for how long, and for which later tasks.
Reuse without citation drift
When a later report reuses a claim:
- retrieve the claim record;
- confirm the source still exists and is current enough;
- confirm the later question uses the same definition;
- inspect corrections or contradictions;
- recalculate if inputs changed;
- cite the original source, not only the prior summary;
- record the new use; and
- assign a new review decision.
A prior accepted sentence is not automatically valid in a new context.
Handle corrections, retractions, and changed pages
Sources change after acceptance.
| Change | Research response |
|---|---|
| Typographical correction | Confirm whether claim or calculation changes |
| Material correction | Reopen affected claims and downstream artifacts |
| Retraction/withdrawal | Remove as support; reassess conclusion |
| Company page update | Compare snapshot and current claim |
| Filing amendment | Use amended filing and record supersession |
| Dataset revision | Recalculate affected measures |
| URL disappears | Find canonical archive/version; do not cite a copy blindly |
| Access becomes restricted | Preserve permitted metadata and reassess reproducibility |
Maintain a dependency index:
Source > evidence records > claims > calculations > findings > recommendation > downstream artifacts
When source S-021 changes, the owner should be able to identify every affected claim and decision artifact without searching prose manually.
For material changes:
- freeze the prior accepted artifact;
- classify the source event;
- reopen dependent claims;
- recalculate models;
- re-review the synthesis;
- notify downstream owners;
- issue a correction or retraction where required;
- preserve the event record; and
- update the evaluation set.
§ 10What Does a Complete Workflow Example Look Like?
Consider a fictional research request:
Compare 6 named AI workflow vendors for a U.S. operations team. Focus on role persistence, schedules, task state, memory controls, approvals, API/UI execution, audit evidence, pricing structure, and pilot fit as of July 29, 2026.
Protocol
| Element | Definition |
|---|---|
| Entities | 6 named vendors |
| Geography | U.S. buyer view |
| Date | Evidence current through July 29, 2026 |
| Required source | Current vendor docs, pricing, terms, security/trust where public |
| Secondary source | Independent technical research only for category context |
| Exclusion | Affiliate lists, snippets, copied directories |
| Unit | One claimed capability under one documented condition |
| Acceptance | Every comparison cell has evidence, unknown, or not applicable |
| Reviewer | Operations owner, security owner, procurement |
Workflow counts
| Stage | Count |
|---|---|
| Search/query runs | 64 |
| Results logged | 1,420 |
| Unique candidate sources | 238 |
| Excluded at identity/relevance pass | 146 |
| Full sources inspected | 92 |
| Included sources | 54 |
| Extracted evidence records | 317 |
| Proposed material claims | 126 |
| Accepted without qualification | 78 |
| Accepted with qualification | 34 |
| Rejected | 14 |
| Material contradictions preserved | 11 |
All counts are illustrative.
Evidence matrix
| Dimension | Vendor A | Vendor B | Vendor C |
|---|---|---|---|
| Scheduled work | Documented, conditions linked | Unknown | Documented, plan-limited |
| Task state | Documented | Documented | Demonstrated, docs incomplete |
| Memory control | Retention docs linked | User-level claim only | Admin controls linked |
| Approval | Per-action docs linked | Manual review workflow | Unknown |
| Audit evidence | Event docs linked | Activity history | Scope unclear |
| Pricing | Subscription + usage | Outcome unit | Contract quote required |
“Unknown” is a result. It is better than inferring capability from a marketing phrase.
Decision output
The report should not simply rank 1-6. It should state:
- which vendors meet mandatory gates;
- which evidence remains unverified;
- which 3 pilot tasks each option can be tested on;
- which permissions and data are required;
- which commercial units must be normalized;
- which risks require contract clarification;
- which option is recommended by the decision owner;
- what counterevidence would change the recommendation; and
- the as-of date.
The content-operations workflow can reuse the accepted source pack to create audience-specific assets. It should not alter the research conclusion or detach claims from their sources.
§ 11How Should You Pilot and Measure the Workflow?
Pilot one decision question before automating continuous market monitoring.
Four-week pilot
| Week | Focus | Exit condition |
|---|---|---|
| 1 | Charter, definitions, protocol, source registry | Decision and acceptance contract approved |
| 2 | Retrieval, screening, extraction | Logs reconcile and required sources covered |
| 3 | Claim/contradiction ledger and draft | Reviewer can audit every material statement |
| 4 | Review, decision use, refresh design | Owner accepts, narrows, changes, or stops |
Metrics
| Dimension | Measure |
|---|---|
| Coverage | Required entities/dimensions with accepted evidence |
| Traceability | Material claims linked to inspected sources |
| Source quality | Claims supported by required source tier |
| Citation fit | Citations supporting the exact sentence |
| Contradiction recall | Known material contradictions surfaced |
| Calculation accuracy | Reproduced calculations passing review |
| Definition consistency | Claims using approved category/segment definitions |
| Review burden | Review and correction minutes per accepted claim |
| Timeliness | Accepted artifact delivered by decision deadline |
| Decision utility | Decision owner records use, defer, or reject |
| Refresh | Triggered claims updated within policy |
| Severity | Worst unsupported or misapplied claim |
Fictional pilot result
| Result | Count/rate |
|---|---|
| Material claims reviewed | 126 |
| Valid source links | 124 of 126, 98.4% |
| Exact citation fit | 117 of 126, 92.9% |
| Accepted without qualification | 78 of 126, 61.9% |
| Accepted with qualification | 34 of 126, 27.0% |
| Rejected | 14 of 126, 11.1% |
| Known material contradictions surfaced | 10 of 11, 90.9% |
| Calculations reproduced | 18 of 18, 100% |
| Review/correction time | 760 minutes |
| Review minutes per accepted claim | 6.8 |
| Severe unsupported claims | 0 |
The missed contradiction matters more than a high source-link rate. Diagnose whether the query plan, source selection, extraction, or reviewer rubric failed.
Launch checklist
- One decision and named owner
- Written scope and definitions
- Source hierarchy and required classes
- Reproducible query and retrieval logs
- Inclusion and exclusion codes
- Claim and contradiction ledgers
- Calculation workbook or reproducible method
- Rights, privacy, and retention rules
- Human domain and decision review
- As-of date and refresh triggers
- Evaluation set with known traps
- Stop rule for severe unsupported claims
For a CellCog deployment, verify the current configuration of the AI Research Assistant and the scope and date of any result shown on the CellCog benchmarks page. A benchmark measures a stated task, environment, model, harness, date, and scoring rule. It does not establish performance for a different buyer’s market-research question.
Run one question with a source ledger and contradiction ledger. If the reviewer cannot reconstruct the material claims, do not scale the workflow.
Q1Can AI perform market research without human review?
It can collect and organize evidence, but consequential research should retain human review for definitions, source sufficiency, calculations, uncertainty, alternatives, and recommendation. Automatic collection is not automatic authority.
Q2How many sources should a market research report include?
There is no universal count. Coverage depends on the question, entities, dimensions, and evidence types. Define required source classes and saturation or stopping rules. Ten strong sources can answer a narrow question; one hundred weak sources may not support one load-bearing claim.
Q3Should search snippets be cited?
No. Use snippets for discovery. Open the canonical source, verify identity and date, inspect the supporting passage or data, and cite the direct destination. If the underlying source is unavailable, label the claim unverified or exclude it.
Q4How should conflicting market-size estimates be handled?
Compare category boundaries, populations, geographies, periods, currency, inflation treatment, and methods. Normalize only compatible measures. Otherwise report the estimates separately, explain the difference, and use a transparent model or range for the decision.
Q5Can a prior research report be reused?
Yes, but revalidate the source, date, definition, correction status, and new decision context. Cite the original evidence and record the new review. An accepted historical report is not a permanently current source.
Q6What is the best first research pilot?
Choose one bounded comparison with 3-6 entities, 5-8 decision dimensions, primary-source requirements, one domain reviewer, and a 2-4 week decision deadline. Require a source ledger, contradiction ledger, and claim-level acceptance.
