Anthropic released Claude Fable 5.1 on September 1, 2026. By that afternoon, CellCog’s Agent Max and Team Max tiers were running it. If you run Grok Bot, nothing changed for you today, and according to xAI’s own documentation nothing you do can change it: there is no model picker, and none is planned.
This is not a post about Grok Bot being bad. Grok 4.6 is a serious coding model that reached GitHub Copilot, Amazon Bedrock and Microsoft Foundry within two weeks of its release, and Grok Bot’s product design is genuinely good, as we said in our explainer. It is a post about a structural fact the always-on agent category has not talked about enough: when the company selling you the agent is also the company training the model, the model decision is theirs, permanently. That is the whole argument for buying the agent from the application layer.
On this page · 7 sectionsOpen
- Anthropic released Claude Fable 5.1 on September 1, 2026. CellCog’s Agent Max and Team Max tiers were running it the same afternoon. Grok Bot users got no new model and, per xAI’s documentation, will not get to choose one.
- xAI’s own docs, verbatim: ‘Grok Bot has no model picker, for members or admins. We do not plan to allow admin or user choice for models that are used with Grok Bot.’ Each request routes to a fixed set of models that xAI selects and does not name.
- People are asking for a choice anyway: an August 26 post claiming Grok Bot had been cracked open into a shell for Claude, Codex, Cursor and OpenRouter drew 38.8K views. It is an unverified user claim; the demand it signals is the point.
- The only exact-version head-to-head we found is CursorBench 3.2.0: Fable 5.1 at 73.4% versus Grok 4.6 at 70.8% (BenchLM public snapshot, September 1, 2026). Other headline numbers come from different benchmark versions and cannot be compared.
- The lock-in is structural, not a quality complaint: when the model vendor and the product are the same company, the model roadmap is theirs. An application layer swaps the engine under your memory, roles and permissions without your setup changing.
- The same rule binds application-layer vendors: OpenAI pulled its models from Cursor on August 28. Multi-vendor routing is the protection, and it is a question worth asking any always-on agent product, including us.
- Does Grok Bot run Fable 5.1?
- No. Grok Bot has no model picker and xAI says it does not plan one; requests route to a fixed set of models xAI chooses. Fable 5.1 is an Anthropic model. As of September 1, 2026 there is no path to run it, or any non-xAI model, under Grok Bot.
- Which model does Grok Bot use?
- xAI does not say. Its docs describe ‘a fixed set of models for its surface, with automatic failover.’ xAI’s newest model is Grok 4.6, released August 12, 2026, a day after Grok Bot’s beta, but xAI never states that Bot runs it.
- Is Fable 5.1 better than Grok 4.6?
- On the one published exact-version comparison, yes by a small margin: CursorBench 3.2.0 puts Fable 5.1 at 73.4% and Grok 4.6 at 70.8%. Everything else is vendor numbers on different benchmark versions. We are not claiming a measured quality change in our own product yet; Fable 5.1 is hours old.
- What does 'application layer' mean here?
- The product that holds your memory, roles, tasks, permissions and org structure, sitting above whichever model does the reasoning. When the two are separate, the engine can change without your setup changing. When they are one company, you get the model they ship, when they ship it.
§ 01What xAI’s documentation says, verbatim
The relevant section of xAI’s Grok Bot documentation for teams and enterprises reads, in full (retrieved September 1, 2026, from docs.x.ai):
Grok Bot has no model picker, for members or admins. We do not plan to allow admin or user choice for models that are used with Grok Bot. Model choice is fully managed by the product.
Each request routes to a fixed set of models for its surface, with automatic failover.
Three things follow. First, xAI never names the model behind Bot; the honest description is “whatever xAI routes for that surface.” xAI’s newest model is Grok 4.6, released August 12, one day after Bot’s August 11 beta, so the natural assumption is that Bot benefits from it, but the docs do not say so and we will not either. Second, “we do not plan to allow” is a roadmap statement, not a beta limitation. Third, the fixed set is xAI’s set. Fable 5.1 is an Anthropic model, and there is no mechanism for it, or any non-xAI model, to reach a Grok Bot.
§ 02The receipts that people want a choice anyway
Structural arguments are cheap. Here is what people said and did in the weeks after Grok 4.6 shipped.
- August 25. xAI’s official @grok account clarified that its Acceptable Use Policy, effective August 14, restricts using the service or its outputs to develop competing machine-learning models or products; non-competing adaptations are allowed, and users own their outputs. Reasonable terms, and a clear line: the outputs are yours, the engine is not.
- August 26. A post from @adiix_official claimed Grok Bot had been “cracked open” fifteen days after launch into a shell for Claude, Codex, Cursor and OpenRouter. It reached 38.8K views. We have not verified the claim and are not endorsing the method. We are noting that a post about putting other models under Grok Bot is what travelled.
- August 28 to 30. r/grok threads on the Grok 4.6 experience. These are unverified individual user reports, quoted as posted: one user called 4.6 “significantly worse than 4.5” for writing (August 29); another wrote that “memory is poor, hallucinations are rampant, and the filters are excessively strict” (August 29); a third complained of a “ridiculous weekly cap on text messages” while praising Grok Build (August 28). None of these say what model Bot uses, and none are measurements. They are what a fixed engine sounds like when some of its users want a different one.
- August 25. The SuperGrok Heavy plan’s marketing changed from “near-unlimited usage” to “highest usage limits” at the same $300 price, per AI Tools Recap’s roundup. The plan is fine. The point is that a bundled model-plus-product changes on the vendor’s schedule, all at once, and the pricing moves with it.
§ 03What a model change looks like from the application layer
Fable 5.1 went generally available on the morning of September 1 (US time). The same afternoon, CellCog’s Agent Max moved from Fable 5 to Fable 5.1 and Team Max moved from the Claude Opus line to Fable 5.1. The full record is the day-one product update; the specs, prices and our rumor scorecard are on the Fable 5.1 tracker.
Two honest caveats. We are not claiming a measured quality improvement in CellCog’s output today. Fable 5.1 is hours old, we treat vendor benchmarks as vendor benchmarks, and our own read will land on the tracker over the coming weeks. And we are not claiming heroics: switching took an afternoon because Fable 5.1 is a clean superset of Fable 5, and because nothing a CellCog user cares about lives inside the model.
That second clause is the argument. What a user owns on CellCog is the employee: its memory that carries from one shift to the next, its role, its task board, its permissions and approval rails on every command that reaches your world (the terminal on your machine, your real browser, your connected tools), and the org structure around it. None of that is model-specific. So when a better engine ships, we can put it under the work the same day, and your pricing does not move: a full shift of real work still runs about $25, and the cost depends purely on how much work you assign.
§ 04The honest benchmark picture
The claim you will hear is “Fable 5.1 is significantly ahead of Grok 4.6.” The receipt is narrower. We looked for an exact-version, same-harness comparison of the two and found one: CursorBench 3.2.0, where BenchLM’s public snapshot (September 1, 2026) lists Claude Fable 5.1 at 73.4% and Grok 4.6 at 70.8%, with Fable 5 third at 70.5%. Anthropic’s own release page reports the same 73.4% for Fable 5.1. That is a real lead and a modest one.
What you will see elsewhere is not comparable. Anthropic reports Fable 5.1 at 55.8% on Terminal-Bench 4.0; Artificial Analysis lists Grok 4.6 at 88.4% on Terminal-Bench 2.1. Different versions, different task sets. Putting those two numbers in one sentence is how misleading comparison posts get written, and neither xAI nor Anthropic has published a head-to-head.
Our own receipt is the one we can stand behind: CellCog Max ranked #1 on Deep Research Bench (July 2026). We can stand behind it because it was earned at the application layer: the routing, the multi-agent debate, the memory, not one vendor’s model. When the model underneath improves, that number has a chance to improve with it. When a vendor stalls, we route elsewhere.
§ 05The lock-in nobody advertises
Vendor lock-in in software usually means your data is stuck. This is a different kind: your judgment is stuck. With a bundled agent you accept the vendor’s model, its rate limits, its content policy, its acceptable-use terms and its release calendar as one package, and you accept the next package too. If Grok 4.6’s successor is better for your work, good. If it is better for someone else’s work, you get it anyway.
The application layer does not escape vendors; it spreads the bet. And the rule binds us as much as anyone. The same week Fable 5.1 shipped, OpenAI announced it would stop supplying models to Cursor after Cursor’s acquisition by SpaceX, naming Astra as a model Cursor will never receive. An application-layer company with one vendor is a bundled product with extra steps. CellCog routes across Anthropic, OpenAI and Google models by mode and tier for exactly that reason: the Flash tiers run on Gemini 3.7 Flash, Agent Creative on Claude Opus 5, the Max tiers now on Fable 5.1. The day-one switch is what that looks like when it pays off; the Cursor story is what it looks like when it does not.
§ 06Five questions to ask any always-on agent product
- Which model runs my work, and where is that written down?
- Can the model change without my setup changing, and who decides?
- When a frontier model ships, what is your record on adopting it, measured in days?
- Which parts of the product are mine (memory, roles, permissions, org structure), and which belong to the model?
- What happens to my agents if you and the model vendor part ways?
Grok Bot answers the first two in its documentation: xAI’s set, and no. The other three are worth asking everyone, including us. Our answers are in the comparison, and if you are already running Bots, the switching guide covers what carries over.
§ 07What we are watching
- Any change to xAI’s no-picker policy, or any public naming of the models behind Bot. Either would change this post, and we will say so here.
- Independent replication of Fable 5.1’s published numbers, and any same-version comparison with Grok 4.6 beyond CursorBench.
- Whether the r/grok reports about Grok 4.6 turn into measurements.
Updates land here and on the tracker, same day.
Q1What exactly do xAI's docs say about model choice in Grok Bot?
The teams-and-enterprises page of xAI’s Grok Bot documentation, retrieved September 1, 2026, says: ‘Grok Bot has no model picker, for members or admins. We do not plan to allow admin or user choice for models that are used with Grok Bot. Model choice is fully managed by the product. Each request routes to a fixed set of models for its surface, with automatic failover.’ It does not name the models.
Q2Did Grok Bot get worse when Grok 4.6 shipped?
We do not know, and neither does anyone outside xAI, because xAI does not say which model Bot runs. What exists is a cluster of unverified user reports on r/grok between August 28 and 30 about Grok 4.6 in the app: ‘significantly worse than 4.5’ for writing, ‘memory is poor, hallucinations are rampant’, a ‘ridiculous weekly cap on text messages’. They are individual reports, not measurements, and they are about Grok, not Bot specifically. Grok 4.6 is also a widely adopted coding model: it reached GitHub Copilot on August 19 and Microsoft Foundry on August 26.
Q3What changed at CellCog on September 1?
Agent Max moved from Fable 5 to Fable 5.1 and Team Max moved from the Claude Opus line to Fable 5.1, the same afternoon Anthropic released it. Pricing did not change: a full shift of real work still runs about $25, and the cost depends purely on how much work you assign. The full record is the day-one product update.
Q4Doesn't every AI product depend on a model vendor?
Yes, and that is why the protection is multi-vendor routing rather than independence. OpenAI announced on August 28 that it would stop supplying models to Cursor after Cursor’s acquisition by SpaceX. An application-layer product with one vendor is a bundled product with extra steps. CellCog routes across Anthropic, OpenAI and Google models by mode and tier so that any one vendor’s decision changes an engine, not the product.
Q5What should I ask an always-on agent vendor about models?
Five questions: which model runs my work and where is that written down; can it change without my setup changing; what is your record on adopting a new frontier model, in days; which parts of the product are mine (memory, roles, permissions, org) versus the model’s; and what happens to my agents if you and the vendor part ways.
