Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact
The AI Employee Library

Field notes on standing AI workers

Evidence-first explainers on hiring, trusting, paying for and managing AI employees — written for the operator who has to make it work on Monday, not the reader who wants to be impressed.

188Articles
3Sections
10Topics
14 SEPLast updated
All articles — page 2 25–48 of 188
04 SEP 2026 K2 Horizon: Six Open Models, Licenses, and What Is Missing MBZUAI released K2 Horizon on September 3, 2026: six Apache 2.0 models from 0.9B to 375B, with training data and code promised. The model cards say which parts are here and which are coming. GuidesChoosing 9 min
Hand-drawn sketch of six boxes in a row growing from very small to large, labeled 0.9B, 3.7B, 7B, 32B, 36B and 375B, the three smallest drawn as solid open crates with papers inside and the two largest drawn with dashed lids and an amber label reading COMING
04 SEP 2026 Gemini 4 Release Date: Leaks vs What Google Has Said Google has confirmed Gemini 4 is in its most ambitious pre-training run, and nothing else. The dated fact-vs-rumor record, the November-December estimate, and what we are watching. GuidesChoosing 16 min
Hand-drawn sketch of a small engine labeled 3.8 FLASH with a SEP 2 tag, a large kiln labeled PRE-TRAINING with its dial turned up under the words MOST AMBITIOUS RUN YET, and a dashed outline of a bigger engine labeled GEMINI 4 next to a wall calendar with an amber question mark
04 SEP 2026 GPT Image 2.5: Flare vs Sunburst, Cost, Which to Use Launched September 8, 2026: ChatGPT Images 2.5 on every plan, Flare and Sunburst in the API at GPT Image 2 token rates, 50 percent lower latency. The dated record, the scorecard, and our read. GuidesChoosing 11 min
Editorial illustration of a teal jet engine for Flare and a jeweler's loupe over a camera lens for Sunburst, above a scoreboard comparing position, speed, price and independent score, with CellCog's agent choosing per request and Flare as its default highlighted in amber
03 SEP 2026 Muse Spark 1.3 Is Out: Meta Claims Frontier Parity, and Its Own Table Mostly Backs It Meta released Muse Spark 1.3 on September 2, 2026: level with Claude Opus 5 on agentic work by Meta's own table, ahead on long context, at $1.25/$4.25. The max mode in the table is not public yet. GuidesChoosing 14 min
Hand-drawn diagram of a three-step version ladder labeled 1.1, 1.2, and Muse Spark 1.3 with a dial reading MAX, two doorways labeled $1.25 / $4.25 and $0.10 / $0.20 contributor, and a terminal window labeled Muse Code pointing to a cloud labeled Model API
03 SEP 2026 MiniMax H3 Max Renders Video Faster Than It Plays. That Is How AI Employees Get a Face. fal's MiniMax H3 Max post-train renders 5 seconds of video in under 3 seconds and tops two leaderboards at $0.08 a second. Once generation outruns playback, an employee can have a face that answers. GuidesChoosing 11 min
Hand-drawn diagram of a five-frame film strip labeled 5 sec clip under a stopwatch reading 3 sec, with an arrow labeled faster than playback pointing to a badge with a face outline labeled AI employee and a speech bubble labeled live
03 SEP 2026 Best Super Agents: September 2026 Rankings, With Receipts Eight general-purpose AI agents that turn one instruction into finished work, ranked for September 2026 with prices read off each vendor's live pricing page, CellCog included and receipted. GuidesChoosing 14 min
Hand-drawn diagram of one instruction arrow entering a box labeled super agent and eight finished deliverables fanning out of it: a report, slides, code, a chart, a video frame, an image, a spreadsheet and a website
02 SEP 2026 Qwen3.8-Max-0902: Same Price, Much Better at Coding and Office Work, Still Behind Opus 5 Alibaba upgraded Qwen3.8-Max in place on September 1, 2026. Same price, same 1M context, sharply better coding and office-work scores. What changed, what it costs, and where Claude Opus 5 still leads. GuidesChoosing 8 min
Hand-drawn diagram of a model block labeled Qwen3.8-Max with an arrow to a second block labeled 0902, a price tag reading $2 / $6 unchanged, and an amber upward arrow on a bar labeled coding
02 SEP 2026 Gemini 3.8 Flash vs 3.7: Benchmarks, Price, Cyber Variant Google released Gemini 3.8 Flash on September 2, 2026: same $0.75/$3.75 intro price as 3.7 Flash, an 8-point DeepSWE jump to within 0.3 of Claude Opus 5, and a cyber variant for defenders. GuidesChoosing 16 min
Hand-drawn diagram of a small fast engine labeled 3.8 Flash with an effort dial reading low, medium, high, a price tag reading $0.75 and $3.75 beside a calendar page reading Dec 31, and a small shield labeled Cyber
02 SEP 2026 Flash Tiers Now Run on Gemini 3.8 Flash, the Day Google Shipped It Google released Gemini 3.8 Flash on September 2, 2026. The same day, CellCog's Flash tiers moved to it: same setting, same pricing, better agent scores. Nothing to change on your side. Product UpdatesChangelog 7 min
Hand-drawn diagram of two boxes labeled Agent Flash and Team Flash with arrows converging into an engine labeled 3.8 Flash, a ghosted dashed engine labeled 3.7 Flash behind it, a dial at medium, a tag reading same price, and an amber stamp reading Day One
01 SEP 2026 Grok Bot Can't Run Fable 5.1. That's the Whole Argument for the Application Layer Anthropic shipped Fable 5.1 on September 1, 2026, and CellCog's Agent Max and Team Max ran it that afternoon. Grok Bot has no model picker, by design. That is the argument for the application layer. GuidesChoosing 9 min
Hand-drawn diagram of a locked box labeled GROK BOT with one fixed gear inside, next to an open box labeled APPLICATION LAYER with a socket receiving an interchangeable cartridge labeled FABLE 5.1
01 SEP 2026 GPT-6 Astra vs Claude Fable 5.1: Same Price, Different Cache Rate, and OpenAI's Head-to-Head Table OpenAI launched GPT-6 Astra at Claude Fable 5.1's price, $10 and $50 per million tokens, then published a head-to-head table. What differs: cached input at 4x, access, and who ran the tests. GuidesChoosing 17 min
Hand-drawn diagram of a shipping crate labeled FABLE 5.1 with a SHIPPED stamp and a price tag reading $10 / $50, beside a telescope pointed at a star labeled ASTRA with a calendar reading SOON and a shield labeled CRITICAL
01 SEP 2026 DeepSeek Opens V4-Flash-Vision-Exp: MIT Weights, Real Specs, and the Multimodal Agent Bet DeepSeek's first multimodal V4 model went open-weight under MIT on August 31: built on V4-Flash, tuned for multimodal agent work, and benchmarked within reach of Opus-4.8 on its own card. GuidesChoosing 7 min
Hand-drawn diagram of a box labeled V4-FLASH with an eye symbol attached labeled VISION, an arrow to a download tray labeled OPEN WEIGHTS, and an amber tag reading MIT
01 SEP 2026 Agent and Team Tiers Now Run on Fable 5.1, the Day Anthropic Shipped It Anthropic released Claude Fable 5.1 on September 1, 2026. The same day, CellCog's Agent and Team tiers, Core and Max, moved to it. Creative and Flash are unchanged. Nothing to change on your side. Product UpdatesChangelog 7 min
Hand-drawn diagram of four boxes labeled Agent Core, Agent Max, Team Core, and Team Max with arrows converging into one engine box labeled Fable 5.1, two smaller boxes labeled Creative and Flash marked unchanged, and a small stamp reading Day One
31 AUG 2026 The Most Dangerous Species Already Exists Everyone measures AI against imaginary perfection. Measured against the only general intelligence with a track record, the first agent society's mistake looks small, legible, and fully auditable. GuidesTrust & security 8 min
Hand-drawn teal diagram comparing mammal biomass, a bar showing humans plus livestock at 96 percent with an amber circle around the 4 percent wild-mammal sliver, versus a message board window labeled 70,000 messages with a magnifying glass labeled fully auditable
31 AUG 2026 OpenClaw 2.0 Is Here: What's New, What Breaks, and What It Signals OpenClaw's largest release ever: 16,000+ pull requests from 933 contributors. The features that matter, the three migrations to plan for, and what the direction says about where agents are heading. GuidesChoosing 19 min
Hand-drawn sketch of a small box labeled 1.X with an arrow into a larger structure labeled 2.0, with branches labeled SETUP, BROWSER, MULTIPLAYER, MEMORY, and SECURITY
31 AUG 2026 Meta's Project OT: The AI-Native Restructuring That Imploded, Explained Meta planned an AI-native company: pods of builders supervising agents, some teams shrunk up to 60 percent. Then incidents rose and the second wave was cancelled. The evidence ledger and the lesson. InsightsMulti-agent 7 min
Hand-drawn teal diagram of Meta's AI-native plan, a builder figure above a row of agent glyphs, a broken arrow circled in amber, and a rising incidents line labeled plus 40 percent with a clipboard labeled plan shelved
31 AUG 2026 Managing AI Agents: The Real Hours, From the Best Public Data The best public numbers say managing an AI agent takes 3 to 4 hours per week, about what a human rep takes. That overhead is real, and it is mostly a bill for structure the agent does not have. InsightsCost & ROI 7 min
Hand-drawn teal diagram with a clock labeled 3-4 hours per agent per week circled in amber over a row of robot glyphs and a human with a checklist, an arrow labeled the fix, and a stack of layers labeled memory, task board, handovers
31 AUG 2026 Connect Your Own MCP Servers: Bring Any Tool to Your CellCog Agents Paste an MCP server's URL and your agents can discover and call its tools - internal company tools included - behind the same approvals as everything else. Product UpdatesChangelog 5 min
Hand-drawn diagram of an MCP server box plugging into a socket, its tools flowing into an AI agent's tool search card
31 AUG 2026 Cloud Browsers: Every AI Employee Now Has Its Own Browser Your AI employees can now browse as themselves: a real Chrome of their own that keeps logins and tabs between shifts, a live view you can watch and drive, and a one-click login handoff. Product UpdatesChangelog 6 min
Hand-drawn teal diagram of a browser window inside a cloud labeled CLOUD BROWSER, linked to an AI employee glyph below and to a LIVE VIEW monitor watched by a stick figure labeled YOU, above a shift timeline, with an amber circle around the address-bar padlock labeled STAYS SIGNED IN
30 AUG 2026 Claude Opus 5.1 and Sonnet 5.1: Release Date, Leaks, and What We Actually Know Opus 5.1 and Sonnet 5.1 are community labels on two leaked test strings, not announced models. The dated fact-vs-rumor record of the marshmallow and melon leak. InsightsMulti-agent 8 min
Hand-drawn sketch of a solid box labeled Opus 5 next to two dashed boxes labeled Opus 5.1 and Sonnet 5.1 with amber question marks, and doodles of a marshmallow and a melon slice pointing at the dashed boxes with uncertain dotted arrows
29 AUG 2026 OpenAI Cut Off Cursor: What Still Works, What to Use OpenAI gave maximum contract notice: model supply to Cursor ends November 12, and Astra will not arrive at all. The verified record, Cursor's 5 percent answer, and the routes that still work. GuidesChoosing 9 min
Hand-drawn sketch of a power plug labeled OpenAI pulled from a socket labeled Cursor, a calendar page reading Nov 12, and three signpost arrows labeled API key, Codex extension, and gateway
29 AUG 2026 GPT-6 Astra: Access, Price, and the $200 Pro Pause OpenAI launched GPT-6 Astra on September 3, 2026 and published its benchmark table a day later. The dated record: price, specs, rollout, OpenAI's numbers against Fable 5.1, and our rumor scorecard. GuidesChoosing 30 min
Hand-drawn sketch of a telescope pointed at a large star labeled ASTRA, a calendar page with a question mark, and a shield labeled CRITICAL between the star and a row of small buildings
28 AUG 2026 What Is Tencent Hy4? 770B Open Model, Specs and Pricing Tencent open-sourced Hy4 preview on August 28: 770B total, 49B active, a 1M-token context, Apache 2.0 weights. The specs, the pricing, and the launch signal that matters most. GuidesChoosing 6 min
Hand-drawn sketch of a large box labeled 770B containing a small amber box labeled 49B active, with arrows pointing to code, office, and research icons, and an open padlock labeled Apache 2.0
28 AUG 2026 What Is Claudeforce? Salesforce's Claude Agent, Explained A source-grounded explainer of Claudeforce: what Salesforce and Anthropic actually announced on August 26, what ships now versus September, and what the deal validates about agents at work. GuidesChoosing 7 min
Hand-drawn sketch of a box labeled Claude and a cloud labeled Salesforce joined by a plug labeled Claudeforce, with skill cards flowing to a seller at a desk
CellCog Research 188 articles · 3 sections · 585k words