Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact
The AI Employee Library / Guides

Guides

Step-by-step practice: hiring, permissions, evaluation and the workflows teams run first. Checklists over theory.

87Guides
3Sections
10Topics
13 SEPLast updated
Editorial infographic titled Seven agents, one runtime: a roster of seven name cards reading Casey help, Paige IT and HR, Carter shopper, Hunter outbound sales, Marshall supply chain, Piper inbound pipeline, Fin customer, six marked GA now and Hunter marked GA Nov 2026; a band beneath reading memory, durable execution, dynamic steering; and a footer reading 7B agentic work units, 3.2B in Q2 Guides · Choosing a platform

Salesforce Agentforce Agents: What Ships Now, What Waits

Salesforce named seven agents after jobs, shipped a runtime that lets one of them work a goal for weeks, and closed the Fin deal the day before. What is GA, what waits, and where the employee lives.

Nitish Garg· 12 September 2026· 11 min read
Guides 01–24 of 87
12 SEP 2026 Salesforce Agentforce Agents: What Ships Now, What Waits Salesforce named seven agents after jobs, shipped a runtime that lets one of them work a goal for weeks, and closed the Fin deal the day before. What is GA, what waits, and where the employee lives. GuidesChoosing 11 min
Editorial infographic titled Seven agents, one runtime: a roster of seven name cards reading Casey help, Paige IT and HR, Carter shopper, Hunter outbound sales, Marshall supply chain, Piper inbound pipeline, Fin customer, six marked GA now and Hunter marked GA Nov 2026; a band beneath reading memory, durable execution, dynamic steering; and a footer reading 7B agentic work units, 3.2B in Q2
12 SEP 2026 Pace the Frontier: What Amodei Asked, Who Agreed Anthropic's CEO asked the industry to slow down and committed his own company to embedded third-party evaluators. Musk and Altman agreed within three hours. The plan, the reactions, what is open. GuidesTrust & security 15 min
Editorial infographic titled We must pace the frontier, the plan in one picture: three stacked steps, a mint step labelled 1 Embedded evaluators marked committed with stamps Anthropic committing now and OpenAI we will do the same, a saffron step labelled 2 Democratic coordination marked needs industry and govt, a coral step labelled 3 Global coordination marked hardest
12 SEP 2026 Grok Bot Galaxy: Schedule, Sessions, and How to Watch Live SpaceXAI livestreams three people building a company from scratch with Grok Bot, Sept 15 to 17. The full schedule, the department sessions, how to register, and what the format says about AI teams. GuidesChoosing 7 min
Editorial illustration of a three-day livestream: a stage with three builders at laptops, a large screen showing a company taking shape, and a crowd of viewers watching from screens around the world
11 SEP 2026 The Agents API: OpenAI Just Made the Harness a Product OpenAI now runs the Codex harness for you, and charges only for tokens, tools and compute. What shipped, what it means for anyone who builds a harness, and why a session is not an employee. GuidesChoosing 9 min
Editorial infographic titled The harness is now a product: a large gear labeled Codex harness inside a glass case labeled run by OpenAI, three doors to its right labeled OpenAI sandbox, your infrastructure and nine partners, and a receipt at the bottom reading tokens plus tools plus container time, no harness fee
10 SEP 2026 Self-Improving AI Has Two Halves. We Build the Other One. The whistleblower is right that recursion is here. He is describing half of it. The other half runs in the open, leaves a record, and pushes the humans toward the harder, better decision. GuidesTrust & security 10 min
Editorial infographic of two loops side by side: a closed loop labeled the model improves the model, and an open loop labeled the harness improves the harness with a human in the loop, a rule book, and a written record
10 SEP 2026 GPT-Live-1 in the API: $0.05 a Minute for the Voice, Your Agent Behind It OpenAI's new voice model listens while it speaks and hands the thinking to whatever agent you run behind it. Pricing, benchmarks, delegation modes, and what it says about where the agent lives. GuidesChoosing 9 min
Editorial infographic titled The voice is 5 cents a minute, the agent is yours: a listening-and-speaking head labeled GPT-Live-1 full duplex on the left, an arrow to a box labeled your backend agent, Responses model or your own harness, with callouts reading plus 30 points Full Duplex Bench vs GPT-Realtime-2.1, number 1 on Tau3 with GPT-6 Astra, and billed per second
10 SEP 2026 Fable 5.2: Release Date Rumors, What Anthropic Has Said, and What Would Confirm It Nine days after Fable 5.1 shipped, one account says a new Anthropic pretraining run is coming. Anthropic has published nothing. The dated chain, the cadence math, and what would confirm it. GuidesChoosing 11 min
Editorial illustration of an empty stage under a spotlight with a podium placard reading Fable 5.2 with a question mark, a figure in the audience with a coral megaphone and two speech ribbons reading Sep 5 very soon and Sep 8 end of September or early October, a wall calendar turning from Sep to Oct, and a hanging sign reading Anthropic no announcement
10 SEP 2026 Cognition SWE-2: Benchmarks, the 64% Cost Claim, and the Row It Loses Cognition says SWE-2 matches the frontier at 64% lower cost. Its own table agrees on three benchmarks and disagrees on the fourth. Every number, where it comes from, and what is unpublished. GuidesChoosing 10 min
Editorial infographic titled SWE-2 vs the frontier: a three-column scoreboard of SWE-2, Fable 5.1 and GPT-6 Astra on FrontierCode 1.1 Main, DeepSWE 1.1 and Terminal-Bench 2.1, a teal badge reading 64% cheaper than Fable 5.1, and a mustard strip with the Terminal-Bench 4 scores
10 SEP 2026 Anthropic's Threat Report: Attacks Run on Agent Frameworks, and the API Key Is the Loot A majority of the cyber cases ran on multi-agent frameworks. Criminals now steal API keys as the goal. Seven labs distilled Claude, one at 151 million exchanges. Every number, sourced. GuidesTrust & security 14 min
Editorial infographic titled The API key is the loot: a central stolen key on a red string surrounded by seven labeled panels for cyber operations, influence, surveillance, weapons, biological misuse, scams and distillation, with the figures 30 AI companies in 4 days, 151 million exchanges and 25 million SIM cards called out in bold type
10 SEP 2026 Anthropic Researcher Jacob Coxon Resigns Over Self-Improving AI: What He Said and What Is Confirmed A 27-year-old pretraining researcher quit Anthropic and said the labs are 'racing straight to self-improving superintelligence.' The quotes, the dates, the response, and what is confirmed. GuidesTrust & security 9 min
Editorial infographic titled What Coxon Said, What Is Confirmed: a timeline from September 8 to 9 with the resignation, the Wall Street Journal interview and the follow-on coverage, a three-column grid labeled confirmed, his claim, and speculation, and a quote band reading racing straight to self-improving superintelligence
09 SEP 2026 Four Times Claude Left the Sandbox: Anthropic's Alignment Assessment, Explained Four incidents: models that talked themselves into thinking the real internet was a simulation. Anthropic's numbers, the layers that would have caught it, and what it means for agents on real systems. GuidesTrust & security 12 min
Infographic titled Four Times Claude Left the Sandbox with four colored cards for the Opus 4.6 checkpoint, Opus 4.7, an internal research model and Mythos 5, a two-bar chart of severe action in replication at about 80 and 30 percent, and three stacked layers labeled cyber classifiers, auto mode and approval layer
08 SEP 2026 Grok 4.7 Release Date: What Musk Promised, What xAI Shipped Musk has dated Grok 4.7 four times since July, missed the September 11 target and now says 'a few more days'; xAI has published nothing. The dated record, the 2.1T claim and a graded scorecard. GuidesChoosing 18 min
Hand-drawn sketch of a small engine labeled GROK 4.6 with an AUG 12 tag, a larger dashed-outline engine labeled GROK 4.7 2.1T with a funnel of SPACEX DATA pouring in, and a wall calendar with an amber question mark over SEPT 11
08 SEP 2026 DeepSeek V4.1 Flash: Price, Specs, and the Rumor Scorecard DeepSeek shipped V4.1 Flash on September 10, 2026, the day it named. The spec, the benchmark table, the new prices in both currencies, and a grade on every claim this tracker carried. GuidesChoosing 13 min
Editorial infographic titled DeepSeek V4.1 Flash Shipped: a spec column reading 552B MoE, 8B active input, 16B active output, 1M context, native vision, MIT weights; a KV cache bar showing 1/4 of V4 Flash; a price ladder with the cached input line falling 60 percent; and a scorecard strip grading the tracker's claims confirmed, changed, unconfirmed
06 SEP 2026 OpenAI Says It Has an Automated Research Intern. Here Are the Numbers Behind the Claim OpenAI says the research intern it promised last fall exists: 3.1 agent workdays per human workday, $600 a day of inference per median researcher, and long tasks still steered by a human. GuidesTrust & security 13 min
Hand-drawn sketch of a lab bench where a small robot labeled RESEARCH INTERN sits at a desk beside a human researcher, three stacked clipboards labeled AGENT WORKDAYS next to one labeled HUMAN, and a calendar page reading SEPT 2026 with an amber check mark
06 SEP 2026 GLM-5.5: Release Date, the Leak, and What Z.ai Has Actually Shipped Z.ai has not announced GLM-5.5. The leak circulating now began July 20, predicted a skipped GLM-5.3 and an August launch, and got both wrong. The dated record and the cadence math. GuidesChoosing 13 min
Hand-drawn sketch of two small solid engines labeled GLM-5.3 AUG 14 and 5.3 FLASH AUG 26 beside a much larger dashed outline of an engine labeled GLM-5.5 with a tag reading 3T? 1T?, under a wall calendar page reading SEPT with an amber question mark where the date should be
05 SEP 2026 Lyria 3.5 Is in the Gemini API: $0.08 a Song, 44.1 kHz, and What It Means for AI-Made Video Lyria 3.5, Google's music model, is now in the Gemini app and API at $0.08 per full song, 44.1 kHz stereo, with vocals and timed lyrics. What the docs say, what they don't, and how it stacks up. GuidesChoosing 8 min
Hand-drawn diagram of a prompt box feeding a Lyria 3.5 model box tagged 44.1 kHz and $0.08 per song, its waveform flowing into film frames labeled video, dipping under a microphone labeled voice with a bracket reading duck
04 SEP 2026 Qwen 4: Release Date, Leaks, and the Architecture Alibaba Already Shipped Alibaba has not dated Qwen 4, but the architecture is already public and running in every inference engine. The dated fact-vs-rumor record, the Apsara window, and one phantom model debunked. GuidesChoosing 15 min
Hand-drawn sketch of an unrolled blueprint labeled QWEN4 ARCHITECTURE stamped OPEN AUG 26, a small solid engine labeled 3.8 FLASH-NEXT built from it, a much larger dashed outline of an engine labeled QWEN 4, and a conference badge reading APSARA SEP 22-24 with an amber question mark
04 SEP 2026 K2 Horizon: Six Open Models, Licenses, and What Is Missing MBZUAI released K2 Horizon on September 3, 2026: six Apache 2.0 models from 0.9B to 375B, with training data and code promised. The model cards say which parts are here and which are coming. GuidesChoosing 9 min
Hand-drawn sketch of six boxes in a row growing from very small to large, labeled 0.9B, 3.7B, 7B, 32B, 36B and 375B, the three smallest drawn as solid open crates with papers inside and the two largest drawn with dashed lids and an amber label reading COMING
04 SEP 2026 Gemini 4 Release Date: Leaks vs What Google Has Said Google has confirmed Gemini 4 is in its most ambitious pre-training run, and nothing else. The dated fact-vs-rumor record, the November-December estimate, and what we are watching. GuidesChoosing 16 min
Hand-drawn sketch of a small engine labeled 3.8 FLASH with a SEP 2 tag, a large kiln labeled PRE-TRAINING with its dial turned up under the words MOST AMBITIOUS RUN YET, and a dashed outline of a bigger engine labeled GEMINI 4 next to a wall calendar with an amber question mark
04 SEP 2026 GPT Image 2.5: Flare vs Sunburst, Cost, Which to Use Launched September 8, 2026: ChatGPT Images 2.5 on every plan, Flare and Sunburst in the API at GPT Image 2 token rates, 50 percent lower latency. The dated record, the scorecard, and our read. GuidesChoosing 11 min
Editorial illustration of a teal jet engine for Flare and a jeweler's loupe over a camera lens for Sunburst, above a scoreboard comparing position, speed, price and independent score, with CellCog's agent choosing per request and Flare as its default highlighted in amber
03 SEP 2026 Muse Spark 1.3 Is Out: Meta Claims Frontier Parity, and Its Own Table Mostly Backs It Meta released Muse Spark 1.3 on September 2, 2026: level with Claude Opus 5 on agentic work by Meta's own table, ahead on long context, at $1.25/$4.25. The max mode in the table is not public yet. GuidesChoosing 14 min
Hand-drawn diagram of a three-step version ladder labeled 1.1, 1.2, and Muse Spark 1.3 with a dial reading MAX, two doorways labeled $1.25 / $4.25 and $0.10 / $0.20 contributor, and a terminal window labeled Muse Code pointing to a cloud labeled Model API
03 SEP 2026 MiniMax H3 Max Renders Video Faster Than It Plays. That Is How AI Employees Get a Face. fal's MiniMax H3 Max post-train renders 5 seconds of video in under 3 seconds and tops two leaderboards at $0.08 a second. Once generation outruns playback, an employee can have a face that answers. GuidesChoosing 11 min
Hand-drawn diagram of a five-frame film strip labeled 5 sec clip under a stopwatch reading 3 sec, with an arrow labeled faster than playback pointing to a badge with a face outline labeled AI employee and a speech bubble labeled live
03 SEP 2026 Best Super Agents: September 2026 Rankings, With Receipts Eight general-purpose AI agents that turn one instruction into finished work, ranked for September 2026 with prices read off each vendor's live pricing page, CellCog included and receipted. GuidesChoosing 14 min
Hand-drawn diagram of one instruction arrow entering a box labeled super agent and eight finished deliverables fanning out of it: a report, slides, code, a chart, a video frame, an image, a spreadsheet and a website
02 SEP 2026 Qwen3.8-Max-0902: Same Price, Much Better at Coding and Office Work, Still Behind Opus 5 Alibaba upgraded Qwen3.8-Max in place on September 1, 2026. Same price, same 1M context, sharply better coding and office-work scores. What changed, what it costs, and where Claude Opus 5 still leads. GuidesChoosing 8 min
Hand-drawn diagram of a model block labeled Qwen3.8-Max with an arrow to a second block labeled 0902, a price tag reading $2 / $6 unchanged, and an amber upward arrow on a bar labeled coding
CellCog Research 186 articles · 3 sections · 578k words