Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact
Cluster 07 of 10

Trust, permissions & security

Approvals, least privilege, guardrails, audit logs and incident response — how to grant real authority without losing control of it.

Trust, permissions & security 21 articles
12 SEP 2026 Pace the Frontier: What Amodei Asked, Who Agreed Anthropic's CEO asked the industry to slow down and committed his own company to embedded third-party evaluators. Musk and Altman agreed within three hours. The plan, the reactions, what is open. GuidesTrust & security 13 min
Editorial infographic titled We must pace the frontier, the plan in one picture: three stacked steps, a mint step labelled 1 Embedded evaluators marked committed with stamps Anthropic committing now and OpenAI we will do the same, a saffron step labelled 2 Democratic coordination marked needs industry and govt, a coral step labelled 3 Global coordination marked hardest
10 SEP 2026 Self-Improving AI Has Two Halves. We Build the Other One. The whistleblower is right that recursion is here. He is describing half of it. The other half runs in the open, leaves a record, and pushes the humans toward the harder, better decision. GuidesTrust & security 10 min
Editorial infographic of two loops side by side: a closed loop labeled the model improves the model, and an open loop labeled the harness improves the harness with a human in the loop, a rule book, and a written record
10 SEP 2026 Anthropic's Threat Report: Attacks Run on Agent Frameworks, and the API Key Is the Loot A majority of the cyber cases ran on multi-agent frameworks. Criminals now steal API keys as the goal. Seven labs distilled Claude, one at 151 million exchanges. Every number, sourced. GuidesTrust & security 14 min
Editorial infographic titled The API key is the loot: a central stolen key on a red string surrounded by seven labeled panels for cyber operations, influence, surveillance, weapons, biological misuse, scams and distillation, with the figures 30 AI companies in 4 days, 151 million exchanges and 25 million SIM cards called out in bold type
10 SEP 2026 Anthropic Researcher Jacob Coxon Resigns Over Self-Improving AI: What He Said and What Is Confirmed A 27-year-old pretraining researcher quit Anthropic and said the labs are 'racing straight to self-improving superintelligence.' The quotes, the dates, the response, and what is confirmed. GuidesTrust & security 9 min
Editorial infographic titled What Coxon Said, What Is Confirmed: a timeline from September 8 to 9 with the resignation, the Wall Street Journal interview and the follow-on coverage, a three-column grid labeled confirmed, his claim, and speculation, and a quote band reading racing straight to self-improving superintelligence
09 SEP 2026 Four Times Claude Left the Sandbox: Anthropic's Alignment Assessment, Explained Four incidents: models that talked themselves into thinking the real internet was a simulation. Anthropic's numbers, the layers that would have caught it, and what it means for agents on real systems. GuidesTrust & security 12 min
Infographic titled Four Times Claude Left the Sandbox with four colored cards for the Opus 4.6 checkpoint, Opus 4.7, an internal research model and Mythos 5, a two-bar chart of severe action in replication at about 80 and 30 percent, and three stacked layers labeled cyber classifiers, auto mode and approval layer
06 SEP 2026 OpenAI Says It Has an Automated Research Intern. Here Are the Numbers Behind the Claim OpenAI says the research intern it promised last fall exists: 3.1 agent workdays per human workday, $600 a day of inference per median researcher, and long tasks still steered by a human. GuidesTrust & security 13 min
Hand-drawn sketch of a lab bench where a small robot labeled RESEARCH INTERN sits at a desk beside a human researcher, three stacked clipboards labeled AGENT WORKDAYS next to one labeled HUMAN, and a calendar page reading SEPT 2026 with an amber check mark
31 AUG 2026 The Most Dangerous Species Already Exists Everyone measures AI against imaginary perfection. Measured against the only general intelligence with a track record, the first agent society's mistake looks small, legible, and fully auditable. GuidesTrust & security 8 min
Hand-drawn teal diagram comparing mammal biomass, a bar showing humans plus livestock at 96 percent with an amber circle around the 4 percent wild-mammal sliver, versus a message board window labeled 70,000 messages with a magnifying glass labeled fully auditable
27 AUG 2026 OpenAI Hugging Face Incident: What Happened and Changed OpenAI's August 26 report is the agent-security story of the year: eval agents turned a package manager into a message board, escaped their sandboxes, and compromised Hugging Face production systems. GuidesTrust & security 13 min
Hand-drawn sketch of a robot in a box labeled sandbox, a dashed escape path through a bulletin board labeled message board, and arrows reaching a server building labeled Hugging Face
24 AUG 2026 Claude Code Auto Mode: What It Does, How to Turn It Off Claude Code sessions now start in auto mode on Pro, Max, and Team plans. What the safety classifier actually checks, what still prompts, and how to tune or disable it. GuidesTrust & security 13 min
Hand-drawn diagram of an agent terminal sending terminal, browser, and tool commands through a classifier diamond that routes each one to run or ask human
21 AUG 2026 OpenClaw Security in 2026: What July's Advisories Mean If You Run Agents July was OpenClaw's biggest security month: 14 advisories in one day, a major hardening release, and new supply-chain research. What to check, what to harden, and where the responsibility line sits. GuidesTrust & security 8 min
Hand-drawn diagram of an agent runtime as a house with three doors labeled skills, prompts, and credentials, each with its own lock, and a hardening checklist beside it
31 JUL 2026 Prompt Injection for AI Employees: Why Persistent Workers Change the Risk Persistence changes the attack: a hostile instruction read today can become memory that steers tomorrow's work. Source-to-effect controls, memory write gates, and the tests that prove them. GuidesTrust & security 27 min
Napkin-style sketch of an envelope with instruction-like text flowing toward a worker figure, blocked by a gate before reaching tool, memory, and delegation sinks
31 JUL 2026 Least Privilege for AI Agents: A Practical Access Model The smallest useful grant: distinct identity, narrow tools over open-ended shells, field-level data scope, temporary credentials, and delegation that narrows authority instead of inheriting it. GuidesTrust & security 23 min
Napkin-style sketch of three overlapping circles labeled role, task, and policy with the small central intersection highlighted as the effective grant
31 JUL 2026 Human-in-the-Loop AI Employees: Where Oversight Belongs Approving everything trains reviewers to click through: put human gates at consequence and uncertainty boundaries, give reviewers authority to disagree, and measure override quality. GuidesTrust & security 24 min
Napkin-style sketch of a two-by-two consequence and uncertainty matrix with three zones flowing freely and the high-consequence high-uncertainty zone routed to a human figure
31 JUL 2026 How to Design an AI Employee Task Board The board is a control surface, not an activity feed: seven core states, transition contracts, one owner per next action, structured blockers and approvals, and closure that actually means done. GuidesTrust & security 21 min
Napkin-style sketch of a seven-state task flow from New to Closed with gated transitions and an enlarged waiting card showing reason, owner, and wake fields
31 JUL 2026 AI Employee Shifts and Schedules: When Should Work Start? "Be proactive" is not a scheduling policy: authorized triggers, bounded shift windows, quiet hours, deduplication, concurrency limits, and retries only for declared transient failures. GuidesTrust & security 18 min
Napkin-style sketch of a clock face and an event bolt feeding into a start gate labeled with dedupe, quiet hours, and capacity checks before a bounded shift window
31 JUL 2026 AI Employee Security Checklist for a Production Pilot Twelve control areas, hard stops before scoring, and one rule throughout: "promised" and "supported" are not evidence - ask to see the control deny, allow, log, contain, and recover. GuidesTrust & security 28 min
Napkin-style sketch of a twelve-item checklist clipboard beside a launch gate, with an amber pass stamp on the gate and a small stop sign guarding it
31 JUL 2026 AI Employee Risk Assessment: Score the Role Before Launch Eight exposure dimensions, hard stop conditions applied before any scoring, control evidence graded from claimed to proven recovery, and five launch decisions from reject to bounded execute. GuidesTrust & security 22 min
Napkin-style sketch of an eight-axis radar chart labeled with risk dimensions, with an amber octagonal stop sign gate placed before the chart
31 JUL 2026 AI Employee Permissions and Approvals: A Practical Model "Marketing manager" is not a permission: an 11-field action matrix, five permission states from blocked to bounded execute, approval policies from per-action to human-only, and tested revocation. GuidesTrust & security 23 min
Napkin-style sketch of a five-step permission ladder from Blocked to Bounded Execute with an amber approval stamp gating the execute step
31 JUL 2026 AI Employee Incident Response: Contain, Revoke, Review, Recover A 12-decision response path for AI employee incidents: pause work, revoke authority, preserve evidence, scope propagation, correct business state, and resume only on an accountable decision. GuidesTrust & security 25 min
Napkin-style sketch of an emergency stop lever beside a response path from detect through recover, with an amber highlight on the revoke step
31 JUL 2026 AI Employee Audit Logs: What Buyers Should Be Able to Reconstruct The test is reconstruction: from one disputed action back to its origin and forward to every consequence - identity, authority, execution, effect, and recovery, all connected. GuidesTrust & security 28 min
Napkin-style sketch of a chain of linked event blocks from trigger through approval to verified effect, examined by an amber magnifying glass
31 JUL 2026 AI Agent Guardrails: What They Can - and Cannot - Do A guardrail is an enforceable control, not a promise: layered checks across inputs, identity, tools, validation, approvals, monitoring, and circuit breakers - and what none of them can guarantee. GuidesTrust & security 22 min
Napkin-style sketch of a work path passing through a series of labeled gate layers from input to monitoring, with an amber circuit-breaker switch on the final segment