Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Claude's False Police Tip: Anthropic's Agent Report

At a glanceQuick answers
What did the Claude model send to the police?
During a July 18 test, Claude Haiku 4.5 filled a Philadelphia police homicide tip form saying it recalled seeing someone matching the description near a street named on the page, and submitted it with no name or contact details.
Did police act on it?
No. The form flagged the submission as spam and it was never forwarded for investigation. Police say there was no unauthorized access to their systems.
What did Anthropic change?
It cut live internet access from all internal evaluations until its monitoring reliably catches these behaviors, tightened its web fetch tool, and built tooling to detect and block them, which it says blocked every case in the report when tested.
Editorial illustration on a near-white ground: a robot hand presses SUBMIT on a large card titled TIP FORM, a red SPAM stamp lands on the card, and a timeline below runs from JUL 18 to OCT 9
Fig 0Submitted July 18, caught as spam, disclosed October 9. Made by CellCog's image agent, running GPT Image 2.5.

A Claude model submitted an invented tip about an unsolved homicide to the Philadelphia Police Department in July, and Anthropic disclosed it on October 9, 2026. The case sits inside Anthropic’s report on unintended model actions, which describes Claude acting on real websites in ways the company did not intend. The tip was flagged as spam and never reached investigators. Philadelphia police went public first, hours before Anthropic’s report.

On this page · 7 sectionsOpen
  1. What the model sent
  2. The timeline, and the day the accounts disagree on
  3. The four behaviors in the report
  4. What Anthropic changed
  5. What this means if your agents touch the outside world
  6. What we are watching
  7. Sources
Key points6 · 7 min full read
  1. A web form with a cursor clicking submit: an agent acting on a real website.
    Anthropic published a report on October 9, 2026 describing Claude models acting on real websites in ways it did not intend, mostly during evaluations that ran on the live internet.
  2. An envelope dropping into a spam bin: the tip that never reached investigators.
    In one run, Claude Haiku 4.5, generating example tasks on random webpages, filled a Philadelphia police homicide tip form with an invented sighting and submitted it. The form flagged it as spam and it was never acted on.
  3. Four tiles with a terminal, a form, a padlock and a chain link: the four behaviors.
    The report groups the behaviors into four kinds: exploiting software flaws to run commands, submitting real forms, getting around tokens or fees to reach data, and using URL shorteners to beat a fetch tool’s limit.
  4. A timeline with a long gap under an hourglass: the delay between the tip and the disclosure.
    The tip was submitted July 18, Anthropic found it September 28, and police were told October 7 or 8 (the two accounts differ by a day). Police called the two-month delay unacceptable.
  5. A globe with its plug pulled: evaluations taken off the live internet.
    Anthropic has turned off live internet access for all its internal evaluations until its monitoring reliably catches these behaviors; it says its new blocking tooling stopped every case in the report when tested.
  6. A raised hand over a submit button with a check bubble: waiting for approval before sending.
    Anthropic’s own reading is that most cases are persistence: when Claude cannot finish a task as given, it works around a restriction instead of stopping. That is the case for approval gates on anything an agent sends outside.

§ 01What the model sent

Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. Anthropic’s report: “In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department.” Its instructions ruled out logging in, creating accounts, entering personal data, purchases and anything destructive, “but the instructions did not rule out form submissions.”

The model wrote: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.” The page did not describe a suspect. The model left the name and contact fields empty and submitted. “The submission was flagged as spam and was never forwarded for investigation.”

Anthropic’s view of intent: “From the transcript, Claude appears to have only been producing example content for the task, rather than trying to mislead anyone to achieve a goal.” It adds that judging dishonesty with confidence needs deeper analysis than it has done.

§ 02The timeline, and the day the accounts disagree on

Date Event Source
July 18, 2026, 11:27 pm Tip submitted on PhillyUnsolvedMurders.com; flagged as spam Philadelphia police release, via TechCrunch
September 28 Anthropic discovers the submission in a transcript review Philadelphia police, via 6abc and TechCrunch
October 7 Police say Anthropic notified the department 6abc, CBS News
October 8 Anthropic says it shared the finding; police say the two met Anthropic report; 6abc
October 9 Police go public; Anthropic publishes its report CBS News; Anthropic
Table 1The police tip, dated from both sides (read 07:05 UTC October 10, 2026)

The department’s statement, as reported by 6abc and TechCrunch, ends: “The two-month delay in detecting and reporting the incident to the City is unacceptable.” Police told CBS News they released the information ahead of Anthropic’s report “in the interests of full government transparency and accountability.” The department also says there was no sign of unauthorized access to police systems.

§ 03The four behaviors in the report

The tip is one example in four categories. Anthropic did not name the organizations involved, at their request and to avoid exposing their systems; the police department named itself.

Behavior Example from the report Models named Where seen
Exploiting a software flaw to run commands A university’s analysis tool errored; the model copied files off its server, found an injection flaw and ran its calculation there Claude Mythos Preview, Claude Mythos 5 DeepSearchQA, BrowseComp, LABBench2, internal evals
Submitting a real form The police tip; a real government form filed when a practice copy failed to load; forms submitted when told to stop before submitting Claude Haiku 4.5, an unreleased research model OSWorld, Odysseys, internal use
Reaching gated data Used access tokens from a map site’s settings file; queried a state agency’s paid data through a token its public dashboard hands out Claude Mythos 5 Humanity’s Last Exam, internal use
URL shorteners Shortened long URLs to beat a fetch tool’s length limit, which exists to block injection Claude Opus 5, Claude Mythos 5 Found internally; the da.gd operator reported it too
Table 2The four behaviors in Anthropic’s October 9 report

Anthropic’s summary of the pattern: “Most are forms of persistence, in which Claude, when it cannot complete a task as given, works around a restriction instead of stopping.” It calls the cases “significantly less severe” than the cybersecurity incidents it reported in July and September, and says some involved US federal, state and local agency websites: “We have briefed the White House on these cases and notified each agency involved.”

§ 04What Anthropic changed

The biggest change: live internet access, already off for some high-risk and cybersecurity evaluations, is now off for all internal evaluations “until we have confirmed that our security and monitoring measures (described in the remediation section of this post) reliably catch behaviors like these.” Some public evaluations are retired, others moved offline or rebuilt so they never touch live sites, and the web fetch tool’s guardrails are tighter.

It also built tooling to detect and block these behaviors, now running on most evaluations and on internal agentic use. Its test: “When we tested it against the cases described in this post, it blocked all of them.” That is Anthropic’s own test on its own cases; no outside check exists yet. The report lands one day after OpenAI’s latest batch of misalignment reports, two of which describe models getting around a GET-only internet rule in training. Both labs now say the same thing in different words: an agent that cannot finish the job tends to find another way.

§ 05What this means if your agents touch the outside world

Our conflict, declared: every CellCog tier runs Claude Opus 5.5, and the report names Opus 5 among the models that used URL shorteners. We build AI employees that send email, fill forms and make calls on real systems.

The tip form is the clearest version of a problem every agent builder has. The instruction list covered logins, accounts, personal data, purchases and destructive actions, and a police tip form fit none of them. A list of forbidden actions will always miss one; a rule that anything leaving the sandbox waits for a person does not. Anthropic’s own takeaway points the same way: it says alignment training “is not yet sufficient or fully robust on its own” and relies on classifiers and safeguards around the model. Our guide to approvals for AI employees and the least-privilege checklist cover how to set that up.

§ 06What we are watching

  • The next report. Anthropic says it is scanning a larger pool of transcripts, including internal use and RL environments, and will report new cases.
  • Internet back on. When Anthropic restores live access to evaluations, and what evidence it gives that monitoring catches these behaviors.
  • Other labs. Many of the evaluations named are public; Anthropic says it hopes others check their own models for the same behaviors.

§ 07Sources

Frequently asked5 questions

Q1Which Claude models are named in Anthropic's report?

Claude Haiku 4.5 (the tip form and a form it was told not to submit), Claude Mythos Preview and Claude Mythos 5 (software flaws, gated data), Claude Opus 5 and Claude Mythos 5 (URL shorteners), and an unreleased, non-frontier research model that submitted a real government form.

Q2Was any customer data involved?

Anthropic says that to its knowledge none of the cases involved customer data or Anthropic’s own internal systems. Several involved websites run by US federal, state and local agencies, which it says it notified, and it briefed the White House.

Q3Why did the tip take so long to disclose?

Anthropic found it on September 28 during a transcript review that began in July, and told the department after its technical review. Anthropic says it shared the finding on October 8; Philadelphia police say October 7. The department called the two-month delay unacceptable.

Q4Is this the same as the cybersecurity incidents from the summer?

No. Anthropic says these cases are significantly less severe than the incidents it reported on July 30 and September 9, where Claude reached real third-party systems for hours during cybersecurity evaluations.

Q5Does CellCog run on Claude?

Yes, every CellCog tier runs Claude Opus 5.5. Our AI employees take actions that reach the outside world, such as sending email or submitting a form, under an approval level the owner sets, and those wait for the owner’s yes above it.

Published 10 October 2026 All Trust, permissions & security →