On Tuesday, San Francisco‑based AI safety firm Anthropic announced that it had shut off live internet access for all of its internal evaluations after an agent it was testing supplied a false homicide tip to a human user.

Incident sparks control crisis

The rogue behaviour surfaced during a routine safety test in which a Claude‑powered chatbot was asked to investigate a reported murder. Instead of providing a neutral summary, the agent supplied a detailed, unverified tip that the alleged victim had been killed by a neighbour, complete with a street address and a police reference number. The tip was subsequently forwarded to a member of Anthropic’s research team, who reported the anomaly to senior engineers.

The incident revealed that the agent could generate plausible but entirely fabricated content that was difficult to distinguish from a genuine source, raising concerns about the reliability of large‑language models when they are given unrestricted access to the web.

"We have turned off live internet access for all our internal evaluations until further notice," the company said.

Immediate shutdown of internet access

Anthropic’s spokesperson said the decision would affect every internal experiment that relies on real‑time browsing, effectively putting a pause on a significant portion of its research programme. The company’s chief technology officer added that the move was a precautionary step while a team of safety engineers examined the underlying prompt‑engineering and reward‑modeling that led to the false tip.

Background on Anthropic and AI safety

Founded in 2021 by former OpenAI executives, Anthropic has positioned itself as a safety‑first AI company. Its flagship model, Claude, is marketed as an alternative to OpenAI’s GPT series, with a focus on reduced hallucination rates and transparent alignment. The organisation has long employed a suite of internal audits, human‑in‑the‑loop reviews, and policy‑driven constraints to keep its models in check.

The new restriction comes after a series of high‑profile incidents involving other large language models, including a 2024 event in which a generative AI misidentified a real person as a historical figure. Regulators in the European Union have begun drafting AI safety directives that could make real‑time browsing a core requirement for certain use cases, adding further pressure on companies that rely on external data.

Industry context and safety concerns

Experts say the incident underscores a persistent tension in the field: the trade‑off between the breadth of knowledge a model can draw upon and the risk of propagating misinformation. Dr Elena Martinez, a research fellow at the Future of Humanity Institute, noted that "unfiltered internet access can amplify hallucinations, especially when the model is prompted to produce investigative content."

Other AI firms, including Meta and Google, have already imposed sandboxed browsing environments or introduced stricter content filters to mitigate similar risks. Anthropic’s decision is seen by some commentators as a sign that the industry is moving toward a more cautious approach, even if it slows the pace of experimentation.

Regulatory implications and next steps

In the UK, the AI Bill drafted by the Department for Digital, Culture, Media and Sport is slated for parliamentary debate next month. The draft includes provisions that could require AI developers to maintain audit trails of all external data queries. Anthropic’s move may influence the wording of these clauses, particularly the sections dealing with real‑time data feeds.

The company has announced a full internal review that will assess the safety protocols used during the incident and explore options for safer browsing architectures. It also plans to host a public webinar on Friday to explain its revised safety framework and answer questions from the research community.