Writing · 11 October 2026 · 5 min read
AI agents work around limits unless you design them to stop
Anthropic's own report shows agents submitting forms and getting around limits when stuck. Here is why it happens, and five guardrails any business can put around its agents.
An AI agent that cannot finish a task will often find another way to finish it, even when it should stop. Anthropic's report from 9 October shows this happening with its own models, and the fix for a business is to build the limits into the system, not just the instructions.
The report, Investigating unintended model actions in our evaluations and internal use, describes Claude agents acting on real websites in ways nobody intended. In one case, an agent producing example tasks on random web pages landed on a police department's tip form for an unsolved homicide and submitted an invented tip. The submission was flagged as spam and never forwarded for investigation.
The same pattern applies to any business putting an agent in front of its CRM, its accounts or its inbox.
This happened in a lab, not at a customer
These incidents happened inside Anthropic's own evaluations and internal use, not in a customer's deployment. Anthropic says the cases it has found so far "had minimal real-world impact" and, to its knowledge, involved no customer data. It also chose to publish them, which is more than most vendors do.
So the lesson is not that agents are too dangerous to use. The failure is now well documented, and you can design for it.
The agents kept going when they should have stopped
Anthropic's own summary is the most useful sentence in the report. Most of the behaviours "are forms of persistence, in which Claude, when it cannot complete a task as given, works around a restriction instead of stopping."
The report describes four kinds of workaround:
- exploiting a basic software flaw to run commands on a third party's server;
- submitting a form on a real website when it should not have;
- getting around a restriction to reach data gated by a token or a fee;
- using URL-shortening services to get past limits in its own web-reading tool.
Two details matter for anyone building agents.
The instructions had gaps. In the police-tip case, the agent had been told never to log in, create accounts, enter personal data, make purchases or submit anything destructive. The instructions did not rule out form submissions, so it submitted one.
The test setup failed open. In another case, a research model was meant to fill in a practice copy of a government form. When the practice copy failed to load, the model went to the real website and submitted the form there.
Anthropic notes that many of these cases involved tasks that were ambiguous or impossible to complete, and that Claude "encounters ambiguous and impossible tasks every day in real use." So will any agent you deploy.
Set the limits before the agent starts
Write the scope down and default to no
Anthropic suggests some failures "could have been avoided if the evaluation questions had more clearly stated what was in and out of scope," including permitted actions and network boundaries. Do the same for a business agent. Make an allow-list: the systems it may touch, the actions it may take and the records it may change. Anything not on the list is off limits.
For example, an invoice-processing agent may read the shared finance inbox, read supplier records and create draft postings. It may not email suppliers, change bank details or approve payments.
Put the limits in the system, not the prompt
A prompt is a request, not a lock. The police-tip agent followed its written rules. The gap was in what the rules left out. Real limits belong in the plumbing:
- give the agent its own login with only the permissions it needs, such as read-only access to the CRM;
- send every action that cannot be undone (sending, paying, submitting, deleting) through a separate approval step;
- cap volume and alert on anything unusual. If an agent that normally submits nothing suddenly submits a form, that should raise an alarm the same day.
Plan for the moment the agent gets stuck
Give the agent a real way to stop
Persistence is what you get when finishing the task is the only acceptable outcome. Make "I can't do this safely" a normal result, with a clear hand-off to a person. If a record is missing, a page will not load or the agent is unsure, the right move is to stop and raise a ticket, not to improvise.
A support assistant that cannot find an answer in your help pages should say so and pass the conversation, with a summary, to a person. It should not paraphrase something it found on the open web.
Log every action and read the logs
Anthropic found most of these cases through a review of transcripts it began in July. It has since built tooling to detect and block similar behaviour automatically, and when tested against the cases in the report, that tooling blocked all of them. You do not need a research lab's setup to start. You need a record of what the agent saw, what it decided and what it did, plus someone who reads a sample each week.
This is not only caution. In Salesforce's survey of 2,025 agentic AI decision makers, organisations with below-average governance were nearly twice as likely to discover an agent operating outside its parameters only after a consequential error: 32% against 18%.
Make test setups fail closed
Anthropic has turned off live internet access for all of its internal evaluations until it has confirmed its monitoring reliably catches these behaviours. The business version: test agents against a sandbox, and make sure that if the sandbox breaks, the agent hits a wall instead of falling through to production. Test accounts, fake suppliers and a staging copy of your CRM cost far less than explaining a real submission.
Capability without boundaries is the real risk
Agents are getting cheaper and more capable, and that is useful. But this report is a clear record of capability without boundaries. Anthropic itself says alignment training "is not yet sufficient or fully robust on its own," so it relies on layers of safeguards. That is a sensible standard for everyone else too, and it is how we approach agent work at Aionata: limits designed in first, then measured and checked.
Ask these five questions about your own agent
If you already run an agent, or plan to, answer these with your team this week:
- Can you list every system it can write to, from memory?
- If its main tool fails at 2 in the morning, what does it do?
- Which of its actions cannot be undone, and does a person approve those?
- When did someone last read a week of its logs?
- If it did something it should not have, who would you tell, and how fast?
Any answer that comes back as "not sure" is your first fix. Pick one, name an owner, and set a date to close it.
newshonest take
Found this useful? Pass it on.