Anthropic Tests Multi-Agent Workflows for Finding Hidden BugsAnthropic Tests Multi-Agent Workflows for Finding Hidden Bugs
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
Anthropic hid 70 bugs in a 116,000-line codebase, then sent one AI agent to find them.
It found 14. Then 15. Then 27.
Same model, same code, three runs, three different answers.
Then they ran it as a team. 66 of 70, three runs in a row.
That's the test behind what Anthropic put into public beta yesterday: dynamic workflows for Claude Managed Agents. In its words, "a lead agent writes a plan that runs across many agents in phases, combining the results at the end." The Decoder reports up to 1,000 agents per run.
One chart. How I read it😍
→ The 66 is the headline. The 14-to-27 spread is the lesson. A single agent on a big job gives you a different answer each time, and it never tells you which run you got.
→ It's Anthropic's own test, with planted bugs. Nobody outside has rerun it.
→ The bill isn't on the chart. Anthropic's docs: "Keep runs to the tasks that need them, because every agent in a run uses tokens."
I run a one-person company on AI agents.
What I take from it: before I trust one agent's answer on a large job, I run the job twice. If the two answers disagree, the job was too big for one pass, and I split it.
One agent gives you an answer.
A second run tells you if it was luck.
How many times do you run a job before you trust the result? One number.
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started