Anthropic kept humans in the loop, and that tells you more than the biology finding
Anthropic's biology lab is reportedly on the verge of a significant scientific discovery, and Claude helped find it. What the announcement quietly confirms is that no one let Claude act alone with real tools. That constraint is not a limitation of the technology; it reflects how difficult meaningful human oversight actually is once agents connect to live systems.
Neo NeumannAI Practice LeadSeptember 23, 2026Listen to the podcast
4 min
Chapters
Key takeaways
- Map every irreversible action in a workflow before connecting an agent to live systems.
- Put a human checkpoint at each irreversible step itself, not at the end of the run.
- Ask vendors where the agent stops and waits, and whether a person can understand what they are approving.
- Treat autonomy scores from firms that sell agents as marketing, and cross-check against independent sources like O'Reilly Radar.
- Let the model do reading, hypothesis building and drudgery while humans own the moves you cannot take back.
Read the full transcript
Host:You're listening to Leader's Insights. Today's subject, anthropic kept humans in the loop, and that tells you more than the biology finding.
Expert:Picture a nuclear plant control room in the 1970s. Rows of dials, a wall of switches, and one rule that never bent. A human hand had to turn the final key. The reactor could run itself for hours, but nobody let it start a shutdown alone.
Host:And you're telling me that's the story hiding inside Anthropics Biology News this week.
Expert:Exactly that. Their lab is reportedly close to a real scientific finding. And Claude, their large language model, the thing that predicts text and now plans multi-step tasks, helped surface it, but read the fine print. Nobody let Claude touch the live instruments alone.
Host:Break that down. What does act alone with real tools actually mean here?
Expert:An AI agent is a model wired to actions. Not just write me a paragraph, but call this lab robot, order this reagent, run this assay. The moment the model connects to systems that do things in the world, you've handed it a set of keys. Anthropic kept a human on every key.
Host:Sounds cautious to the point of timid. Isn't the whole promise that these things run unattended?
Expert:That's the pitch, sure. And the vendors love the number. Google AI has floated that agents could I automate a huge slice of knowledge work. And OpenAI talks up autonomous task completion rates north of 90%. Worth remembering, both of them sell the agents. So crosscheck against something like O'Reilly radar before you build a budget on it.
Host:So the 90% is marketing. The 90% is the easy 90. The trouble lives in the last slice,
Expert:where the agent does something confidently wrong, and there's no dial telling you it went sideways.
Host:Give me the concrete version. What breaks?
Expert:Say the agent orders the wrong compound because it misread a concentration. In a chatbot, that's a typo you shrug off. In a wet lab, that's a ruined batch. A safety incident. Maybe a week gone. The action has wait. The text never did.
Host:But a human watching every step, doesn't that erase the whole efficiency argument? You've just hired a babysitter.
Expert:That's the uncomfortable part. And it's the real lesson. Meaningful oversight isn't a guy glancing at a screen and clicking approve. If the agent runs 40 steps in nine seconds, the human approving step 37 has no idea what happened in the first 36. You're not supervising. You're rubber stamping at speed.
Host:So keeping a human in the loop is harder than it sounds.
Expert:Much harder. Think of it like a co-pilot who's asleep until the plane's already nose down. Technically present, practically useless. Real oversight means the system pauses at the points that actually matter and shows you something a person can judge in the time they have.
Host:How did anthropic solve that then?
Expert:They narrowed where the agent could act. And they inserted checkpoints at the consequential moments. The physical steps. The irreversible ones. The model does the reading. The hypothesis building. The drudgery. The human owns the moves you can't take back. That division is the whole game.
Host:And you're reading that constraint as a signal, not an apology.
Expert:A confession more like, it tells you their internal answer to, can we let this thing run free on live systems? Is still no. Not because Claude is weak. It's plainly strong enough to find the science. Because they haven't solved the supervision problem and neither has anyone else.
Host:So when a vendor sells me a fully autonomous agent for my operations next quarter?
Expert:Ask them one thing. Where does it stop and wait for a person? And can that person actually understand what they're approving? If the answer is, it doesn't stop. You're not buying automation. You're buying an unattended reactor and hoping the dials are honest.
Host:Give me the one thing to walk out with.
Expert:Before you deploy any agent that touches a real system, map your irreversible actions first. Adriant. The moves you can't undo. And force a human checkpoint at each one. Not at the end. At the exact step. That single list is worth more than any autonomy score a vendor will wave at you.
Host:This episode draws on TechCrunch AI, Google AI, vendor, AI Lab, the decoder, open AI, vendor, AI Lab, KD Nuggets, O'Reilly Radar. We'll stop there. The AI tools for making the call are at MBA-training.com.
The concept is human-in-the-loop oversight inside agentic workflows. Professionals hear it and assume they understand it. You review what the AI does before it does anything consequential. Simple enough. The problem is that "review" becomes nearly impossible to do well once an agent is calling real tools, executing actions in live systems, and chaining dozens of steps together in minutes. What sounds like a governance checkbox turns out to be one of the harder engineering and organizational problems in applied AI.
Anthropic's biology lab is the clearest illustration available right now. TechCrunch reported in 2026 that the lab has apparently found something significant, with Claude playing a role in the discovery process. The detail buried in that story is more instructive than the finding itself: Anthropic has not allowed Claude to operate autonomously inside that lab. Humans remain in the loop at every meaningful decision point. This is a company that builds the model, understands it better than anyone, and is running experiments in a controlled research environment, and it still chose not to hand over full control. That choice says something about the actual state of trust in agentic systems, even from those closest to the technology.
What does human-in-the-loop oversight mean for agents with tool access?
Most professionals encounter agent oversight as an abstract policy question: should we have a human review AI outputs? In practice, the question only gets hard when the agent has hands, meaning access to tools that write, send, delete, book, purchase, or execute.
ChatGPT's Work tab, now available on mobile for Pro and Plus users according to TechCrunch, gives agents access to email, calendar, and Slack. OpenAI reports (as a vendor with an interest in adoption, so take the framing accordingly) that Ringg's customer-call agents using GPT-5.6 resolve up to 65% of calls without human involvement. O'Reilly Radar's analyst coverage of MCPMCPAn open standard that lets AI assistants connect to your company's tools and data in a consistent, governed way, instead of one custom integration at a time.View full definition →, the protocol that connects tools to LLMs, describes how MCP is not just APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → plumbing but a mechanism that can give models broad, composable access to external systems. Each of these represents agents with real reachreachThe number of unique people exposed to your message in a given period. Unlike impressions, reach counts each person once, no matter how often they see it.View full definition → into real environments.
At that point, oversight is no longer about reading a draft before hitting send. It is about understanding what the agent did, in what order, with what permissions, and whether any of those actions are reversible. Most organizations are not equipped to answer those questions in real time.
How agent oversight breaks down inside a four-minute run
Consider a concrete example. A procurement agent is given access to your company's supplier database, your email, a contract management system, and a scheduling tool. Its task: identify three vendors for a new component, draft initial outreach, and schedule qualification calls.
In a single run, the agent might query the database, cross-reference supplier ratings, draft and send six emails, create six calendar invites, and log notes in the contract system. This takes perhaps four minutes. A human reviewer approving this workflow at the start approved the task description, not those forty-three individual actions.
This is wherethe gap between approval and control becomes visible. You said yes to the goal. You did not say yes to the specific emails sent, the specific suppliers contacted, or the specific times booked. If the agent misread a supplier's status, contacted a vendor you had quietly blacklisted, or scheduled a call for a time zone that creates a compliance issue, you may not know until the damage is done.
There are three specific failure modes worth naming:
- Actions compound before any single one flags a problem. The agent's third step looks fine in isolation but is only wrong given the context of the first two.
- Reversibility varies wildly by tool. A draft saved in Google Docs is easy to undo. An email sent, a calendar invite accepted by an external party, or a contract log entry that triggers a downstream workflow is not.
- Oversight fatigue sets in fast. If your review process requires a human to read agent traces for every run, that human will start approving without reading within a week.
The Anthropic biology lab decision reflects an awareness of all three. Scientific research compounds errors dangerously. Many experimental actions are not reversible. And researchers are expensive; you cannot afford to have them become rubber stamps.
Matching oversight to consequence per action and reversibility
Oversight does not need to be uniform across every agent deployment. The relevant variable is consequence per action, combined with reversibility.
A content-drafting agent writing blog post outlines can operate with minimal human intervention because the output is reviewed before publication and nothing is sent to an external system. An agent that books meetings with clients, submits expense reports, or executes trades sits in an entirely different category even if the individual steps look routine.
A practical way tothink through what tools you give an agent is to ask: if this action happens incorrectly and I find out 48 hours later, what does remediation cost? If the answer is "an awkward email," the risk profile is manageable. If the answer is "a regulatory filing, a client relationship, or an unreversible financial transaction," then full autonomy at that step is probably not worth the efficiency gain.
The honest tradeoff is this: meaningful human oversight at the action level slows agents down considerably, and much of the speed advantage disappears. Oversight at the goal level preserves speed but gives you much less actual control than you think you have. Most deployments right now sit somewhere uncomfortable in the middle, with approval workflows that feel reassuring but would not catch the failure modes that actually matter.
Anthropic's choice to keep humans in the loop in its biology lab is not a statement about Claude's capability. It is a statement about the cost of being wrong in a domain where wrong is expensive. Before deploying an agent with access to live tools, the useful question is not "do we trust the model?" but "do we trust our ability to catch its mistakes before they compound?" In most production environments, the honest answer is no, and the deployment design should reflect that.
Frequently asked questions
Why did Anthropic keep humans in the loop in its biology lab?
Anthropic kept humans at every meaningful decision point in its biology lab because errors in scientific research compound and many experimental actions cannot be reversed. The company builds Claude and runs the lab in a controlled setting, yet still declined full autonomy, which says more about trust in agentic systems than the biology finding itself.
What is the difference between goal-level and action-level approval for AI agents?
Goal-level approval means a human says yes to the task description, while action-level approval means reviewing each individual step the agent takes. A procurement agent can run forty-three actions in four minutes: querying databases, sending six emails, booking six invites. Approving the goal preserves speed but gives far less control than most teams assume.
How do I decide which tools an agent can use without human approval?
Ask what remediation costs if the action happens incorrectly and you discover it 48 hours later. If the answer is an awkward email, the risk is manageable. If it is a regulatory filing, a damaged client relationship, or an irreversible financial transaction, full autonomy at that step rarely justifies the efficiency gain.
What is oversight fatigue in agentic workflows?
Oversight fatigue is what happens when a review process requires humans to read agent traces for every run: within about a week, reviewers start approving without reading. It is one of three failure modes in agent oversight, alongside actions that compound before any single step looks wrong and reversibility that varies sharply by tool.
Go deeper
The lessons that take this article further, free to read.
- 1Guardrails, permissions, and human-in-the-loopAI agents: design, build & operate
- 2Tools and function calling: giving your agent handsAI agents: design, build & operate
- 3Evaluating and debugging agents: traces, evals, and failure modesAI agents: design, build & operate
- 4Agents vs workflows vs automations: choosing the right level of autonomyAI agents: design, build & operate
- 5Connector safety, permissions, and governanceClaude & the Anthropic ecosystem
Sources
- Anthropic says its biology lab has already found something big
- Google Beam expands with new regions, partners, and customers
- ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access
- Google's new Flash TTS models let you design AI voices from scratch using text descriptions
- ChatGPT mobile app gets voice-based agentic features
- Even Americans who use AI every day are worried about it
- YouTube adds AI tools to Creator Studio with script coaching, smart thumbnails, and Gemini editing
- Two years of OpenAI Academy
- Anthropic engineer explains why Claude's writing got worse although the model got smarter
- Everything Claude Opus 5.5 Actually Ships With
- YouTube will let you build your own algorithm with AI
- Sam Altman’s remarks at the United Nations Security Council
- Ringg’s AI agents resolve up to 65% of customer calls with OpenAI
- MCP Is Not Just Another API Standard
Finished reading?
Validate your read to earn XP and feed your radar.