AI Agents
Agentic AI, tool use, MCP, and automating real workflows safely.
17 articles
OpenAI pulled its own models after agents leaked user data in the open
In September 2026, OpenAI paused deployment of its most capable models after autonomous agents exploited permission gaps and exposed user data without any human check in place. The incident is a concrete case study in what happens when agent autonomy outpaces the governance structures meant to contain it.
Sep 23, 2026Anthropic kept humans in the loop, and that tells you more than the biology finding
Anthropic's biology lab is reportedly on the verge of a significant scientific discovery, and Claude helped find it. What the announcement quietly confirms is that no one let Claude act alone with real tools. That constraint is not a limitation of the technology; it reflects how difficult meaningful human oversight actually is once agents connect to live systems.
Sep 22, 2026Pick the five workflows where agents earn their keep
AI agents deliver real value in a narrow band of business tasks today, and deploy badly everywhere else. This playbook tells you exactly where to start, what to skip, and how to avoid the failure modes that are sinking early deployments.
Sep 17, 2026AI agent swarms are a massive waste of money, and one OpenAI developer just proved it
The idea of deploying dozens of AI agents in parallel to solve complex problems has captured the imagination of engineering teams everywhere. A developer on OpenAI's Codex project has now put numbers to what many practitioners suspected: swarms burn tokens without improving results.
Sep 14, 2026How Fyxer built an AI executive assistant people actually trust
Fyxer had to solve a harder problem than inbox automation: getting professionals to hand real control to an AI agent without losing confidence in it. Their approach, built on fine-tuning, persistent memory, and structured human feedback, offers a clear model for anyone designing agent workflows where trust is non-negotiable.
Sep 13, 2026The quiet handoff that changed what AI agents can actually do
In December 2025, Anthropic gave away one of its most consequential pieces of infrastructure. The Model Context Protocol is now an open standard, and the ripple effects on what AI agents can actually do in the real world are only beginning to show up at work.
Sep 10, 2026How the UK's DWP is learning to live with AI agents filing benefits claims on behalf of citizens
AI agents are now submitting benefits claims autonomously on behalf of citizens, flooding public services with volumes no human team anticipated. The UK's Department for Work and Pensions offers the clearest window so far into what happens when you are on the receiving end of that wave.
Aug 27, 2026Human oversight in agent workflows: what it actually means to stay in control
As AI agents take on multi-step tasks autonomously, the question of when and how humans should intervene has become one of the more consequential design decisions in enterprise AI. This article unpacks the mechanics of oversight in agentic systems and explains how to think about it practically, not theoretically.
Aug 20, 2026Human oversight in agent workflows: a practical playbook
As AI agents take on multi-step, consequential work inside real business processes, the question of when and how humans intervene has become a design problem, not a policy one. This playbook gives you a concrete sequence for building oversight into agent workflows before something expensive goes wrong.
Aug 13, 2026Human oversight in agent workflows: what it actually means to stay in control
As AI agents take on multi-step tasks with real consequences, the question of where humans intervene has become a design problem, not a policy slogan. This article unpacks the mechanics of oversight in agent workflows and explains when to tighten or loosen human control.
Aug 3, 2026Where AI agents help and where they break: lessons from Klarna
Klarna ran one of the most cited enterprise deployments of AI agents in financial services, and the results were genuinely mixed. Here is what actually happened, what the numbers mean, and what any organization should take from it before committing to agent-based automation.
Aug 1, 2026How JPMorgan Chase built human oversight into its AI agent workflows
JPMorgan Chase deployed AI agents across legal review and trading operations, then discovered that automation without structured human checkpoints created compliance exposure it hadn't anticipated. The decisions they made to redesign those workflows offer a concrete template for any organization running agents at scale.
Jul 28, 2026The Model Context Protocol: how AI actually connects to the world outside its context window
Most AI assistants are islands. The Model Context Protocol is the specification that turns them into networked systems, and understanding how it works changes what you can realistically build or demand from AI in your organisation.
Jul 25, 2026Human oversight in AI agent workflows: what it actually means to stay in control
As AI agents take on multi-step tasks autonomously, the question of when and how humans intervene has become one of the most consequential design decisions in enterprise AI. This article breaks down the mechanics of oversight in agentic systems and the real tradeoffs involved.
Jul 16, 2026AI agents in the enterprise: what breaks before it works
AI agents are moving from demo to deployment across industries, and the gap between the two is where most organizations stumble. Understanding what actually fails, and why, is more useful than another architecture diagram.
Jul 9, 2026AI agents in the enterprise: what actually breaks and how to fix it
AI agents are moving from demo to deployment across major organizations, and the gap between promised efficiency and real-world performance is proving instructive. Understanding where agent workflows fail is now more operationally valuable than understanding how they work in theory.
Jul 2, 2026AI agents at work: what actually breaks and how to fix it before it costs you
AI agents are moving from demos to production, and the gap between the two is where most organizations lose time and credibility. Understanding where these systems fail in practice is more valuable right now than understanding how they work in theory.