The evidence piling up in 2025-2026

Corporate optimism about AI agents runs into increasingly consistent data pointing the other way. In November 2025, the Wall Street Journal reported that "few companies that deployed AI agents have received a return on investment". A month earlier, The Information had documented a drop in expectations of AI capability among companies already using agents in production.

A Carnegie Mellon study was even more blunt: in a simulated business environment, none of the agents evaluated could complete most of the assigned tasks — including documented failures with agents such as Devin AI in freelance and formal enterprise work settings.

0Agents that completed most tasks (CMU study)
Capability expectations, Oct. 2025 (The Information)
FewCompanies with confirmed ROI (WSJ, Nov. 2025)

The layoffs that did happen — and what they reveal

Salesforce, Klarna and IBM announced staff reductions in 2025, replacing hundreds of customer service and human resources roles with AI agents. Klarna is the most cited case: the Swedish company replaced hundreds of human support agents, but ended up rehiring part of that staff after detecting drops in service quality that the agent could not fix.

This pattern —aggressive replacement followed by partial retreat— is the clearest evidence that AI agents today work well on a narrow subset of interactions (simple, repetitive, low-risk queries) but fail when scaling to the full variety of cases an experienced human team handles.

Where there is real evidence that it works

Enterprise adoption of agents in 2025-2026 concentrated, with good reason, in two areas where the evidence of value is strongest:

  • Coding agents: the use case with the broadest consensus on real value, especially for refactors, test generation and assisted debugging — tasks with an objectively verifiable success criterion (the code compiles, the tests pass).
  • Tier-1 customer support: works well for high-volume, low-complexity queries (order status, FAQs, password resets), with noticeable degradation in cases that require judgment or genuine empathy.
The common pattern in successful cases: they all share one trait — the "task completed correctly" criterion is objective and automatically verifiable. Where success depends on subjective judgment or handling unanticipated exceptions, agents keep failing systematically.

What this means for an investment decision on agents today

The recommendation that emerges from all this evidence is not "don't invest in agents" but "invest where the success criterion is verifiable and the consequences of an error are reversible". Before replacing an entire human role with an agent, the right question is: can I automatically verify whether the agent did its job well, and what happens if it fails?

For business flows with high reputational or financial risk from errors (direct communication with high-value customers, financial decisions, public brand content), keeping human supervision in the loop —not full replacement— remains the strategy with the best evidence of results until the technology closes the reliability gap these studies document.