The Woman Who Launched IBM Watson Just Fired Half Her AI Agents
Sol Rashidi helped launch IBM Watson — one of the most ambitious AI systems in history. She's held Chief AI Officer roles across Fortune 500 companies. And last month, she fired half her AI agents.
Her reason? "It was actually taking me more time managing them and course correcting than it would for me to actually train an early career adult."
That's not a skeptic talking. That's one of the most experienced AI leaders on the planet telling you the technology isn't ready for the responsibility we're giving it.
The Problem: Most AI Agents Are Interns Pretending to Be Employees
Rashidi's experience isn't an outlier. It's the baseline.
Only 1 in 3 AI agents performs consistently at enterprise scale. The other two fall into a predictable failure pattern: they work beautifully in demo, degrade quietly in production, and eventually cost more in oversight than they save in automation.
The failure modes are consistent across industries:
-
Silent drift. An agent that worked perfectly for two months starts producing subtly wrong outputs. No error message. No crash. Just gradually worse results that nobody catches until a client complains.
-
Context collapse. Agents lose track of what they're supposed to be doing in long-running tasks. They follow the letter of their instructions while violating the spirit. Like an intern who fills out every field on a form — including the ones that say "DO NOT FILL."
-
Overconfidence. Agents don't know what they don't know. They'll produce a confident, well-formatted, completely wrong answer. A human employee would say "I'm not sure about this." An agent says "Here's your analysis" with a straight face.
The analogy Rashidi uses is precise: AI agents are interns. They can do defined tasks with clear inputs and outputs. They need supervision. They make mistakes that require experience to catch. And you wouldn't give an intern access to your production database without a manager reviewing their work.
Yet that's exactly what most companies are doing.
The Solution: Evaluate Agents Like You'd Evaluate a New Hire
The companies getting real value from agents have stopped treating them as autonomous workers. They treat them as managed resources with structured evaluation.
Here's what that looks like in practice:
30-day trial periods. Every agent gets a defined evaluation window with specific success criteria. Not "does it work?" but "does it produce output that meets our quality bar, at the speed we need, without supervision escalation beyond X hours/week?"
Error rate tracking. Every agent output gets logged and sampled. If the error rate exceeds a threshold — typically 5-10% for production work — the agent gets pulled, retrained, or replaced. Most companies don't track this at all.
Supervision cost accounting. The real cost of an agent isn't the API call. It's the human time spent reviewing, correcting, and managing its output. Smart teams track this explicitly. When supervision cost exceeds the cost of doing the work manually, the agent gets fired.
Escalation protocols. Instead of letting agents handle everything autonomously, define what they can do independently and what requires human approval. A junior employee doesn't sign off on a $50K purchase order. Neither should an agent.
This isn't complicated. It's basic workforce management. The problem is that companies got so excited about "AI agents replacing jobs" that they forgot to apply the same rigor they'd use hiring an actual person.
The Benchmarks: What Enterprise Agent Performance Actually Looks Like
- 1 in 3 agents performs consistently at enterprise scale (Sol Rashidi, industry data)
- 74% of enterprises have rolled back AI agent deployments from production (multiple surveys, 2026)
- 85% piloting, 5% shipping — the agent production gap (Cisco, 2026)
- 6.4 hours/week spent by workers fixing AI output ("botsitting," Glean survey)
- Caveat: These numbers are self-reported by enterprises already invested in AI. The actual failure rate for companies just starting their agent journey is almost certainly higher. Early adopters have better infrastructure than the mainstream.
The narrative that "agents will replace knowledge workers" runs headfirst into the reality that most agents can't reliably do knowledge work without supervision.
The Impact: The Agent ROI Math Doesn't Work (Yet)
Let's say you deploy 10 AI agents across your organization. Each one costs $500/month in API calls and tooling. That's $5,000/month — cheap.
But if only 3 of the 10 perform consistently, and each of the 7 underperformers requires 10 hours/month of human supervision at $75/hour, that's $5,250/month in supervision costs. Your $5,000 agent fleet now costs $10,250/month. And the 7 underperformers are still producing lower-quality work than the humans they were supposed to augment.
The math only works when you have:
- Agents with proven, measured reliability (not just demo performance)
- Structured evaluation that catches failures before they compound
- Clear task boundaries that match what agents can actually do well
Most companies have none of these. They have a ChatGPT subscription and a prayer.
The Bottom Line
Sol Rashidi didn't fire her agents because she's anti-AI. She fired them because she's pro-results. When managing an AI agent costs more than managing a human employee, the agent goes. That's not a technology failure — it's a management failure.
The companies winning with agents aren't the ones deploying the most. They're the ones who treat agents like employees: evaluate them rigorously, measure their output, and fire the ones that don't perform. The rest are paying for expensive interns they're too afraid to let go.
If the woman who launched IBM Watson thinks half her agents aren't worth keeping, what makes you think yours are?
Sources: Inc — "It Was Faster to Train a Grad" · Cisco AI Agent Reliability Study 2026