A thread from Vas Rao, CEO of enterprise AI consultancy Varick Agents, crossed 186,000+ views this week with an argument that's uncomfortable precisely because it isn't new: most of the money enterprises are spending on AI is being wasted on the same mistake companies made with computers in the 1990s. The thread pairs a 36-year-old Harvard Business Review quote with a UK government trial result that reads like it was written to prove the quote still holds, then offers a concrete framework for what to do instead.
Here's the argument, the precedent it leans on, the framework at its center, and where we think it holds up versus where a consultancy's own thread deserves a skeptical read.

TL;DR
| Question | Answer |
|---|---|
| What's the core claim? | Enterprises spend fortunes applying AI to unchanged, inefficient processes — the result is a faster bad process, not a better one. |
| What's the 1990 precedent? | Michael Hammer's Harvard Business Review argument that IT spending underdelivered because companies mechanized old processes instead of redesigning them — "paving the cow paths." |
| What's the headline evidence? | A UK government Microsoft Copilot trial: 1,000 licenses, 1.14 actions/user/day, no measured productivity gain, yet 72% user satisfaction. |
| What's the fix the thread proposes? | Map the process first (happy path, exceptions, systems of record, touch time vs. elapsed time), then sort every step into deterministic software, an AI judgment call, or a human-in-the-loop checkpoint. |
| What's the strongest evidence? | The UK government's own published trial report — independently checkable. |
| What's the weakest evidence? | Varick's own client results — self-reported, unnamed companies, no independent audit. |
| Does this apply below enterprise scale? | Partially — the "time is in the queues, not the work" observation and the three-bucket sort both generalize; the $500M-revenue engagement model obviously doesn't. |
The 1990 precedent, and why it's the strongest part of the thread
The most effective move in the whole piece is opening with someone else's decades-old words instead of a new claim. Michael Hammer, a former MIT computer science professor, wrote in a 1990 Harvard Business Review article:
"Heavy investments in information technology have delivered disappointing results, largely because companies tend to use technology to mechanize old ways of doing business. They leave the existing processes intact and use computers simply to speed them up."
Swap "information technology" for "AI" and the sentence needs no other edit to describe 2026. Hammer's own sharpest example, cited in the thread: an insurance application that took 22 days to move through a business, of which the total human working time inside those 22 days was 17 minutes. Even a tool that makes every step twice as fast only recovers roughly 8 minutes — a number someone can honestly report as "50% faster" while the actual bottleneck (queues, handoffs, waiting on another team) goes completely untouched.
That distinction — work time versus elapsed time — is doing almost all the argumentative work here, and it's genuinely useful independent of anything else in the thread. It's a concrete, checkable idea: before optimizing a process with AI, measure how much of its total duration is actual work versus queueing, because AI aimed at the wrong half of that ratio can make your metrics look better while changing nothing that matters.
The UK Copilot trial: the 2026 evidence that echoes 1990 almost exactly
The thread's headline data point is a UK Department for Business and Trade trial: 1,000 Microsoft 365 Copilot licenses, rolled out over three months, producing 1.14 Copilot actions per user per day. PowerPoint generation got faster (18 minutes down to 11) but at lower quality; Excel analysis got measurably slower, also at lower quality; email time savings were, in the department's own words as quoted in the thread, "extremely small." The department's stated conclusion: it did not find robust evidence that time savings translated into improved productivity. And yet 72% of users reported being satisfied or very satisfied — people liked using the tool while producing essentially the same output as before.
This is the one figure in the whole thread worth independently verifying, because it's attributed to a specific government department's own published trial rather than to Varick's private client work — and it's the kind of number that should make anyone budgeting AI seats for a large team pause before treating "seats deployed" or "usage satisfaction" as a proxy for ROI.
The framework: sort every step before you build anything
The thread's practical contribution is a three-way sort for any process step, and it's the part most transferable outside a $500M-revenue consulting engagement:
| Bucket | When to use it | Why |
|---|---|---|
| Deterministic | A fixed rule with no judgment involved (amount under a threshold + matches PO → pay it) | Plain code is cheaper, auditable, and never hallucinates — an LLM call is the wrong tool here |
| Agentic | A judgment call with many past examples and low enough risk to measure (route-to-team, flag/don't-flag, match/no-match) | Enough historical signal exists to both automate the decision and measure whether the agent is getting it right |
| Human-in-the-loop | Too risky or too novel to automate, but evidence-gathering can still be automated | The agent assembles everything a human needs to see, cutting a 40-minute investigation down to a 30-second decision — the human keeps the final call |
The thread's own example — a global reconciliation process across 300+ bank accounts, reportedly redesigned from an 11-step, multi-region, month-long process down to three steps (automated statement ingestion, rule-based-plus-agent-assisted matching, and a human exception queue) — is the clearest illustration of what "collapsing a process" looks like when you apply this sort correctly rather than just wrapping the existing 11 steps in an AI layer.
Where this connects to a framework explainx.ai has covered from a different angle: Boris Cherny's Steps of AI Adoption maturity model maps how far an organization has progressed toward AI-native operation; this three-bucket sort is closer to a per-workflow design tool for getting there. They answer different questions — one is "how mature are we," the other is "what should this specific step actually be" — and a real transformation program needs both.
Where we think the argument is genuinely useful
- The work-time-vs-elapsed-time distinction is real and underused. It's a five-minute exercise any team can run on their own workflows before reaching for an agent, and it will correctly kill a lot of proposed automations that target the wrong 17 minutes.
- The "don't consolidate your ERP first" point is sound and well-evidenced. The thread's TSB (£318M migration, £48.65M regulatory fine, 5.2 million customers affected) and Zimmer Biomet ($172M claim against Deloitte) examples are real, publicly reported incidents, and the underlying point — agents tolerate disagreeing systems of record in a way deterministic integrations never could — is a legitimate argument for orchestrating across existing systems rather than betting a transformation on a multi-year data-platform consolidation first.
- The three-bucket sort is genuinely reusable at any scale, not just the enterprise engagements Varick sells. It's a good discipline check for anyone building an agent: are you giving it a decision it should make, a decision it should just execute deterministically, or a decision it should prepare but not make?
Where the thread deserves a skeptical read
This is, transparently, a sales pitch from the CEO of a company that sells exactly the service the thread recommends — and that context matters for how to weigh two specific claims:
- The client results are unverifiable as presented. Month-end close "18-22 days to 7-9," AP exceptions "600-800 to under 50," "80%+ of the work handled entirely by an agent," over "$100M in aggregate value" — none of these name a client, publish a methodology, or come with third-party audit. That doesn't make them false; it makes them exactly as reliable as any consultancy's own case-study numbers, which is to say: a claim, not a finding, until independently checked.
- The framing of competitors is one-sided by design. The thread's critique of traditional consultancies (good at slides, bad at model selection) and AI labs (good at agents, bad at process work) conveniently describes a gap only Varick's stated positioning fills. The gap itself — process expertise plus AI-engineering expertise rarely sitting in the same room — is a real, widely observed problem in the industry; the claim that one specific firm has uniquely solved it is the part to treat as marketing.
None of this means the underlying diagnosis is wrong. The UK Copilot trial numbers didn't come from Varick, and Hammer's 1990 precedent didn't either — those two pieces of evidence stand on their own regardless of who's citing them in service of a sales thread.
What this means if you're building agents, not running a transformation program
The scale is different, but three things transfer directly to anyone shipping agents for a team, a startup, or a single workflow rather than a Fortune 500:
- Measure elapsed time, not just work time, before you automate. If the actual computation or generation step is a small fraction of a workflow's total duration, speeding up that step is the 8-minutes-out-of-22-days trap. Look for the queue first.
- Sort before you build. For each step in the workflow you're automating, decide explicitly: is this a rule, a judgment call worth measuring, or a decision that needs a human with better evidence in front of them? Building an agent for a step that should have been a five-line
ifstatement is expensive, harder to audit, and no more reliable. - Don't treat usage or satisfaction as a proxy for value. The UK trial's 72% satisfaction number next to "no robust evidence of productivity improvement" is the cleanest illustration available of why adoption metrics and outcome metrics are not the same thing — a lesson that applies whether you're rolling out Copilot to 1,000 people or shipping one agent to your own team.
For teams actually managing agent costs and reliability at scale rather than redesigning enterprise workflows, Uber's cost-optimization playbook and Databricks' four cost levers are useful companion reading — both are examples of the "measure the actual thing that matters" discipline this thread argues for, applied to spend rather than process design.
Honest limitations
- This post analyzes a single viral thread, not an independently verified case study. Varick's own client numbers are presented as the company describes them; we have not audited them.
- The UK Copilot trial description comes secondhand, via the thread, not from us reading the department's original published report directly — treat the specific figures as reported-by-the-thread pending independent verification of the primary document.
- "Applied AI doesn't work" is itself a marketing framing for a specific consultancy's service offering. The diagnosis (process before tooling) is well-supported by the historical precedent and the Copilot trial; the prescription (hire a firm like Varick specifically) is the part that's self-interested by construction, and readers should weigh it accordingly.
- The three-bucket framework isn't unique to this thread. Deterministic/agentic/human-in-the-loop sorting is a increasingly common pattern across agent-harness design more broadly — see explainx.ai's multi-agent orchestration patterns guide for the same underlying idea from a systems-design angle rather than a business-process one.
Related on explainx.ai
- Boris Cherny's Steps of AI Adoption — Claude Code's 0-4 maturity model
- How Uber runs coding agents cost-effectively at scale
- Databricks on managing AI coding costs at scale — 4 cost levers
- Multi-agent orchestration patterns — complete guide
- Agent harness design: DAG planner, worker, critic under budget pressure
- Forward Deployed Everything — how every role is becoming customer-embedded
- What running an AI agent actually costs per month
Primary source: Vas Rao (@vasuman) on X, September 5, 2026.
This post reflects a public X thread as of September 5, 2026, and explainx.ai's editorial analysis of it. Client-specific figures are as stated by Varick Agents and have not been independently audited by explainx.ai; the UK government Copilot trial figures are as characterized in the thread and should be checked against the department's own published report for exact wording.
