> ## Content Index
> Fetch the complete content index at: https://www.philipmward.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Maybe AI won't replace us all: the cost ceiling almost nobody's pricing in
- URL: https://www.philipmward.com/the-cost-ceiling/
- Published: 2026-08-19T17:00:00.000Z
- Updated: 2026-08-19T17:00:03.000Z
- Description: The "AI replaces every job" story quietly assumes inference is free and AI is reliable enough to trust without a human. Both are wrong, and the gap is a real cost ceiling.
- Author: Philip Ward
- Tags: AI, AI economics, future of work, token efficiency

My AI bill went from $100 a month to $200 a day.

Same me, same work. My company moved me onto an enterprise plan, which meant usage-based billing, and it turns out I use a lot. The flat-rate math broke almost immediately. So they pivoted me and the other power users back onto a Team plan, where we still get the old $100-a-month pricing.

There's a catch. A Team plan caps at 150 seats (Anthropic's published limit: minimum 5, maximum 150; above that you're on Enterprise). A handful of heavy users at a small shop can duck onto the flat rate. A company with three thousand engineers who all now live inside an AI coding tool cannot. They pay for what they burn.

I run the cost side of a production AI system for a living. And the longer I watch that side of the house, the *less* worried I am that AI is about to replace everyone. Not because the models can't do the work. Because of what the work costs the moment you need it done reliably, autonomously, at scale.

## What the "AI takes all the jobs" story gets wrong

The replace-everyone story runs on two assumptions it never says out loud: that inference is basically free, and that AI is reliable enough to trust without a human watching. Both are wrong. Pull either one and the story turns from "AI does your job" into "AI helps with your job," which is a completely different economy.

That's the argument. The rest is showing why each assumption breaks, and being honest about the ways it might un-break. I won't give you a date or a number for when the ceiling bites, because nobody honestly can. This is a moving contest, not a countdown.

## Isn't AI inference basically free now?

Per token, it's astonishingly cheap and getting cheaper. That part is real. Epoch AI tracks the cost of a given level of capability falling more than tenfold since early 2023, and still falling fast. If cost were only about price-per-token, there'd be no ceiling to talk about.

But a token is the wrong unit. You don't buy tokens, you buy *finished tasks*, and finished tasks got hungrier faster than tokens got cheaper.

Here's the mechanism most cost intuition misses. I wrote a whole post on it, the [context tax](https://www.philipmward.com/the-context-tax-why-your-long-running-ai-agent-gets-expensive-and-the-one-pattern-that-fixes-it/): a long-running agent re-pays for its entire context on every single turn. Ask it to do ten steps and it re-reads everything it knows ten times over. A chatbot answering one question is one cheap call. An agent doing your actual job is that call, paid again and again, for hours.

The numbers move the way you'd expect. Reasoning models and agent workflows burn 5–30× the tokens of a plain chatbot query. By one industry estimate, enterprise AI spend rose about 320% in 2025 even as per-token prices kept falling: the demand curve simply outran the price curve. Cheaper tokens, bigger bills.

![Schematic line chart titled "Cheaper tokens, bigger bills." A green line, price per token, falls steeply from early 2023 to today (labeled minus 94% since early 2023, per Epoch AI). An amber line, your total AI bill, rises just as steeply over the same period (labeled usage up 5 to 30 times per agent task). The two cross in the middle, annotated "unit price falls, the bill still climbs." Marked illustrative and schematic — two different trends, not one dataset; price from Epoch AI, spend an industry estimate.](https://storage.ghost.io/c/67/8c/678c88a8-eaf2-4b46-8968-b5902985493b/content/images/2026/08/chart-the-cost-ceiling-cost-paradox.png)

So the honest version of "inference is cheap" is narrower than it sounds. The cheap mode is cheap. The autonomous, long-horizon, do-the-whole-job mode is the context tax paid on a loop, and that is exactly the mode you need to *replace* a person instead of assist one.

## Why does reliable automation cost so much more than the demo?

Because the last mile, the part where you actually remove the human, is where the cost lives. And it scales with two things: how complex the work is, and how much a mistake costs.

Low stakes, low complexity: cheap. Let AI draft a first-pass FAQ reply and worst case it's a little off and someone fixes it. Fine. Push into work where being wrong is expensive, though, and two things happen at once. You need far more verification to trust the output, and the hard cases escalate to a human anyway. Now you're paying for the AI and the person it was supposed to replace.

Klarna ran this experiment in public. In early 2024 its CEO announced the company's AI was doing the work of 700 support agents, resolving issues in under two minutes, on track to save $40 million. By May 2025 he'd walked it back: quality had dropped, and Klarna was hiring humans back into a hybrid model. The bot was fine on routine questions and bad on the high-stakes tail of disputes, fraud, and hardship cases. The projected savings got eaten by cleanup and rehiring.

I pay this tax on purpose. My own system doesn't trust its own output. It runs a second agent to audit the first one's work, and a third to judge the audit, before anything ships. In one measured run that audit loop re-paid the same evidence to its sub-agents about six times over across three rounds (measured in our setup, on a single skill's executor thread, not a controlled benchmark). That redundancy is the price of trusting an autonomous result. It's worth paying. It also isn't free, and it's the part the demo never shows you.

And the productivity gain that's supposed to justify all of this is shakier than the pitch. When METR ran a randomized controlled trial on experienced open-source developers in 2025, the ones using AI were 19% *slower*, while believing they'd been 20% faster. If the real uplift on genuinely hard work is that uncertain, "AI will just do the whole job" is standing on soft ground.

## Is compute even the real limit?

Push past price and reliability and you hit a wall that ignores both: power.

The bottleneck on AI right now is power. Not the chips — the grid to run them. Interconnection queues in the big US data-center markets run four to seven years, high-power transformers have multi-year lead times, and Gartner expects 40% of AI data centers to be power-constrained by 2027\. You can make a model ten times more efficient and still not get the megawatts hooked up in time. That's a supply ceiling, and it behaves nothing like a price ceiling: you can't discount your way past a transformer that doesn't exist.

This is where the cost story and the environmental story turn out to be one story. Every token is compute, compute is power and cooling water, and the physical build-out is the point where "cheaper" and "greener" stop being a tradeoff. (I make that whole case in [Cheaper Is Greener](https://www.philipmward.com/efficiency-sustainability-equation/).)

My own bet on the fix is nuclear, specifically small modular reactors. But betting and building are different sports, and the US is behind. China's Linglong-1 is on track to be the world's first commercial land-based SMR this year; America's first has no ground broken, and the enriched fuel it needs isn't produced here at scale yet. You can't add national-scale power in a quarter.

## But won't falling prices just erase the ceiling?

This is the strongest counter, and it deserves a real answer.

Prices are collapsing, and the smart-money bet is that volume swamps everything. Goldman Sachs projects AI token demand rising 24× by 2030 and reads the falling cost-per-token as a coming "margin inflection," the point where the economics finally tip toward the providers. If they're right, the ceiling recedes faster than the work rises to meet it, and this post ages badly.

Maybe. Here's why I don't think it's settled. Price-per-token is falling, but tokens-per-job is rising, physical supply is capped for years, and there's a subsidy hiding in today's prices. AI providers mostly aren't making money on inference yet. Anthropic ran deeply negative gross margins in 2024 and is only now climbing back toward profitability, and both it and OpenAI have missed their own margin targets as inference costs ran hotter than forecast. When a market is priced below cost to win share, the sticker isn't the true cost. And the true cost is already surfacing: in April 2026 Anthropic pulled bundled tokens out of its enterprise seat deals and moved everyone to metered usage, which by licensing-analyst estimates can double or triple the bill for heavy users (The Register).

And the headline price misses something bigger: most people never paid per token at all. They paid a flat monthly subscription, the all-you-can-eat plan that made AI feel nearly free. It's the plan I opened this piece on, and coming off it is exactly why my own bill briefly leapt sixtyfold, well past the double-or-triple a metered enterprise seat sees. That era is closing everywhere. In 2026, Anthropic, GitHub Copilot, and Microsoft all moved from flat-rate toward metered billing, and even OpenAI, the most reluctant, started metering its agent usage. And even a launch that reads "cost-neutral" can climb underneath you: Claude Sonnet 5 launched at an introductory $2/$10 per million tokens, and on September 1 that reverts to the standard $3/$15, a clean 50% step-up, while its new tokenizer counts the same text as roughly 30% more tokens. The sticker still reads "unchanged." The bill doesn't. Anthropic is just the most transparent about a shift happening across the whole category, and it's the tool I run on.

So falling price pulls the ceiling up while rising usage, capped supply, un-hidden cost, and the death of the flat rate push it back down. Which force wins, and when, I genuinely don't know. That's the real shape of it.

## Who does the ceiling actually protect, and for how long?

Not everyone, and not evenly. This is where I'd temper any comfort you're taking from the argument.

The ceiling bites hardest where labor is expensive. To replace a $50-an-hour US knowledge worker, AI has to beat $50 an hour. To replace *everyone*, it also has to beat a roughly $3-an-hour worker somewhere cheaper, and that's a much lower bar for a human to keep clearing. (Those wage numbers are my read of the incentives, not a measured claim.) So the honest read is narrower than "AI won't replace jobs": cost decides *which* jobs go, and in what order. The expensive middle is the most exposed; the cheapest labor may never be worth automating at all.

An old MIT finding sharpens it. Svanberg and co-authors looked at computer-vision tasks and estimated that only about 23% of the wages paid for them would be cost-effective to automate today, because the all-in cost of the system beats the wage in only a quarter of cases. It's vision-specific, so don't read it as an LLM number. The principle travels, though: being technically automatable and being worth automating are very different things, and the gap between them is money.

The canary is already chirping. Uber's CTO told The Information the company burned through its entire 2026 AI budget in four months after rolling out an AI coding tool, at $150–250 per engineer per month, with power users running $500–2,000 and one of his own two-hour sessions costing $1,200\. When Careerminds surveyed 600 HR leaders who'd made AI-driven layoffs, only 8.4% would do it again unchanged, and nearly a third found rehiring cost *more* than the layoffs had saved. That's my own story at national scale: heavy use, sticker shock, pivot back to the cheaper arrangement. The difference is that I could pivot. Above 150 seats, you can't.

## So what do we do with the reprieve?

If the ceiling holds long enough to matter, it doesn't cancel the disruption. It slows it, and buys time to adapt. Which leaves the question I actually can't answer.

Big shifts usually arrive one generation at a time. That's a pattern in how change tends to land; it says nothing about how long this particular ceiling holds. The people most exposed are the ones who take their identity from their careers, and plenty of us do. The younger cohorts already ranking work lower read almost like the world pre-adapting to this. But the part I keep turning over is what people do with the freed-up time. The hopeful version is a renaissance of craft and care. My honest worry is that removed friction gets filled with the cheapest thing on offer, and the cheapest things are engineered to be addictive: the feeds, the games, the machines that pay out on a schedule. I hope I'm wrong. I'm genuinely not sure.

Either way, the move I'd stop making is pricing inference at zero and reliability at a hundred percent. The bill is real. I've seen mine.

## FAQ

**Is AI actually getting more expensive?**  
No. The price per token is falling fast: more than tenfold cheaper since early 2023, per Epoch AI. What's rising is the number of tokens a real task consumes. Reasoning models and autonomous agents burn 5–30× a simple chatbot call, and the flat-rate subscriptions that hid the cost are being replaced by metered billing, so total bills climb even as unit prices drop.

**Why is "AI will replace every job" an overstatement?**  
It assumes two things that aren't true: that inference is nearly free, and that AI is reliable enough to run without a human. Full replacement needs the expensive, autonomous, heavily-verified mode of AI, not the cheap single-call mode the assumption is built on.

**What's the difference between AI assisting and AI replacing, in cost terms?**  
Assisting is one cheap call with a human still in the loop. Replacing means paying for autonomy plus the verification and human-escalation needed to trust the output, and that cost climbs with how much a mistake would cost. The gap between the two is the ceiling.

**Will falling AI prices eventually remove the ceiling?**  
Maybe. Prices are dropping and demand is exploding: Goldman Sachs projects 24× token growth by 2030\. But tokens-per-task is rising, power supply is capped for years, and today's prices are partly subsidized below true cost. Which force wins is genuinely unsettled.