Tuesday. 9:30. Manolo walks in pushing a box on wheels. It takes him longer to get through the door than to say hello.
—Jose, I've got it. I'm canceling OpenAI. I bought this. Fifty thousand a month in savings.
It's a Mac Studio M3 Ultra. The beast version. The one that costs as much as a car. They don't sell them anymore, but he's managed to get a refurbished one for more than it cost new.
—Manolo, how much did this cost you?
—Twelve thousand euros, less than a Dacia Sandero, and think of the savings. Fifty thousand a month, Jose.
—Manolo, those fifty thousand a month, what do they get spent on?
—On AI.
—Be specific.
—On people using it.
—And that's fifty thousand a month?
—No. That's more like ten thousand.
—And the other forty thousand?
Manolo goes pale. He looks at the box as if searching for help.
—Well, the agents stuff. The thing that classifies invoices. The thing that watches the logs. The thing that processes the emails.
—And that runs on its own?
—Twenty-four hours.
—Bingo, Manolo. You've just discovered the only distinction that matters. And it's not OpenAI versus your Mac Studio. It's the human bill versus the automated bill. Two different businesses. And you've spent six months treating them as if they were the same one.
Silence.
—What's the difference?
—The difference is that you use them for different things and, conveniently, Anthropic and OpenAI sell them to you under different pricing models. If you don't understand that, you get everything wrong. Like now.
I grab the pen.
—First world. Your people. Programmers with Claude Code. Product with ChatGPT. This is interactive use. A person in front of a screen. For this, the right price isn't per token. It's per seat. Anthropic sells Team Premium at a hundred dollars per seat, annual contract. Includes Claude Code, integrations, admin controls. For your ninety engineers: nine thousand a month. Flat.
—Nine thousand? For ninety licenses?
—Nine thousand. Flat. And look at this: Anthropic says ninety percent of Claude Code users stay under thirty dollars a day. If you paid that via API, it would come to between five hundred and two thousand per engineer per month. But paying flat, you don't care.
—And why has nobody told me this?
—Because the guy selling you the Mac Studio earns a commission per unit. He sells you hardware. He doesn't get a cut of Anthropic's subscriptions.
—Damn.
—Wait, that's only the first world. Want to hear what Microsoft did this month?
—What?
—It canceled Claude Code in its big divisions —Windows, Outlook, Teams, Surface— because the variable bill was hitting two thousand dollars per engineer per month. Guess what they switched to?
—Local AI?
—No, Manolo. GitHub Copilot Enterprise. Thirty-nine dollars per seat. Flat. Centralization, not decentralization. Subscription, not home-office hardware.
Manolo looks at the box. The box looks back at him.
—And the second world?
—Your remaining forty thousand. The agents. The stuff that runs on its own. Here the flat plans are useless, because they're designed for humans, not for an agent burning tokens twenty-four hours a day. If you hook an agent up to a Max account, they shut it down by day three. Here, yes, API. And this is where the money goes.
—And can it come down?
—A lot. Prompt caching: ninety percent off on whatever repeats. Batch API: fifty percent off on anything asynchronous. Haiku instead of Opus wherever nobody will notice —twenty-five times cheaper—. This is what got away from Uber this year when it burned through its entire annual AI budget in four months. They had to put in a cap of fifteen hundred dollars per employee.
—Damn.
—And finally, Manolo, we have the local AI thing. Which does exist, does make sense, but is not what you've bought.
—What?
—There's an excellent case for local AI: anonymizing prompts before sending them to OpenAI. Detecting names, account numbers, medical data, replacing them with placeholders, sending the clean version to the API, and restoring the originals on the way back. This is worth gold for compliance. It's worth so much that one day it will save you from a fine.
—And the Mac Studio helps me with that?
—Here's the subtlest mistake. If you put the anonymization filter on one central Mac Studio, you've created three problems. One: a bottleneck. Every prompt from every employee going through the same machine. Two: added latency. Every call to Claude gets delayed waiting for your filter. Three: a single point of failure. The day it goes down, there's no AI in the company.
—So what then?
—The right answer is the opposite. Distributed, not centralized. Every employee carries a Qwen 7B or a Llama 3.2 8B on their own MacBook. Those models run perfectly fine on 36 or 48 GB. Local anonymization, on the user's machine, no added latency, no bottleneck, no single point of failure. And the additional cost, zero, because the MacBooks are already paid for.
—Damn.
—Yes, Manolo. You've spent twelve thousand euros on a Mac Studio when the right solution was sitting in the laptops your people already had on their desks.
Long silence.
—And the Mac Studio?
—For what it's actually good for: nightly batch agents. Classifying invoices at three in the morning. Processing logs. Things that can wait and benefit from a dedicated workhorse. You give Pablo the Mac Studio and ask him to move everything asynchronous onto it. We shave another three or four thousand a month off the API.
—Noted.
—And tomorrow you talk to Anthropic. Team Premium for the engineers. You cancel whatever personal API keys they have open.
—Noted.
—And we roll out to every laptop the app we've built. The one that installs the local model and mounts the anonymization on top. Every employee, their own system, on their own machine.
—Noted.
—And the next time a PDF arrives promising to save you ninety-nine percent, you read it with this sentence in your head: the problem is not the vendor. It's not knowing which use gets which thing. And where each piece goes.
Manolo closes his notebook. Gets up. Pushes the Mac Studio box out the door with a little less pride than when he came in.
Twelve thousand euros, one masterclass, and the morning's single lesson:
> In AI, the most expensive thing isn't the model. It's not knowing what kind of problem you have.
I head out for a coffee.
Thanks for reading.
We all know a Manolo. Sometimes we are the Manolo. A new one every Tuesday. Subscribe so you don't miss it.
