The hidden cost of renting your AI

From the Desk of Cale · Cost, Fixed-cost AI · ~5 min read

Per-token pricing is a great way to start and a hard way to budget. The variable cost of someone else’s AI infrastructure isn’t actually variable to you. You just absorb the outputs: pricing changes, capacity limits, decisions made at a scale you’ll never see. Fixed, right-sized private infrastructure flips that. You know what you’re paying, and you know why.

The day-one price is not the real price

I’ve been thinking about taxi meters. A taxi is the right call for a trip to the airport, and nobody argues with the meter for one ride. But if you found yourself taking that same taxi to work every morning, you’d start doing math on a car payment before the end of the month. Per-token AI is a taxi meter, and a lot of companies are now commuting in it. Nobody chose the commute, either. The meter was already running when the habit formed.

The day-one price looks great because day-one usage is tiny. A handful of questions, a summarized document or two. Fractions of a cent each. Then the tool turns out to be useful, which is the whole point, and useful tools get used. People stop asking one question and start pasting in whole contracts. Someone wires it into a workflow that runs on every support ticket. Then come the agents, and an agent doesn’t make one call per task. It reads the document, checks its own work, calls a tool, reads the result, tries again. Every one of those steps is a metered ride. The per-token price never moved. Your consumption did, while everyone was busy being productive.

I sat with a client this spring and put a year of their AI invoices side by side on one screen. That was the entire exercise. No spreadsheet wizardry, no consultants. The line went one direction, and nobody in the room could name a month where anyone decided to spend more. That’s the tell. Metered costs don’t get decided. They accumulate.

Variable to them, fixed to you (and vice versa)

Here’s the part that took me embarrassingly long to see clearly. The provider’s costs are mostly fixed. The data centers are already built and the payroll is already set. What’s variable, to them, is you. Metered pricing is how they convert their fixed cost into your variable one. That’s a rational move on their side of the table. It’s just worth noticing which side of the table you’re on. When a business absorbs volatility, it charges for the service. When it passes volatility through, you’re the one providing that service, and nobody’s paying you for it.

So when their world shifts, the shift gets passed through. A new model generation lands and the price per token changes. Demand spikes and rate limits show up at exactly your busy hour. An older model gets retired, and the workflow your team spent a quarter tuning now runs on something that behaves differently. There’s no villain in any of that. A business planning for millions of customers makes ordinary capacity decisions, and you’re one line in the plan. You don’t get a vote. You get an email.

Now put yourself in the budget meeting. Finance asks what AI will cost next year. The honest answer under metered pricing is “it depends on how much people use it,” which is another way of saying the better it works, the less we can predict. That’s a strange incentive to hand a mid-sized team. Success becomes a cost overrun. I’ve watched a manager quietly discourage adoption of a tool his own company was paying for, because every new enthusiastic user made his forecast worse. Every budget is a guess, but this one is a guess about other people’s guesses.

What predictable looks like

Predictable starts with a boring question: what do you actually run? Not someday, today. Count the real workloads. The document review, the drafting, the internal search, the two or three automations that matter. Most teams find the list is shorter than they feared and steadier than the invoices implied. That steadiness is the asset. You size for your own team and the work it actually does. The whole internet is somebody else’s capacity problem. Once you know the workload, you can size the hardware to it, and once the hardware is sized, the cost is flat. You know the number in January and it’s still the number in October.

That’s the shape of what we build at Modular Technology Group. Our Private AI Workspaces start with Wildcat on shared infrastructure, step up to Panther on dedicated infrastructure, and top out at Grizzly on fully dedicated hardware, which you can host with us or stand up in your own building. The hosted tiers all live in a US-based, FedRAMP-certified data center. Every one of them bills as a flat monthly number. No per-token billing anywhere. We run our own work on the same stack, so when the meter argument comes up, we’re not speculating. We live on the fixed side of it. Right-sizing is a conversation. Some teams land on shared infrastructure and stay there happily. Some need the hardware where they can see it. Either way the number gets chosen, and you’re in the room when it happens.

The total-cost math is worth saying plainly. Owning can look more expensive on day one, the way a car payment looks worse than one cab fare. But a fixed cost changes the direction of every incentive after that. Under a meter, each new user is a liability. On infrastructure you own, each new user makes every task cheaper, because the same monthly number is now doing more work. You quit rationing the tool and start pushing it. Adoption stops being a cost problem and turns back into what it should have been, which is a productivity story. And because the models sit behind an interface you control, no single provider’s pricing decision can reach into your budget. When a better open model ships, it slots in. The bill doesn’t notice.

Here’s something you can do this week, and it costs nothing. Pull your last twelve months of AI invoices and put them in one place. Ask two questions. What happens to this line if usage doubles, and who decided the current number? If the answers are “it doubles” and “nobody,” you’re renting, and now you know what the rent really is. Then you get to decide what the number should be instead. Your data, your rules, from dirt to desktop.

If your AI line item has quit behaving, or you just want a second set of eyes on the own-versus-meter math, I’m always glad to compare notes. No pitch. Bring the invoices.

Own it, don’t rent it. Your data, your rules.