Zendesk and others are moving from per-seat subscriptions to token-based pricing, because AI’s per-token marginal cost breaks the flat-rate model. Local ‘AI PCs’ are the hedge — fixed cost, near-zero marginal. But the real lesson is one layer down: whoever holds the meter holds the power, and owning your stack is how you escape it.
Zendesk is moving from per-seat subscriptions to token-based pricing, and it’s the visible edge of an industry shift. The stated reason is noble. The real reason is arithmetic: when your product runs on a model that charges per token, a flat per-seat price is a promise to lose money on your best customers.
Zendesk is moving from per-seat subscriptions to token-based, consumption pricing, and it isn’t alone — it’s the visible edge of a shift running through enterprise software. The stated reason is almost noble: “We believe software value should align directly with customer success, not headcount.” The real reason is arithmetic. When your product runs on an AI model that charges you per token, a flat per-seat price is a promise to lose money on your best customers.
Why subscriptions break for AI software
The per-seat SaaS subscription was one of the great business-model inventions of the last two decades. It worked because the marginal cost of one more user doing one more thing was essentially zero. Software scaled; the bill didn’t.
AI breaks that assumption at the root. Every summarisation, every draft, every code completion, every data extraction consumes tokens the vendor pays a cloud provider for. Marginal cost is no longer zero — it’s a real, metered, per-action expense. A heavy user on a flat per-seat plan can now cost more to serve than they pay. Flat-rate AI subscriptions were, in effect, loss leaders to win early adopters; now that enterprises depend on the tools, the vendors need the revenue to actually cover the compute.
So the model is migrating from “pay for access” to “pay for usage,” because the vendor’s own costs are usage-based and someone has to absorb the variability. Increasingly, that someone is the customer.
The hedge: bring the compute in-house
The counter-move is the interesting part. If cloud AI charges per token, the way to escape the meter is to stop using the cloud for routine work. Local models running on a neural processing unit — an “AI PC,” or a Mac Mini quietly running an open model in a cupboard — carry a fixed hardware cost and then a marginal cost of roughly zero. For a $1,500 machine under heavy daily use, break-even against per-query cloud charges can arrive in months.
That’s a genuine strategic fork for anyone deploying AI at volume: rent it per-token and keep costs variable, or own the compute and make them fixed. The right answer depends on your usage pattern — but the mere existence of a credible local option changes the negotiation, because “we’ll run it ourselves” is now a real threat rather than a bluff.
The part nobody selling you software wants to discuss
Here’s the uncomfortable question underneath both moves. When a vendor switches you to consumption pricing, your costs become unpredictable and rise with your success — you’re penalised precisely when the tool is working. And when you consider bringing compute in-house to escape that, you discover how little of your own stack you actually control. Your data, your workflows, your customer relationships all live inside a product priced by someone whose interests just diverged from yours.
The deeper issue isn’t the pricing model. It’s ownership. Consumption pricing is only frightening when the meter belongs to someone else and you can’t leave. The businesses least exposed to this shift are the ones that own their stack — their application, their data, and increasingly their inference — so that a cloud vendor’s pricing decision is a line item to optimise rather than an existential surprise.
This is a large part of the case for self-hosted, source-available platforms. A system like VBWD is built on exactly that premise: you run the application yourself, the subscription and entitlement billing is yours to configure (per-seat, tiered, usage-based, whatever fits — because it’s your code, not a vendor’s price sheet), and its central LLM connection lets you point AI features at a hosted model or one you run locally. When the market swings between subscription and consumption pricing, owning the layer means you get to choose which one you charge and which one you pay — rather than absorbing whichever a vendor imposes. The honest caveat is the usual one: owning the stack means running it, which is real operational work and not right for every team. But when your software costs just became variable and someone else holds the meter, the value of holding your own shifts sharply.
The read
The move from subscription to consumption pricing is a rational response to AI’s real, per-token marginal cost — vendors can’t keep subsidising heavy users on flat plans, and they won’t. For buyers it means AI costs that scale with success and are harder to forecast, which is uncomfortable precisely for the customers getting the most value.
The local-compute hedge is real and worth evaluating on the numbers. But the durable lesson is one layer down: pricing models will keep swinging as the economics of AI settle, and the businesses that ride those swings comfortably are the ones that own their stack rather than rent it. Whether you pay per seat or per token matters far less than whether you can walk away from the meter. Most companies are about to find out they can’t.
Analysis based on reporting on AI pricing shifts as covered on 28 July 2026. Vendor pricing details and cost figures are as reported and vary by product and usage. Commentary, not financial or purchasing advice. The VBWD reference illustrates the own-your-stack approach and is not an endorsement.
Learn more about VBWD
VBWD is a self-hosted, source-available platform for building subscription products, marketplaces, and AI-powered apps. Explore it further:
- 🌐 Website and documentation: vbwd.cc — see the plugins, architecture, and developer docs.
- 💻 Source code and plugins on GitHub: github.com/VBWD-platform
- 🎥 Watch VBWD in action: demo video 1 and demo video 2
- 💼 Follow the project on LinkedIn: linkedin.com/company/vbwd
VBWD is source-available — get the SDK on GitHub.