# AI Pricing Models in 2026: The AI Price Ladder from Tokens to Outcomes

> AI pricing models in 2026, explained step by step: token prices, tariff mechanics, cost per solved task, outcome pricing. With market data and examples.

- Canonical URL: https://olivergausmann.com/en/insights/ki-preismodelle-preistreppe
- Published: 2026-08-10 · Updated: 2026-08-10
- Author: Dr. Oliver Gausmann — https://olivergausmann.com/en/author/oliver-gausmann

## At a glance

- Why token list prices mislead in 32 percent of model comparisons and cost per solved task is the honest unit.
- How the five steps of the AI price ladder shift risk from buyer to vendor, all the way to verified outcomes on the rate card.
- Why you should negotiate the counting unit before the price and date every AI cost calculation.

## TL;DR

AI pricing models are being rebuilt in 2026 because AI, unlike classical software, burns real compute on every single use: AI product vendors now target median gross margins around 50 percent, well below the 80 percent SaaS is used to [1]. The AI price ladder sorts the field into five steps, from price per token to cost per solved task to paying for outcomes. List prices won't carry you far up that ladder, since in 32 percent of studied model pairs the cheaper-listed model ends up costing more overall [2]. Hybrid AI pricing models jumped from 25 to 37 percent of vendors within a year [1]. If you're buying, negotiate the counting unit first and the price second.

I maintain a cost calculator that tracks 75 AI models, and it keeps humbling me. Twice in two weeks my price data went stale without a single test firing: one model listed at five times its actual price, another vendor had added a second price tier for long inputs, noted only in a footnote under its price table. Then, on July 30, OpenAI cut the price of a three-week-old model by 80 percent [3].

Anyone buying or pricing AI is operating in a market where the price tags spin faster than any budget cycle. It looks chaotic, but there's an order to it, and it fits on five steps. I call it the AI price ladder: price per token, tariff mechanics, price per capability, cost per solved task, price per outcome. With every step, risk shifts a little further from buyer to vendor.

## Every AI answer costs money, which is why every pricing model is wobbling

A sold software license cost its vendor almost nothing to serve. Twenty years of subscription pricing rested on that marginal-cost logic. AI breaks it, because every answer consumes compute on expensive hardware, and that bill lands on the vendor every month.

The survey data now shows the fallout. AI product vendors target median gross margins around 50 percent, and only 12 percent believe the classic 80 percent is achievable [1]. A second, independently collected sample projects 52 percent for 2026 [4]. On top of that, 70 percent of vendors say their customers pay for AI out of existing software budgets [1]. The money itself is real enough: enterprises roughly tripled their generative AI spending to an estimated 37 billion dollars in 2025, according to Menlo Ventures [25]. The room for pricing mistakes is small, yet the industry is openly searching: 76 percent of software vendors have launched AI features, and most report revenue impact below 10 percent so far [5].

One concrete case will carry us up the ladder: a company with 10,000 customer inquiries a month that it wants AI to answer. Every step of the AI price ladder can be priced against that single workload.

## What does a token really cost?

The bottom step sounds reassuringly simple. Language models bill in tokens, small chunks of text, with list prices per million. Our 10,000 inquiries at roughly 3,700 tokens each cost about 37 dollars a month on a budget model, a figure straight from Anthropic's own pricing documentation [6]. So far, any intern can do this math.

The tariff behind it, step two of the ladder, has more moving parts. Recurring inputs can be cached; reading from that cache costs a tenth of the input price at Anthropic, but writing to it carries a surcharge of 25 to 100 percent [6]. At DeepSeek, a cache hit costs around one percent [7]. If you can wait, batch processing halves the bill [6]. For long inputs, several vendors add a second price tier, often noted only in a footnote under the price table: Google doubles the rate on Gemini 2.5 Pro above 200,000 tokens, retroactively for every token in the call [8], and OpenAI raises rates on several models above 272,000 tokens [9]. Then there's thinking: modern models generate invisible reasoning tokens that show up on the bill like any others, and whose volume swings by up to 9.7 times on the identical task [2]. The output side of your bill has become, quite literally, incalculable.

The third step of the ladder is price per capability, and it carries a price dynamic that runs in both directions at once. The price of a fixed capability level falls 5 to 10 times per year [10], and measured at specific performance thresholds, the median decline is 50 times per year [11]. Meanwhile the price of the frontier, the best available model, rises 3 to 18 times per year [10]. That's why companies experience both at once: yesterday's task keeps getting cheaper while the newest ambition keeps getting pricier.

## The honest unit is the solved task

On the fourth step, the question changes: what does it cost to get a task done? Research has had a plain formula for it since 2025: cost per solved task equals price per attempt divided by success rate [12]. A model that costs half as much but succeeds a third as often is the expensive one on this step.

A study published in March quantified how often list prices mislead: in 32 percent of model pairs, the cheaper-listed model causes higher total cost, by up to 28 times in the extreme [2]. One model listed 80 percent below its competitor came out 38 percent more expensive across all tasks [2]. The study attributes this to those swinging reasoning tokens plus extra working steps [2]. The list price fails as a price signal exactly where it promises the most.

For our 10,000 inquiries, the 37 dollars from earlier are only a good number if the budget model actually resolves them. If every fifth answer is unusable and a human cleans up, the real bill lives somewhere else. There's even a number for when the strongest model wins despite its price: once a single error costs around 10 cents, the frontier model almost always comes out ahead, per a Caltech analysis at mid-2025 prices [13]. That sounds low, but it follows from the proportions: what the frontier model costs extra per single request is small against what one avoided error saves. With customer inquiries, one lost customer clears that bar easily.

This step is why I built my cost calculator around exactly this unit, and I took the cost levers behind it apart in cutting AI costs through model choice.

## Which AI pricing models are winning in 2026?

Above the buying side, the vendor side of the AI price ladder begins, and it's been on the move for two years. In 2024, most software vendors priced AI features the way they'd always priced software, as a per-seat add-on. Then the margin math arrived, and the market has been testing new counting units in public ever since.

The American developer-tools market shows the migration in fast motion. GitHub moved Copilot to usage-based billing on June 1, 2026, replacing flat request allowances with token-based credits while keeping plan prices unchanged [15]. Cursor's June 2025 switch from request counts to a monthly compute budget went badly enough that the company publicly apologized and refunded surprise charges [16]. Salesforce runs two consumption models for its service agent side by side, 2 dollars per conversation or credits at 10 cents per standard action, and buyers have to do the math on which one is cheaper [14].

The fifth rung, the top of the ladder, is already on the rate card. Intercom charges 99 cents per outcome, counted as a confirmed resolution or a completed workflow [17]. HubSpot cut its price in April from 1 dollar per conversation to 50 cents per resolved conversation [18]. Zendesk goes furthest: a second AI model verifies which resolutions actually count before they're billed [19]. In Europe, SAP built a corporate answer, a prepaid currency called AI Units that AI features draw on across its whole product line; one Joule consultant package bundles 22,900 requests per user and month, pooled across all users [20]. For our 10,000 inquiries, two worlds now coexist: a two-digit dollar figure as a token bill, or several thousand dollars as an outcome bill once you scale 50 to 99 cents per resolved case (own calculation). The spread between them is the vendor's margin, risk and verification cost. The ladder sorts the logic; in the market it condenses into four basic types of pricing models.

The four AI pricing models compared
Criterion | Per seat | Usage (tokens, credits) | Per outcome | Hybrid (base plus component)
Counting unit | Workplace | Tokens, actions, credits | Resolved case | Base fee plus usage or outcome
Cost risk sits with | Vendor | Buyer | Vendor | Shared
Predictability for the buyer | High | Low | Medium | Medium to high
Typical dispute | Unused licenses | Usage spikes | Who verifies the outcome? | Contract complexity
2026 example | Microsoft 365 Copilot | GitHub Copilot, SAP AI Units, Salesforce Agentforce | Intercom, HubSpot, Zendesk | Most common type: 37% of B2B SaaS vendors

The adoption numbers match the story. Among roughly 300 surveyed AI companies, 58 percent bill by subscription, 35 percent by usage and 18 percent by outcome, with multiple answers allowed [4]. Hybrid pricing jumped from 25 to 37 percent of 230 B2B software vendors within twelve months, making it the most common model [1]. German-speaking Europe shows the same break: classic SaaS there sells 92 percent by subscription, but for AI products the share of usage-based pricing leaps to 69 percent [21]. And 45 percent of software vendors already run two or more pricing models in parallel [5]. Behind all of it sits an old pricing rule with new teeth: put the price on the unit that grows with customer value.

I traced what this does to software contracts, including the double-payment trap of licenses plus credits, in SaaS pricing in the agent era, and the professional-services side in consulting fees in the AI era.

## Three out of four requests don't need a frontier model

One lever runs across every step of the AI price ladder, and it only becomes visible once companies combine models. The average enterprise already uses 3.1 model providers [4], mostly without a system. Yet the sorting logic works like a mailroom: nobody sends every letter by express courier, and of our 10,000 inquiries, most are routine and few are delicate.

Berkeley's RouteLLM work measured what that sorting is worth: a trained router sent only about one in five requests to the expensive model, versus every second one under random assignment, while holding 95 percent of the top model's quality [22]. In a production workload, a calibrated router cut raw model costs by 31 percent at stable quality [23]. As an order of magnitude: a third to two thirds of cost, always relative to a declared quality bar.

Two caveats keep this honest. First, many off-the-shelf routers, commercial ones included, fail to beat a simple baseline, as a benchmark across 400,000 test cases showed [24]. Second, the error-cost threshold from the previous step applies here too: once a mistake costs real money, the single frontier model beats any combination [13]. And for agents that work autonomously across multiple steps, model choice already happens per step, which turns combination into an architecture question.

Negotiate the counting unit before the price, and budget in cost per solved task. Date every price snapshot; AI prices go stale within weeks.

## How do you bring order into your AI pricing?

- Date every price snapshot. Write the retrieval date of vendor price pages onto every AI calculation, and treat any undated number as stale. My own calculator went stale twice in two weeks despite daily use, so this catches people who do it for a living.

- Measure cost per solved task for one week. Run a real workload on a cheap and a strong model in parallel and count how many results are usable without rework. The AI cost calculator does the math behind it, including retries, caching and human cleanup.

- Negotiate the counting unit before the price. As a buyer, ask what exactly the vendor counts, who verifies a resolved outcome, and what happens above the quota. As a vendor, pick the unit that grows with your customer's value, and enter hybrid, base fee plus variable component; the hard cut can come later.

- Define your error-cost threshold. Put a number on what a wrong result costs per task type. Above the threshold, the best available model does the work; below it, the cheapest model that holds your quality bar.

One word on the switch itself, because it's communication work. Cursor showed in 2025 how a defensible price change turns into reputation damage when customers first notice it on the invoice [16]. If you change your AI pricing model, announce it, publish worked examples, and give existing customers time.

## My Take

The number I keep coming back to is the 37 percent of AI companies planning to change their pricing model within twelve months [4]. That's a market admitting, in public, that every current price tag is a draft. If I were signing an AI contract this quarter, I'd treat it that way: short term, defined exit, an adjustment clause both sides can live with.

The next battleground I see is the meter itself. Once vendors charge per verified outcome, whoever controls the verification controls the revenue, which is why a second model counting resolutions strikes me as more consequential than any single price cut. I'd expect buyers to demand audit rights on that meter the way they once demanded audit rights on license counts. Maybe that takes longer than I think, procurement habits are slow to move.

What I can report first-hand: in late July I benchmarked five popular AI cost calculators (own survey), and none of them could show cost per solved task. Everyone counts tokens. The market argues about the top rung of the AI price ladder with tools that still stand on the bottom one, and whether we'll even talk about tokens in five years... I'll leave that open.

## Sources

1. Growth Unhinged: The 2026 State of B2B SaaS and AI Monetization (230 Unternehmen, April/Mai 2026) — https://www.growthunhinged.com/p/the-state-of-b2b-monetization-in-2026
2. Chen et al.: The Price Reversal Phenomenon (arXiv, März 2026) — https://arxiv.org/abs/2603.23971
3. CNBC: OpenAI cuts API prices, 30.07.2026 — https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html
4. ICONIQ Growth: State of AI 2026, Bi-Annual Snapshot (rund 300 Führungskräfte) — https://www.iconiq.com/growth/reports/2026-state-of-ai-bi-annual-snapshot
5. Simon-Kucher: Global Software Study 2025 — https://www.simon-kucher.com/en/who-we-are/newsroom/software-leaders-brace-ai-shake-while-chasing-growth
6. Anthropic: Preisdokumentation (Cache, Batch, Referenzrechnung; Abruf 10.08.2026) — https://platform.claude.com/docs/en/about-claude/pricing
7. DeepSeek: API-Preisdokumentation (Cache-Treffer; Abruf 10.08.2026) — https://api-docs.deepseek.com/quick_start/pricing
8. Google: Gemini-API-Preisseite (Kontextstaffel ab 200.000 Token; Abruf 10.08.2026) — https://ai.google.dev/gemini-api/docs/pricing
9. OpenAI: API-Preisseite (Kontextstaffeln ab 272.000 Token; Abruf 10.08.2026) — https://developers.openai.com/api/docs/pricing
10. Gundlach et al.: The Price of Progress (MIT, arXiv, März 2026) — https://arxiv.org/abs/2511.23455
11. Epoch AI: LLM inference prices have fallen rapidly but unequally across tasks (Abruf 08/2026) — https://epoch.ai/data-insights/llm-inference-price-trends
12. Erol et al.: Cost-of-Pass, An Economic Framework for Evaluating Language Models (arXiv) — https://arxiv.org/abs/2504.13359
13. Zellinger, Thomson: Economic Evaluation of LLMs (Caltech, arXiv, Juli 2025) — https://arxiv.org/abs/2507.03834
14. Salesforce-Hilfe: Agentforce-Preismodelle (Conversations und Flex Credits) — https://help.salesforce.com/s/articleView?id=004811240&language=en_US&type=1
15. GitHub Blog: Copilot is moving to usage-based billing (Umstellung zum 01.06.2026) — https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/
16. Cursor Blog: June 2025 pricing (Umstellung, Entschuldigung, Rückerstattungen) — https://cursor.com/blog/june-2025-pricing
17. Intercom: Preisseite Fin (0,99 USD je Outcome; Abruf 10.08.2026) — https://www.intercom.com/pricing
18. HubSpot Company News: Now you pay when the task is complete (14.04.2026) — https://www.hubspot.com/company-news/hubspots-customer-agent-and-prospecting-agent-now-you-pay-when-the-task-is-complete
19. Zendesk Newsroom: Relate 2026, verifizierte Resolutions (19.05.2026) — https://www.zendesk.com/newsroom/press-releases/relate-2026/
20. SAP Learning: Evaluating the Commercial Model (Joule, AI Units) — https://learning.sap.com/courses/introducing-sap-joule-for-consultants/evaluating-the-commercial-model
21. hy × OMR Reviews: SaaS & AI Pricing Report 2026 (4.000+ Profile, 180 Befragte) — https://pricing.hy.co/
22. Ong et al.: RouteLLM, Learning to Route LLMs with Preference Data (ICLR 2025) — https://arxiv.org/abs/2406.18665
23. UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing (arXiv, Mai 2026) — https://arxiv.org/abs/2605.18796
24. LLMRouterBench (arXiv, Januar 2026) — https://arxiv.org/abs/2601.07206
25. Menlo Ventures: 2025 The State of Generative AI in the Enterprise (November 2025) — https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/

## FAQ

### What is the AI price ladder?

A five-step model for sorting AI pricing models: price per token, tariff mechanics (caching, batch, context tiers), price per capability, cost per solved task, and price per outcome. With every step, cost risk shifts further from buyer to vendor.

### What AI pricing models exist in 2026?

Four basic types: per seat, usage-based via tokens or credits, per outcome, and hybrid models combining a base fee with a variable component. Hybrids are now the most common at 37 percent of vendors, and outcome prices have reached the rate card at Intercom, HubSpot and Zendesk.

### Why is the cheapest AI model often not the most economical?

Because total cost depends on success rates and reasoning tokens. In 32 percent of studied model pairs, the cheaper-listed model causes higher total cost, as retries and longer thinking eat up the list-price advantage.

### What is outcome-based pricing for AI agents?

A pricing model where only the verified result gets paid, such as a resolved customer issue. Intercom charges 99 cents per outcome, HubSpot 50 cents per resolved conversation, and Zendesk has a second AI model confirm resolutions before they're billed.
