Tool
AI Cost Calculator
Requirement first, then the number: pick the case, set in the requirement profile what an error costs and who reviews it, and see the cost per solved task for 75 models, with a range and the comparison to the literature.
The article behind it: Cutting AI Costs Without Cutting QualityQuestion
What do you want to calculate?
Background and evidence for: Answering a support ticket
3,700 tokens per conversation from the Anthropic documentation; 88 percent of AI-using companies deploy AI in customer contact (Bitkom, 02/2026).
Typical failure modes: A wrong answer that binds the company. Then the volume boomerang: Commonwealth Bank reversed 45 redundancies because the voice bot increased call volumes. Klarna rolled back in 2025 citing quality, and was back at 853 full-time equivalents as of Q3 2025.
Tokens in the GPT-5 reference measure (English text); for German text the calculator applies a per-model factor of 1.46 to 2.6.
Requirements
What does the result have to withstand?
rework: The error binds human time: an escalation, a query, a correction. Assumed: €0.50 to €5.
Source: Market price per resolved support ticket 0.49 to 2.00 USD (Drag, MavenAGI 2026); Zellinger and Thomson (arXiv 2507.03834): from about 0.01 USD error cost the stronger model wins. As of 19 August 2026.
spot-checked: Humans check samples. Caution: in one study 35 to 45 percent of flawed AI drafts were sent unchanged. Assumed: 0.5 minutes per task.
Source: MedStar/Georgetown, npj Digital Medicine 24 April 2025 (automation bias). As of 19 August 2026.
paragraphs: Short context, every model copes. Up to 8,000 tokens.
Source: NoLiMa (arXiv 2502.05167). As of 19 August 2026.
A wrong answer first costs an escalation and a correction, so human time; that is zone 3. As soon as the answer becomes legally binding, the case belongs in zone 4: Moffatt v. Air Canada, 2024 BCCRT 149, ended at 812 CAD for one wrong sentence from the chatbot. Checking is by sample, thinking effort stays low, and the answer has to hold across a few paragraphs.
Result
Answering a support ticket, Typical conversation, 10,000 per month.
Requirements: zone 3 rework, spot-checked, brief, paragraphs, English.
The measurement covers only 2 current models; the rest count as not measured from zone 3.
Show basis
One task: 3,700 input and 400 output tokens, 1 call, attempts per model from the solve rate on the anchor τ²-bench, customer service with policy and tools (tau2-bench (sierra-research/tau2-bench), pass rate. The README of the benchmark repo carries no table and points to taubench.com instead; the figures come from the CodeSOTA aggregator table, retrieved 19 August 2026., as of 19 August 2026) with a retry correction of 1.2. Monthly values assume 10,000 such tasks. 3,700 tokens per conversation from the Anthropic documentation; 88 percent of AI-using companies deploy AI in customer contact (Bitkom, 02/2026).
gpt-5.2 · Workhorse class
Attempts matter here: the ranking uses 1 divided by the solve rate on the anchor τ²-bench, customer service with policy and tools.
For this case the attempts come from the solve rate; the break-even over attempts does not apply.
Against zone 1 with machine-checkable review, your profile costs €10,400.00 per month.
Average model costs across the 38 models in the field sit at 1.4 times the cheapest eligible model. Review and error costs stay out of this comparison, they hit every model alike.
Cost per completed task: €1.06. Monthly total: €10,571.47.
Solve rate 73%, meaning 1.4 attempts on average.
Suitable models for this task
Class floor and solve rate decide, then price. Solve rates of the 2 rated models range from 59 to 73 percent, and cost per solved task varies by a factor of 1.
Read the assessment
A standard ticket runs about 3,700 input tokens and a few hundred output tokens; at this size the price gap between the compact class and the top class is at its widest. Whether the compact class holds shows in your attempts: at one attempt the price gap wins, and as attempts rise the arithmetic tips at the break-even point.
Own assessment, not a benchmark. The numbers above are price arithmetic, not a quality measurement.
The five cheapest in this scenario
Calculated on the basis shown in the report.
- 1Ministral 3 3BEdgebelow class floor
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,003.54
- Solve rate
- —
- Gap to the recommendation
- -52.7%
- 2Ministral 3 8BEdgebelow class floor
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,005.31
- Solve rate
- —
- Gap to the recommendation
- -52.7%
- 3Qwen3.5 9B (Together)hosted at Together AICompactnot measured
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,006.30
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 4Mistral Small 4Compactnot measured
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,006.87
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 5Ministral 3 14BEdgebelow class floor
- Per solved task
- €0.50per task
- Per month, 10,000 tasks
- €5,007.08
- Solve rate
- —
- Gap to the recommendation
- -52.6%
- 37●gpt-5.2Workhorse
- Per solved task
- €1.06per task
- Per month, 10,000 tasks
- €10,571.47
- Solve rate
- 73 %
- Gap to the recommendation
- —
| Rank | Model | Per month, 10,000 tasks (€) | Gapto the recommendation |
|---|---|---|---|
| 1 | Ministral 3 3BEdgebelow class floor | €5,003.54 | -52.7% |
| 2 | Ministral 3 8BEdgebelow class floor | €5,005.31 | -52.7% |
| 3 | Qwen3.5 9B (Together)hosted at Together AICompactnot measured | €5,006.30 | -52.6% |
| 4 | Mistral Small 4Compactnot measured | €5,006.87 | -52.6% |
| 5 | Ministral 3 14BEdgebelow class floor | €5,007.08 | -52.6% |
| Recommendation | |||
| 37 | ●gpt-5.2Workhorse | €10,571.47 | — |
Calculated for English text.
Rows without a solve rate on the anchor use the attempts you set.
What the literature recommends
Workhorse class as the default, compact class as the floor
The floor is the fast class of the strong vendors. Once the answer reaches the customer, the workhorse class is the default. τ²-bench spreads the models from 36 to 79 percent solve rate, and a wrong answer binds the company: Moffatt v. Air Canada ended at 812 CAD for one sentence. For the classes below Haiku and Flash no customer service measurement exists, in either direction. Caching the policies cuts the bill further than a step down in class: a cache read at Anthropic costs one tenth of the input price (pricing documentation, as of 18 August 2026).
- Anthropic, Choosing the right model 19 August 2026
- OpenAI, Model selection guide 19 August 2026
- tau2-bench leaderboard, pass rate 19 August 2026
- Moffatt v. Air Canada, 2024 BCCRT 149 14 February 2024
As of 19 August 2026
The calculator and the literature arrive at the same class here.
- Anthropic, pricing documentation (3,700 tokens per support conversation) 18 August 2026
- Bitkom, study report on artificial intelligence, 02/2026 19 August 2026
- Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 14 February 2024
- ABC News Australia on Commonwealth Bank, 21 August 2025 (no address) 21 August 2025
- Customer Experience Dive on Klarna, 20 November 2025 (no address) 20 November 2025
Fine-tune
Fine-tune: model, attempts, cache, calls
For this case the attempts per model come from the solve rate on the anchor; the slider has no effect.
The fee falls due when credit is topped up and does not sit on the individual call. Other routes, the Vercel Gateway among them, charge nothing for it.
All models
All values in EUR per 1M tokens · list prices, as of 18 August 2026
Cheapest, Frontier class
Mistral Large 3
€0.4319 / €1.30 per 1M tokens
Cheapest, Workhorse class
Qwen3.7-Plus (Together)
€0.2764 / €1.11 per 1M tokens
Cheapest, Compact class
Mistral Small 4
€0.1296 / €0.5183 per 1M tokens
As of
Prices 18 August 2026
Rate 1.1576 USD/EUR, ECB 18 August 2026
Capability (ECI) 18 August 2026
Knowledge
How the costs arise. Every card carries its sources.
What is a token?
The rule of thumb is one token per four characters of English, roughly 750 words per 1,000 tokens. German packs more densely: compounds and umlauts split into more pieces, so the same content costs a little more.
Every price on this page refers to 1M tokens. That sounds like a lot and fills up fast: a single 500 kB PDF already runs about 125,000 tokens, an eighth of it.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 18 August 2026
Input and output
Input covers system instructions, prior conversation and attached documents, and it is billed again on every single call. A typical 10 kB web page runs about 2,500 tokens.
Output includes reasoning steps, even when you never see them. Tasks with long output, drafts for example, shift the largest cost item from input to output.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 18 August 2026
The cache
The per-model prices sit in the model table: Anthropic charges 1.25 times the input price to write and one tenth to read, OpenAI charges the same 1.25 times to write since the gpt-5.6 series, Mistral has read for one tenth since August 2026, DeepSeek for about three percent.
This has a consequence that gets overlooked: on a single call, caching is MORE expensive than none, because it is written and never read. The benefit starts with the second call of the same task.
Every model and every effort level keeps its own cache. Switching mid-task forces a fresh write.
Sources: Anthropic, pricing documentation · Mistral, prompt caching documentation · DeepSeek, pricing page
As of 18 August 2026
Calls per task
A simple chat is one call. Once tools, search steps or intermediate checks join in, the number multiplies, and with it the input that is billed anew on every call.
This value stays easy to overlook: token prices sit on the pricing page, the number of calls sits nowhere and has to be measured.
Own assessment, no external source.
As of 19 August 2026
Attempts until it holds
This is where a cheap model proves whether it is actually cheaper. In research from March 2026, the cheaper-listed model produced the higher total cost in 32 percent of model pairs, mostly through highly variable thinking tokens and more working steps per task.
The break-even in the result works this out for your scenario: the number of attempts at which the cheapest model loses its lead to the priciest.
Where an anchor supplies a solve rate, the calculator derives the attempts from it, with a correction. Yang (arXiv 2605.08563, 8 May 2026) shows that assuming independent tries puts pass@3 17.4 points too high, 98.6 percent against 81.2 percent: a second attempt inherits the failure cause of the first. The calculator therefore applies a factor of 1.2 to the attempts. That factor is a stated assumption drawn from this measurement, not a measurement of our own.
Sources: Chen et al. 2026, arXiv 2603.23971 · Yang, arXiv 2605.08563, 8 May 2026
As of 19 August 2026
The second price tier
The higher rate then applies to ALL tokens of the call, including those below the threshold. That is how the vendors bill, and the thresholds often sit in a footnote below the pricing table.
A calculator that flatly uses the page-one price therefore underestimates long documents by up to half. Anthropic does not tier and offers the full window at the base price.
The "Trait" filter in the model table shows which nine models are affected; in the result, affected rows carry the marker "long-context rate".
Sources: OpenAI, pricing page · Google, Gemini pricing page
As of 18 August 2026
The context window
The range is wide: Claude Haiku 4.5 takes up to 200,000 tokens, the current Opus and gpt-5.6 models around 1M. In the result, models whose window is too small for your input carry the marker "does not fit".
A dash in the model table means the vendor publishes no size for this model on its pricing page. No estimated number stands in for it.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 18 August 2026
Batch processing
Anthropic, OpenAI, Google and Mistral publish the batch discount as its own price list; the open-weights providers in our set list none. Where it is missing, this calculator drops the discount and says so at the number.
For recurring volume work, overnight document processing for example, that is 50 percent off; quality stays the same and only the wait is added.
Sources: Anthropic, pricing documentation · OpenAI, pricing page · Google, Gemini pricing page
As of 18 August 2026
Routers and data residency
Routers apply no markup to the tokens themselves; top-ups carry 5.5 percent by card. Requests are forwarded to whichever provider is selected, and that provider may use them for training or improvement. A fixed subprocessor list requires pinning the routing down.
Forced US processing costs 1.1 times the standard rate at Anthropic. The residency parameter currently knows no EU value; EU processing runs through the cloud platforms offering European regions.
Sources: OpenRouter, pricing page · Anthropic, pricing documentation
As of 18 August 2026
The five model classes
The mapping follows vendor naming: Anthropic tiers Opus, Sonnet, Haiku; OpenAI sol, terra, luna plus the pro tiers; Google Pro, Flash, Flash-Lite; Mistral Large, Medium, Small, plus the Ministral line explicitly for on-device use and Codestral and Devstral for code.
The class sorts, the capability column measures. The two can diverge: DeepSeek V4 Pro is its vendor’s frontier class at a compact-class price, and an edge-class model can top list-price rankings while its small context window and missing capability measurement rule it out for many business cases.
Only models bookable at an active endpoint with a published list price enter the set. Circulating headline prices without a bookable endpoint stay out, however tempting the number looks.
Sources: Anthropic, pricing documentation · OpenAI, pricing page · Google, Gemini pricing page · Mistral, API pricing page
As of 18 August 2026
Capability, difficulty, solve rate
The Epoch Capabilities Index (ECI) aggregates over 50 benchmarks into one capability value per model via item response theory, the same tooling that calibrates exams such as the GMAT. The scale is relative: GPT-5 sits at 150 by definition, Claude 3.5 Sonnet at 130.
The economic consequence is cost-of-pass: cost per solved task equals cost per attempt divided by the solve rate. A model just below the task difficulty gets expensive through attempts; far above it you pay for capability the task does not need.
The calculator uses this for business cases with an anchor benchmark, and only with measured solve rates: where no measurement exists, a dash stands instead of an estimated number, and from zone 3 upwards the row cannot win. Where no anchor holds, the report says so openly.
Sources: Epoch AI, Epoch Capabilities Index (CC-BY) · Ho et al. 2025, A Rosetta Stone for AI Benchmarks, arXiv 2512.00193 · Erol et al. 2025, Cost-of-Pass, arXiv 2504.13359
As of 18 August 2026
Prices and comparability
Verified 18 August 2026 directly against vendor pricing pages, without discounts. Prices change silently: between two checks of this calculator, one model temporarily sat five times too high in the list, with no notice anywhere.
Anthropic documents a tokeniser for Claude 4.7 and later that produces roughly 30 percent more tokens for the same text. A price comparison across vendors therefore only holds as an order of magnitude.
Sources: Anthropic, pricing documentation · OpenAI, pricing page
As of 18 August 2026
What an error costs
Zellinger and Thomson put a number on the threshold: “reasoning models offer better accuracy-cost tradeoffs as soon as the economic cost of a mistake exceeds $0.01”. Model cascades lose their advantage from about $0.1. One precondition belongs with it: the threshold holds as long as the stronger model is also the more accurate one. For code review and research the measurements often run the other way; the threshold logic still holds there, the ranking behind it has to be measured.
In customer service the two figures sit two orders of magnitude apart. A ticket of 3,700 tokens costs fractions of a cent in every model class, while the market price of a resolved ticket runs 0.49 to 2.00 USD (Drag and MavenAGI 2026, secondary sources). Saving on the token price while leaving out escalation optimises the smaller of the two numbers.
Two documented cases show the upper end. In Moffatt v. Air Canada (2024 BCCRT 149, 14 February 2024) the tribunal awarded 812.02 CAD and stated: “It should be obvious to Air Canada that it is responsible for all the information on its website.” Klarna announced in February 2024 the work of a calculated 700 full-time agents and $40 million in profit improvement, and corrected course in May 2025: “As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”
The requirements section is where you set this figure as a zone. The presets run from 0 euro through the cost of one more call up to 100 euro per task. The top zone carries no figure, because the error cannot be priced there, and it switches the result to model group and process.
Sources: Zellinger and Thomson, Economic Evaluation of LLMs, arXiv 2507.03834 · Drag and MavenAGI 2026, market price per resolved ticket (secondary source) · Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 · Klarna, announcement 27 February 2024 and correction 05/2025 (secondary source)
As of 19 August 2026
Checking costs time and catches part of it
A simulation study by MedStar and Georgetown (npj Digital Medicine, 24 April 2025) put 18 patient messages with AI drafts in front of 20 primary care physicians, four of them carrying planted errors. Between 35 and 45 percent of the flawed drafts went out unchanged, every flawed draft was missed by at least 13 of the 20 participants, on average 2.67 of 4 errors went unnoticed, and exactly one participant addressed all four. The authors name the cause as automation bias, “the tendency to over-rely on automation”.
Checking costs measurable time. At UC San Diego reading time per message rose by 21.8 percent (JAMA Network Open, 15 April 2024, p equals 0.008), while reply time fell by 5.9 percent without statistical significance. The calculator therefore applies review minutes times an hourly rate, with 60 euro per hour as a stated assumption.
Attention is a finite resource. Turan (arXiv 2606.08919, 8 June 2026) models reviewers who tire as the escalation load grows: “when the reviewer is modeled as endogenous (fatiguing as escalation load grows), realized safety becomes an inverted-U in the escalation rate: more human oversight can make a system less safe”. Reviewers also agree only moderately on what counts as risky, Fleiss kappa 0.52. The work is a single-author preprint without a human study; it stands here as a caution and does not enter the calculation.
OpenAI writes about GDPval (25 September 2025) that frontier models handle those tasks around 100 times faster and around 100 times cheaper than industry experts, and clarifies in the same passage that these figures cover inference time and API cost only and leave out human oversight, revision and integration. Those are the items the reviewability control brings into the calculation.
Sources: MedStar and Georgetown, npj Digital Medicine, 24 April 2025 · UC San Diego (Tai-Seale et al.), JAMA Network Open, 15 April 2024 · Turan 2026, Oversight Has a Capacity, arXiv 2606.08919 · OpenAI, GDPval, 25 September 2025
As of 19 August 2026
The price reversal
Chen et al. (arXiv 2603.23971, 25 March 2026, 8 models, 12 tasks) measure: “in 32% of model-pair comparisons, the model with a lower listed price actually incurs a higher total cost, with reversal magnitude reaching up to 28x”. On the same query one model burns up to 900 percent more thinking tokens than another, or ten times as many turns. Repeating the same query on the same model, thinking tokens vary by up to 9.7 times; the paper calls that an irreducible noise floor for any predictor.
OckBench (arXiv 2511.05722) measures the same effect at equal accuracy: more than a 25-fold token difference, about 1,600 against about 42,000 tokens. The paper puts it this way: “cheaper token cost does not always imply cheaper task cost; verbose smaller models can pay an Overthinking Tax”.
The effort level works at the same magnitude. Artificial Analysis measured GPT-5 at high effort using 23 times the tokens of minimal, 82M against 3.5M for the whole index, at 68 against 44 index points; the step from medium to high added one point. The generation counts too: Claude Sonnet 5 produces around 40 percent more output tokens per index task than Sonnet 4.6, needs around three times as many agent turns and costs around twice as much per task, at the same list price.
This version therefore calculates with a volume factor per model and effort level and shows the result as a range. Where no measured spread exists for the upper band, that stands at the figure.
Sources: Chen et al. 2026, arXiv 2603.23971 · OckBench, arXiv 2511.05722 · Artificial Analysis, GPT-5 benchmarks and analysis, 7 August 2025 · Artificial Analysis, Claude Sonnet 5 agentic cost, 30 June 2026
As of 19 August 2026
The tokeniser as a volume factor
Anthropic documents a new tokeniser for models from Claude 4.7 on: “This tokenizer produces approximately 30% more tokens for the same text.” At the same price per token that is a surcharge of roughly 30 percent on the volume side, without any price figure changing.
Across vendors the range is wider. On 10 June 2026 TextKit measured German text at 1.71 tokens per word for GPT-5, 2.18 for GPT-4, 2.64 for Claude Sonnet 4.6 and 3.48 for Claude Opus 4.8; English text sits at 1.17 to 1.88. Between the ends lies a factor of 2.0, and it hits input and output alike.
Measured across 24 EU languages (arXiv 2605.24718) the token count per word spreads by a factor of 2.5, from 1.2 in English to 3.1 in Greek. German averages 1.76 and ranges from 1.55 to 1.98 depending on the vendor, so 1.28 times through the choice of vendor alone.
The calculator applies the factor per model family to input and output and shows it as its own column in the model table. A comparison that treats tokens as a vendor-neutral unit favours the vendors with the coarser tokeniser.
Sources: Anthropic, pricing documentation · TextKit, tokens per word across vendors, 10 June 2026 · Tokeniser surcharge across 24 EU languages, arXiv 2605.24718
As of 19 August 2026
Window size and context fidelity
Two vendors report the same long-context value in their model cards, MRCR v2 with eight needles, and that is where the label falls apart. Gemini 3.1 Pro drops from 84.9 percent at 128,000 tokens to 26.3 percent at 1M (model card, 19 February 2026). Anthropic writes about the 1M variant: “on the 8-needle 1M variant of MRCR v2 … Opus 4.6 scores 76%, whereas Sonnet 4.5 scores just 18.5%”. Same label, fourfold difference.
NoLiMa (ICML 2025) measures the drop where the question does not literally overlap with the passage: at 32,000 tokens eleven of thirteen models fall below half their short-context performance. RULER confirms this across 17 models and reports an effective length below the advertised one: Llama-3.1-70B is listed at 128k and carries 64k.
Chroma shows that the drop hits simple tasks as well: “model performance varies significantly as input length changes, even on simple tasks”. Needle search tests less than real analysis work, which sits above it.
In the model table fidelity has its own column for 128k and 1M. Four models in the set carry a value: three of them are reported by the vendors in their model cards, the fourth, Gemini 3.7 Flash, comes from a secondary source and carries a different metric, GDM-MRCR long context. All remaining models carry a dash, because no figure exists. The context length control in the requirements section works with these values.
Sources: Google DeepMind, Gemini 3.1 Pro model card, 19 February 2026 · Anthropic, Claude Opus 4.6, 5 February 2026 · NoLiMa, ICML 2025, arXiv 2502.05167 · NVIDIA RULER, effective context length · Chroma, Context Rot, 14 July 2025
As of 19 August 2026
The same question twice
On 10 September 2025 Thinking Machines Lab traced the cause to the varying batch size on the server: with the load, the order in which the kernel reduces changes as well. Measured, 1,000 identical requests at temperature 0 produced eighty different answers. With batch-invariant kernels all 1,000 answers were identical, and runtime rose from 26 to 55 seconds, and to 42 with an improved attention kernel.
The variation does not stay in the wording. Atil et al. (arXiv 2408.04667, five models, eight tasks, ten runs) report: “We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%.” Yuan et al. (arXiv 2506.09501) measure “up to 9% variation in accuracy and 9,000 tokens difference in response length” on a reasoning model and trace it to floating-point arithmetic at limited precision.
Reliability therefore comes from the architecture around the model. The documented base pattern: the model proposes, a fixed rule or a checking program decides, and only the decision counts. Structured outputs secure the form of the answer; about the content they say nothing, and the vendor notes that such outputs can still contain mistakes. Repeated sampling and voting schemes lower the spread and do not create determinism, because they draw several samples on purpose.
The BSI lists non-reproducibility as a risk category of its own, R7 (Generative AI Models, version 2.0, 17 January 2025): “The outputs of many generative AI models are not necessarily reproducible due to the use of random components.” Measure M20 requires, where the impact may be critical, a review of the output, cross-referencing with further sources and manual post-processing where needed. In the result, the top error-cost zone shows the documented patterns with their limits.
Sources: Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, 10 September 2025 · Atil et al., Non-Determinism of Deterministic LLM Settings, arXiv 2408.04667 · Yuan et al., Numerical Sources of Nondeterminism in LLM Inference, arXiv 2506.09501 · BSI, Generative AI Models, version 2.0, 17 January 2025, risk R7 and measure M20
As of 19 August 2026
Felt and measured effect
METR (Becker, Rush, Barnes, Rein, arXiv 2507.09089, 12 July 2025) had experienced developers work on mature code, randomised with and without AI. Measured, they were 19 percent slower with AI, while they had estimated themselves faster. Between self-assessment and measurement lie around 39 percentage points.
The counter-figures belong with it. A rollout across tens of thousands of developers at Microsoft shows around 24 percent more merged pull requests (arXiv 2607.01418, 1 July 2026). And the Stanford AI Index 2026 records that METR later could not replicate its own finding, “primarily due to a growing reluctance among developers to work without AI, and that developers in late 2025 were likely sped up by AI relative to the original study period” (page 219).
One rule for this calculator follows: it is never calibrated with self-assessments. Every case study resting on a satisfaction survey can be optimistic by that order of magnitude, and vendor figures on time saved carry the same risk. The numbers here come from price lists, model cards and measurement series with a published setup.
Sources: Becker, Rush, Barnes, Rein (METR), arXiv 2507.09089, 12 July 2025 · Microsoft rollout across tens of thousands of developers, arXiv 2607.01418, 1 July 2026 · Stanford AI Index 2026, page 219 (note on the METR replication)
As of 19 August 2026
Which benchmark per business case
GPQA Diamond is built to be “Google-proof”, that is, built so searching does not help. A measure that neutralises search ability cannot assess it. On top of that it is saturated: on 15 August 2026 vals.ai counts “24 of the 133 models tested now score 90% or higher”, and the benchmark author, David Rein, says in an interview that it “stopped discriminating between good and great”. For research, BrowseComp and Deep Research Bench hold, the latter being the only one that reports cost per task.
SWE-bench Verified measures producing a patch when the fault location is known. A review has to find the fault first. CR-Bench (arXiv 2603.11078, 10 March 2026) sets out the difference and introduces its own measures, usefulness rate and signal-to-noise: “these agents often face a fundamental trade-off, they either prioritize precision and risk missing critical vulnerabilities, or prioritize recall at the cost of producing noisy and low-actionable feedback.”
For support tickets τ²-bench holds, because it tests the three properties of the case: a policy the system has to follow, tool use, and a simulated customer who follows up on incomplete answers. The spread on the public leaderboard is wide: Claude Opus 4.5 79 percent, GPT-5.2 73, Sonnet 4.5 63, GPT-4o 36. As a blocking criterion for source-bound output the HHEM hallucination leaderboard serves, read as a minimum threshold, because a low rate can also reflect low capability.
For text work there is a retrievable anchor with a price column, arena.ai creative writing (as of 12 August 2026, 1.18M votes). Confidence intervals of ±7 to ±18 points in the top group (±5 to ±28 across the whole field) make the leading ranks statistically indistinguishable. The anchor therefore works as a threshold: it answers whether a model belongs to the top group, and inside that group the price decides.
Sources: MindStudio, interview with GPQA author David Rein, 5 May 2026 · vals.ai, GPQA evaluation, 15 August 2026 · Deep Research Bench, as of 10 June 2026 · CR-Bench, arXiv 2603.11078, 10 March 2026 · τ²-bench, sierra-research/tau2-bench · CodeSOTA, τ²-bench leaderboard 2026 (secondary source) · Vectara HHEM-2.3, values reported via CodingFleet 2026 (secondary source) · arena.ai, creative writing, as of 12 August 2026
As of 19 August 2026
The five error zones
Error resilience in the requirement profile has five levels. Each level carries an amount per task and a source; zone 5 carries no figure, there the process decides.
- Zone 1: doesn't matter (€0)An error costs nothing; the result is discarded at once or overwritten anyway.Source: Anthropic, Optimizing for cost and intelligence: "high-volume work with checkable outputs"
- Zone 2: cheap to fix (about one call)The error is noticed and costs one more call, nothing else.Source: Erol et al., Cost-of-Pass (arXiv 2504.13359); Anthropic: "retry the failures"
- Zone 3: rework (€0.50 to €5)The error binds human time: an escalation, a query, a correction.Source: Market price per resolved support ticket 0.49 to 2.00 USD (Drag, MavenAGI 2026); Zellinger and Thomson (arXiv 2507.03834): from about 0.01 USD error cost the stronger model wins
- Zone 4: business impact (€50 to €500)The error costs customers, deadlines or legal positions.Source: Cost-of-Pass, human reference GPQA Diamond 58 USD per task (arXiv 2504.13359); Moffatt v. Air Canada 2024 BCCRT 149 (812 CAD); Klarna rollback 05/2025
- Zone 5: must not happen (no figure, blocked)The error cannot be priced or is existential. A rule set in advance decides.Source: BSI Generative AI Models v2.0 (17 January 2025) M20; AI Act Regulation (EU) 2024/1689 Art. 14, 15; Lemonade: "we never let AI perform deterministic actions"
Formula, assumptions and sources
1 · Formula
Cost per task = machine × modifiers. The machine has four line items: fresh input (× calls × attempts), one cache write, cache reads (× every further call) and output (× calls × attempts). The cache write counts once per task and is not multiplied by attempts, because it is billed only once as well. Human rework is deliberately not part of this calculator. It compares models.
Since version 3 the expected calculation per solved task sits above it: K = cost per attempt × attempts + (1 − solve rate) × error costs + review costs. The attempt is the machine calculation above, extended by the volume and tokeniser factors. Attempts come from the solve rate at the case anchor, times 1.2, because repeated tries inherit the same failure cause (Yang, arXiv 2605.08563); that factor is a stated assumption. Error costs are the amount of the selected zone and fall due at the counter-probability to the solve rate. Review costs are review minutes times an hourly rate, assumed at 60 euro per hour. The machine calculation from v2 stays untouched: with error costs 0, review costs 0, normal effort, a neutral tokenizer factor and no anchor, the formula returns exactly the value of the earlier version.
2 · Cache
The cache is calculated the way it is billed: written once per task, then read. That is why it costs more than it saves on a single call. The calculator uses each model’s published cache prices. Anthropic charges 1.25 times to write and one tenth to read, OpenAI charges the same 1.25 times to write since the gpt-5.6 series, DeepSeek reads for about three percent, Mistral has read for one tenth of the input price since August 2026. Google’s hourly storage fee depends on runtime and is NOT included.
3 · Surcharges and discounts
Modifiers: batch × 0.5 (only at vendors that offer it), US residency × 1.1, router × 1.055. Attempts multiply calls. Running a task twice still writes the cache once.
4 · Second price tier
Nine models switch to a more expensive tier above a certain context size: Gemini 2.5 Pro and Gemini 3.1 Pro above 200,000 tokens, the gpt-5.6 series, gpt-5.5, gpt-5.5-pro, gpt-5.4 and gpt-5.4-pro above 272,000. The higher rate then applies to every token of the call, including those below the threshold, because that is how the vendors bill it. Affected rows carry the note "long-context rate" in the ranking. Anthropic does not tier and offers the full window at the base price.
5 · Capability and solve rate
The section "All models" shows the Epoch Capabilities Index per model (Epoch AI, data CC-BY, as of 18 August 2026), a capability value merged across 50-plus benchmarks via item response theory. The scale is relative: GPT-5 sits at 150 by definition, Claude 3.5 Sonnet at 130.
9 of the 12 business cases carry an anchor benchmark from anker.json, each with source, licence and date. Only anchors of type "rate" yield a solve rate per model; attempts follow from it under the cost-of-pass principle: cost per solved task equals cost per attempt divided by solve rate. Anchors of type "threshold" (Elo), "gate" (hallucination rate) and "fidelity" (MRCR, context fidelity across the full window) only filter and are never read as an attempt count; for "Working through a long document" the attempts therefore come from the case's manual value.
Nothing is estimated: what sits in anker.json is measured, and models without a value show a dash. From zone 3 upwards a row without a measurement cannot win. Rates are bounded to 5 through 99 percent. The benchmark approximates the business case; it does not measure the concrete task, and expert mode keeps calculating with manual values. Where a case carries no anchor, the assessment says what the ranking follows instead.
6 · Model set
The set covers 75 models, including deprecated and retired ones. The ranking in the result hides them by default, because nobody should adopt them fresh; anyone running an existing estate brings them back through the checkbox. The section "All models" does the same: the table shows the current ones by default, and the switch "show superseded and retired" above it brings in the rest. Where the context window is too small for the configured input, the row reads "does not fit". A cheap model is no use if the task will not fit. Every model carries one of five classes (frontier, workhorse, compact, edge, code), derived from the vendor’s own positioning; the class sorts, measurement sits in the capability column. Only models bookable at an active endpoint with a published list price enter the set; circulating headline prices without a bookable endpoint stay out.
7 · Evidence depth
Every price, cache rate, batch discount and context tier comes straight from the vendor pricing pages linked below. A context window is published for 72 of the 75 models; only Claude Opus 4, Claude Sonnet 4 and Claude Haiku 3.5 carry a dash, because the vendor publishes no size for them. The deprecated classification follows the vendor’s own model overview at Anthropic; at the other vendors, which publish no such marking, it is my own reading by model generation.
8 · Price status and conversion
Prices are vendor list prices, verified 18 August 2026 against the pricing pages linked below, without discounts. Conversion uses 1.1576 USD/EUR (ECB reference rate of 18 August 2026); that rate is an assumption, not a daily quote. Among the business cases, only the 3,700 tokens for a support conversation is empirically documented (source: Anthropic documentation). All other sizes are plausible example values, not measurements. The dates are deliberately separate: prices as of 18 August 2026, anchors and zone evidence as of 19 August 2026.
9 · Cross-check
Anthropic’s documentation puts 10,000 support conversations of 3,700 tokens each on Claude Haiku 4.5 at about 37 US dollars. That is exactly what this calculator returns in expert mode with output and cache set to zero.
Sources
- Anthropic: pricing documentation (prices, cache rates, context windows, cross-check)
- OpenAI: pricing page (prices, tiers from 272,000 tokens)
- Google: Gemini pricing page (prices, tiers from 200,000 tokens, storage fee)
- Mistral: API pricing (prices, cache discount)
- DeepSeek: pricing page
- Together AI: pricing page (open weights at a hoster)
- OpenRouter: pricing page (router fee)
- ECB: euro reference rate USD
- Epoch AI: Epoch Capabilities Index (capability values and solve rates, data CC-BY)
- Ho et al. 2025: A Rosetta Stone for AI Benchmarks (method behind the capability index, arXiv 2512.00193)
- Erol et al. 2025: Cost-of-Pass (cost per solved task, arXiv 2504.13359)
- Chen et al. 2026: The Price Reversal Phenomenon (price reversal between list price and total cost, arXiv 2603.23971)