# AI Cost Calculator

> Which AI model carries your requirement, what does safety cost? Cost per solved task, twelve scenarios, 75 models. As of 18 August 2026.

- Canonical URL: https://olivergausmann.com/en/insights/ai-cost-calculator
- Prices as of 18 August 2026
- Anchors and zones as of 19 August 2026
- Exchange rate: 1.1576 USD/EUR (ECB reference rate of 18 August 2026); that rate is an assumption, not a daily quote.

## What the calculator does

Requirement first, then the number: pick the case, set in the requirement profile what an error costs and who reviews it, and see the cost per solved task for 75 models, with a range and the comparison to the literature.

Cost per task = machine × modifiers. The machine has four line items: fresh input (× calls × attempts), one cache write, cache reads (× every further call) and output (× calls × attempts). The cache write counts once per task and is not multiplied by attempts, because it is billed only once as well. Human rework is deliberately not part of this calculator. It compares models.

Since version 3 the expected calculation per solved task sits above it: `K = cost per attempt × attempts + (1 − solve rate) × error costs + review costs`. The attempt is the machine calculation above, extended by the volume and tokeniser factors. Attempts come from the solve rate at the case anchor, times 1.2, because repeated tries inherit the same failure cause (Yang, arXiv 2605.08563); that factor is a stated assumption. Error costs are the amount of the selected zone and fall due at the counter-probability to the solve rate. Review costs are review minutes times an hourly rate, assumed at 60 euro per hour. The machine calculation from v2 stays untouched: with error costs 0, review costs 0, normal effort, a neutral tokenizer factor and no anchor, the formula returns exactly the value of the earlier version.

## The twelve cases

### Answering a support ticket

Understand a customer request and propose an answer that an agent approves or that goes out directly.

What matters: What a wrong answer costs with the customer, how many tickets arrive per month and whether someone reads along.

Default sizes: Short ticket, one follow-up (input 1,500 tokens, output 200 tokens); Typical conversation (input 3,700 tokens, output 400 tokens); Long thread with attachments (input 9,000 tokens, output 800 tokens).

Profile preset: F3 (rework), P b (spot-checked), D low (brief), K absaetze (paragraphs).

Class floor: Compact class.

Anchor: τ²-bench, customer service with policy and tools.

#### Background and evidence

3,700 tokens per conversation from the Anthropic documentation; 88 percent of AI-using companies deploy AI in customer contact (Bitkom, 02/2026).

Typical failure modes: A wrong answer that binds the company. Then the volume boomerang: Commonwealth Bank reversed 45 redundancies because the voice bot increased call volumes. Klarna rolled back in 2025 citing quality, and was back at 853 full-time equivalents as of Q3 2025.

Sources for this case:

- Anthropic, pricing documentation (3,700 tokens per support conversation) — https://platform.claude.com/docs/en/about-claude/pricing (18 August 2026)
- Bitkom, study report on artificial intelligence, 02/2026 — https://www.bitkom.org/sites/main/files/2026-02/bitkom-studienbericht-ki.pdf (19 August 2026)
- Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 — https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html (14 February 2024)
- ABC News Australia on Commonwealth Bank, 21 August 2025 (no address) (21 August 2025)
- Customer Experience Dive on Klarna, 20 November 2025 (no address) (20 November 2025)

### Reading an email and drafting the reply

Read incoming emails, pull out the essentials and draft a reply.

What matters: The volume per day, the share someone checks before sending and whether an invented detail gets noticed.

Default sizes: Short email, a few paragraphs (input 800 tokens, output 150 tokens); Typical email (input 2,000 tokens, output 300 tokens); Long thread with many participants (input 6,000 tokens, output 600 tokens).

Profile preset: F2 (cheap to fix), P b (spot-checked), D low (brief), K absaetze (paragraphs).

Class floor: Compact class.

Anchor: HHEM, hallucination when summarising supplied text.

#### Background and evidence

Email drafts are 7 percent of work conversations (Anthropic Economic Index, 06/2026); the matching API task is "Analyze emails and draft replies" at 0.28 percent of all calls.

Typical failure modes: A silently dropped question and a break in register. For summarisation the top 15 models on the HHEM leaderboard show hallucination rates of 1.8 to 5.4 percent, most clustering at 4 to 5 percent (Stanford AI Index 2026).

Sources for this case:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- Anthropic Economic Index, January 2026 (top API tasks) — https://www.anthropic.com/research/anthropic-economic-index-january-2026-report (15 January 2026)
- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)

### Writing or revising a text

Draft and revise reports, documentation or marketing copy.

What matters: How long the source material is, how often you ask for another pass and whether the text goes out unchecked.

Default sizes: Short text, one page (input 1,500 tokens, output 500 tokens); Typical report or article (input 4,000 tokens, output 1,200 tokens); Long text with templates and attachments (input 12,000 tokens, output 2,500 tokens).

Profile preset: F2 (cheap to fix), P b (spot-checked), D medium (normal), K absaetze (paragraphs).

Class floor: Workhorse class.

Anchor: arena.ai creative writing, writing quality by audience vote.

#### Background and evidence

Documents and reports are 20 percent of work conversations (Anthropic Economic Index, 06/2026); 78 percent of AI-using companies use it to produce text, images or code (DIHK, 01/2026).

Typical failure modes: Silent omission and a break in register weigh heavier here than invention. The volume gain is measured, the quality check stays with the person: marketing teams produced 50 percent more ads per head with multimodal AI (Ju and Aral 2025, Stanford AI Index 2026).

Sources for this case:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- DIHK, Digitalisation 2026: companies stay the course, 28 January 2026 — https://www.dihk.de/de/newsroom/digitalisierung-2026-unternehmen-halten-kurs-163290 (28 January 2026)
- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)

### Sales email and proposal draft

Draft a personal approach or a proposal from CRM data and a template.

What matters: How much context goes in per recipient, how many drafts you need per day and what an embarrassing line costs with the customer.

Default sizes: Short proposal (input 3,000 tokens, output 600 tokens); Typical proposal (input 6,000 tokens, output 1,200 tokens); Detailed proposal (input 15,000 tokens, output 2,500 tokens).

Profile preset: F2 (cheap to fix), P c (every task reviewed), D medium (normal), K absaetze (paragraphs).

Class floor: Workhorse class.

Anchor: arena.ai creative writing, writing quality by audience vote.

#### Background and evidence

"Generate personalized B2B cold sales emails" is the most frequent non-technical task in the Anthropic API telemetry at 0.47 percent of all records (01/2026).

Typical failure modes: Wrong terms, invented references and errors of tone. There is no reliable error-rate study for this case; the research package lists it as the weakest evidenced of the six old cases.

Sources for this case:

- Anthropic Economic Index, January 2026 (top API tasks) — https://www.anthropic.com/research/anthropic-economic-index-january-2026-report (15 January 2026)
- DIHK, Digitalisation 2026: companies stay the course, 28 January 2026 — https://www.dihk.de/de/newsroom/digitalisierung-2026-unternehmen-halten-kurs-163290 (28 January 2026)
- Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 — https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html (14 February 2024)

### Classification and routing

Sort requests, tickets or documents into categories and route them.

What matters: The volume, the hit rate and what a wrong assignment triggers in detours.

Default sizes: Short message, one label (input 500 tokens, output 20 tokens); Typical inbound item (input 1,500 tokens, output 30 tokens); Long item with attachment (input 4,000 tokens, output 60 tokens).

Profile preset: F2 (cheap to fix), P a (machine-checkable), D low (brief), K absaetze (paragraphs).

Class floor: Compact class.

Anchor: HHEM, hallucination when summarising supplied text.

#### Background and evidence

"Classify and categorize emails into predefined labels" is a top-10 task in the Anthropic API telemetry at 0.23 percent (01/2026); LinkedIn measures 30.6 percentage points more routing accuracy in a production A/B test.

Typical failure modes: The expensive error is the silently wrong assignment that nobody notices; it is more dangerous than a visibly missing one (arXiv 2606.24420). The way out runs through per-field confidence and hand-back to a person.

Sources for this case:

- Anthropic Economic Index, January 2026 (top API tasks) — https://www.anthropic.com/research/anthropic-economic-index-january-2026-report (15 January 2026)
- LinkedIn, production A/B test, arXiv 2608.10224 — https://arxiv.org/abs/2608.10224 (10 August 2026)
- Google Cloud, 101 real-world generative AI use cases (Gelato), 22 April 2026 — https://cloud.google.com/transform/101-real-world-generative-ai-use-cases-from-industry-leaders (22 April 2026)
- Silently wrong extraction, arXiv 2606.24420 — https://arxiv.org/abs/2606.24420 (23 June 2026)

### Extracting data from documents

Transfer fields from contracts, invoices or forms into tables.

What matters: How long the documents are, whether someone checks the values against the original and what a wrong value in the system costs.

Default sizes: Invoice or form (input 4,000 tokens, output 300 tokens); Typical contract or claim (input 15,000 tokens, output 600 tokens); Extensive file (input 60,000 tokens, output 1,500 tokens).

Profile preset: F3 (rework), P c (every task reviewed), D low (brief), K dokument (document).

Class floor: Compact class.

Anchor: Vals CorpFin v2, reading long credit agreements.

#### Background and evidence

Deutsche Bank measures 97 percent accuracy with dbTextract at 40 percent less handling time; on real credit agreements in Vals CorpFin v2: AI Index 2026 (data cut) no model above 70 percent, best value 68.26; leaderboard 12 August 2026: Claude Opus 5 73.19 percent.

Typical failure modes: The silent omission that no sample finds. Then context fidelity: models answer from general knowledge although the answer sits in the supplied document, even where the instruction forbids exactly that (Vals CaseLaw v2, Stanford AI Index 2026). Reasoning mode raises completeness and lowers correctness (ContractEval, arXiv 2508.03080).

Sources for this case:

- Deutsche Bank, dbTextract (technology page, undated) — https://www.db.com/what-we-do/focus-topics/tech (19 August 2026)
- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)
- Vals AI, CorpFin v2 — https://www.vals.ai/benchmarks/corp_fin_v2 (12 August 2026)
- ContractEval, arXiv 2508.03080 — https://arxiv.org/abs/2508.03080 (5 August 2025)

### Knowledge-base question

Answer employee questions from internal policies, manuals and wikis.

What matters: How much knowledge goes in per question, how often people ask and whether anyone notices an invented rule.

Default sizes: Short question, one policy (input 3,000 tokens, output 250 tokens); Typical question with several passages (input 8,000 tokens, output 400 tokens); Broad question across many documents (input 25,000 tokens, output 800 tokens).

Profile preset: F3 (rework), P b (spot-checked), D low (brief), K dokument (document).

Class floor: Workhorse class.

#### Background and evidence

Explanations are 17 percent of all conversations (Anthropic Economic Index, 06/2026); UBS counts more than 25 million queries for its internal assistant since launch.

Typical failure modes: Two errors that belong apart. Open knowledge questions hallucinate heavily, AA-Omniscience measures 22 to 94 percent across 26 models. Grounded answers fail invisibly: "Deceptive Grounding" measures 7.8 percent of cases in a production system where real evidence is attributed to the wrong entity, rising to 13.6 percent for newer drugs (arXiv 2607.09349).

Sources for this case:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- UBS, Innovation and AI (internal assistant Red) — https://www.ubs.com/global/en/our-firm/what-we-do/technology/innovation-and-ai.html (19 August 2026)
- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)
- Deceptive Grounding, arXiv 2607.09349 — https://arxiv.org/abs/2607.09349 (10 July 2026)

### Research agent over documents

Warning case: the calculator lists this case as a warning case, the measurements speak against an unchecked full rollout.

Work up a market, a vendor or a question in several search steps with sources.

What matters: Many calls and long contexts per job, plus the risk of invented references when nobody reads along.

Default sizes: A few documents (input 8,000 tokens, output 600 tokens); Typical dossier (input 25,000 tokens, output 1,200 tokens); Extensive dossier (input 80,000 tokens, output 2,500 tokens).

Profile preset: F4 (business impact), P c (every task reviewed), D high (thorough), K dokument (document).

Class floor: Workhorse class.

Anchor: BrowseComp, research on the live web.

#### Background and evidence

Only 16 percent of enterprise deployments are true agents (Menlo, 12/2025); 3 to 13 percent of citation links are invented and 5 to 18 percent fail to resolve (arXiv 2604.03173).

Typical failure modes: Invented citations, shallow coverage and regressions when reworking. The strongest models satisfy fewer than 50 percent of the rubrics (arXiv 2601.08536), and when incorporating feedback agents destroy previously correct content in 16 to 27 percent of cases (arXiv 2601.13217). A liveness check on dead links cuts invented sources below 1 percent and costs less than any model upgrade.

Sources for this case:

- Menlo Ventures, The State of Generative AI in the Enterprise, 9 December 2025 — https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/ (9 December 2025)
- Fabricated and dead citations in deep research agents, arXiv 2604.03173 — https://arxiv.org/abs/2604.03173 (3 April 2026)
- Depth of research agents, arXiv 2601.08536 — https://arxiv.org/abs/2601.08536 (13 January 2026)
- Regressions when incorporating feedback, arXiv 2601.13217 — https://arxiv.org/abs/2601.13217 (19 January 2026)
- The Guardian on Deloitte Australia, 6 October 2025 (no address) (6 October 2025)

### Writing or changing code

Write, change or review code, with repository context and a test run.

What matters: The context per task, the number of attempts until the tests pass and whether a team review follows.

Default sizes: Small change (input 5,000 tokens, output 600 tokens); Typical change (input 15,000 tokens, output 1,500 tokens); Large change across files (input 40,000 tokens, output 3,000 tokens).

Profile preset: F3 (rework), P a (machine-checkable), D medium (normal), K dokument (document).

Class floor: Workhorse class.

#### Background and evidence

Coding is 55 percent of departmental AI spend at 4.0 billion USD (Menlo, 12/2025); the Copilot field study measures 26 percent more completed pull requests, while an RCT with experienced developers measures a 19 percent slowdown (arXiv 2507.09089).

Typical failure modes: Lasting quality cost: 18 percent more static analysis warnings and 39 percent more cognitive complexity, even where the speed gain evaporates (arXiv 2601.13597). In review operation 56.3 percent of agent comments were rejected (arXiv 2607.03316, 10,191 pull requests).

Sources for this case:

- Menlo Ventures, The State of Generative AI in the Enterprise, 9 December 2025 — https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/ (9 December 2025)
- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)
- METR, RCT with experienced developers, arXiv 2507.09089 — https://arxiv.org/abs/2507.09089 (12 July 2025)
- Quality costs of AI code, arXiv 2601.13597 — https://arxiv.org/abs/2601.13597 (20 January 2026)
- Agentic code reviews in practice, arXiv 2607.03316 — https://arxiv.org/abs/2607.03316 (3 July 2026)

### Working through a long document

Evaluate a contract, report or audit report of hundreds of pages in one pass.

What matters: Whether the document fits the context window, whether the model is still reliable at the end and the long-context surcharges.

Default sizes: Around 120,000 tokens, below the tiers (input 120,000 tokens, output 1,500 tokens); Around 300,000 tokens, above both tiers (input 300,000 tokens, output 3,000 tokens); Around 700,000 tokens, largest windows only (input 700,000 tokens, output 5,000 tokens).

Profile preset: F3 (rework), P c (every task reviewed), D medium (normal), K korpus (corpus).

Class floor: Workhorse class.

Anchor: MRCR v2, context fidelity across the full window.

#### Background and evidence

Typical size deliberately above the tiering thresholds of OpenAI (272,000) and Google (200,000): the more expensive long-context rates apply here.

Typical failure modes: Usable context length falls short of the advertised one: once the question does not overlap literally with the source passage, 11 of 13 models drop below half their short-context performance at 32,000 tokens (NoLiMa, ICML 2025, arXiv 2502.05167). In LongBench v2 the best directly answering model reaches 50.1 percent.

Sources for this case:

- OpenAI, pricing page (tier from 272,000 tokens) — https://developers.openai.com/api/docs/pricing (18 August 2026)
- Google, Gemini pricing page (tier from 200,000 tokens) — https://ai.google.dev/gemini-api/docs/pricing (18 August 2026)
- NoLiMa, ICML 2025, arXiv 2502.05167 — https://arxiv.org/abs/2502.05167 (7 February 2025)
- LongBench v2, arXiv 2412.15204 — https://arxiv.org/abs/2412.15204 (19 December 2024)

### Meeting notes

Turn a transcript or notes into minutes with decisions and tasks.

What matters: The length of the transcripts, the number of meetings and whether a decision is distributed unchecked.

Default sizes: Short conversation, half an hour (input 6,000 tokens, output 400 tokens); Typical meeting (input 15,000 tokens, output 800 tokens); Long workshop with many participants (input 40,000 tokens, output 1,500 tokens).

Profile preset: F1 (doesn't matter), P b (spot-checked), D low (brief), K dokument (document).

Class floor: Compact class.

Anchor: HHEM, hallucination when summarising supplied text.

#### Background and evidence

Sharp HealthCare reports 83 percent less note-writing effort, Northwestern 24 percent less documentation time and 11.3 additional patients per month (Stanford AI Index 2026).

Typical failure modes: Hallucination, omission and irrelevance in the same note (arXiv 2509.15901). The bottleneck sits in transcript quality; a larger model does not raise it.

Sources for this case:

- Stanford AI Index 2026 — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)
- Failure modes of meeting summarisation, arXiv 2509.15901 — https://arxiv.org/abs/2509.15901 (19 September 2025)

### Translation

Transfer texts between languages, usually with post-editing.

What matters: The text volume, the share of post-editing and the language direction, because German text costs more tokens.

Default sizes: Short passage (input 500 tokens, output 500 tokens); Typical text, a few pages (input 2,000 tokens, output 2,000 tokens); Long document for translation (input 8,000 tokens, output 8,000 tokens).

Profile preset: F1 (doesn't matter), P b (spot-checked), D low (brief), K absaetze (paragraphs).

Class floor: none.

#### Background and evidence

Translation carries the lowest autonomy level of all artefacts in the Anthropic telemetry (06/2026); 41 percent of AI users employ it for this (Bitkom, 04/2026).

Typical failure modes: Break in register, technical terms without a fixed equivalent and silent omission of whole sentences. For the quality drop in German only one relevant source exists (arXiv 2607.19243, single-author preprint).

Sources for this case:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- Bitkom, press release: a third uses AI at least once a week, 28 April 2026 — https://www.bitkom.org/Presse/Presseinformation/Ein-Drittel-nutzt-KI-mindestens-einmal-pro-Woche (28 April 2026)
- Quality drop in German, arXiv 2607.19243 (single-author preprint) — https://arxiv.org/abs/2607.19243 (25 July 2026)

## The requirement profile

### What does an error cost?

- Zone 1 doesn't matter: An error costs nothing; the result is discarded at once or overwritten anyway. Assumed: €0. Source: Anthropic, Optimizing for cost and intelligence: "high-volume work with checkable outputs" — https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence. As of 19 August 2026.
- Zone 2 cheap to fix: The error is noticed and costs one more call, nothing else. Assumed: about one call. Source: Erol et al., Cost-of-Pass (arXiv 2504.13359); Anthropic: "retry the failures" — https://arxiv.org/abs/2504.13359. As of 19 August 2026.
- Zone 3 rework: The error binds human time: an escalation, a query, a correction. Assumed: €0.50 to €5. Source: Market price per resolved support ticket 0.49 to 2.00 USD (Drag, MavenAGI 2026); Zellinger and Thomson (arXiv 2507.03834): from about 0.01 USD error cost the stronger model wins — https://arxiv.org/abs/2507.03834. As of 19 August 2026.
- Zone 4 business impact: The error costs customers, deadlines or legal positions. Assumed: €50 to €500. Source: Cost-of-Pass, human reference GPQA Diamond 58 USD per task (arXiv 2504.13359); Moffatt v. Air Canada 2024 BCCRT 149 (812 CAD); Klarna rollback 05/2025 — https://arxiv.org/abs/2504.13359. As of 19 August 2026.
- Zone 5 must not happen: The error cannot be priced or is existential. A rule set in advance decides. Assumed: no figure, blocked. Source: BSI Generative AI Models v2.0 (17 January 2025) M20; AI Act Regulation (EU) 2024/1689 Art. 14, 15; Lemonade: "we never let AI perform deterministic actions" — https://www.bsi.bund.de/SharedDocs/Downloads/DE/BSI/KI/Generative_KI-Modelle.pdf. As of 19 August 2026.

### Who checks the result?

- machine-checkable: A program checks the result; errors only cost another call. Assumed: 0.0 minutes per task. Source: PAL (arXiv 2211.10435); OpenAI Structured Outputs only guarantee the shape — https://arxiv.org/abs/2211.10435. As of 19 August 2026.
- spot-checked: Humans check samples. Caution: in one study 35 to 45 percent of flawed AI drafts were sent unchanged. Assumed: 0.5 minutes per task. Source: MedStar/Georgetown, npj Digital Medicine 24 April 2025 (automation bias) — https://www.nature.com/articles/s41746-025-01586-2. As of 19 August 2026.
- every task reviewed: Every result is read. Reviewing costs reading time, and reviewers tire under load. Assumed: 2.0 minutes per task. Source: UC San Diego, JAMA Netw Open 15 April 2024, doi 10.1001/jamanetworkopen.2024.6565 (plus 21.8 percent reading time); Turan (arXiv 2606.08919): more escalation can lower safety; OpenAI GDPval: costs without oversight — https://arxiv.org/abs/2606.08919. As of 19 August 2026.

### How much thinking effort?

- brief: On simple tasks standard models beat reasoning models. Source: Shojaee et al., The Illusion of Thinking (arXiv 2506.06941) — https://arxiv.org/abs/2506.06941. As of 19 August 2026.
- normal: Medium thinking level. For Claude Opus 5 and Fable 5 the default level is high; normal lowers the output there to half (Opus 5) or about three quarters (Fable 5) according to the measurement, see mengen.json. Source: Reference level of the volume factors (mengen.json, medium equals 1). Anthropic, Optimizing for cost and intelligence: default level high, medium halves cost per task for Opus 5 (as of 19 August 2026). Artificial Analysis states tokens per index task only in prose, not per model and effort level — https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1. As of 19 August 2026.
- thorough: GPT-5 high uses 23 times the tokens of minimal; medium to high adds one index point. Source: Artificial Analysis, GPT-5 benchmarks 7 August 2025; Claude Sonnet 5 agentic cost 30 June 2026 — https://artificialanalysis.ai/articles/gpt-5-benchmarks-and-analysis. As of 19 August 2026.

The volume factor is measured for 2 models; factor 1 applies to the rest.

### How much text in view?

- paragraphs: Short context, every model copes. Up to 8,000 tokens. Source: NoLiMa (arXiv 2502.05167) — https://arxiv.org/abs/2502.05167. As of 19 August 2026.
- document: From 32,000 tokens eleven of thirteen models drop below half their short-context performance once the question does not literally match the passage. Up to 128,000 tokens. Source: NoLiMa, ICML 2025 (arXiv 2502.05167) — https://arxiv.org/abs/2502.05167. As of 19 August 2026.
- corpus: A 1M window says nothing about fidelity: Gemini 3.1 Pro drops in MRCR v2 from 84.9 percent at 128k to 26.3 at 1M; Opus 4.6 holds 76, Sonnet 4.5 18.5 percent. Up to 1,050,000 tokens. Source: DeepMind Gemini 3.1 Pro model card 19 February 2026; Anthropic Claude Opus 4.6 5 February 2026 — https://deepmind.google/models/model-cards/gemini-3-1-pro/. As of 19 August 2026.

### Language of the task

German text needs more tokens than English for the same content. The surcharge is held per model in the data.

Tokens in the GPT-5 reference measure (English text); for German text the calculator applies a per-model factor of 1.46 to 2.60.

Source: TextKit, Tokens per word (10 June 2026), Anthropic Pricing (new tokenizer from Claude 4.7, about 30 percent more tokens), arXiv 2605.24718 (German 1.55 to 1.98 per vendor). Reference: GPT-5 tokenizer English = 1.0, rounded down conservatively. — https://textkit.tech/blog/tokens-per-word-tokenizer-comparison-2026. As of 19 August 2026.

### Constraints

Processing in the EU, Fast response mode. Fast response mode: 2× input and output, cannot be combined with batch processing. With processing in the EU, hosters of third-party models without a published EU price cannot win the case; providers with their own EU price stay in and are charged their surcharge.

OpenRouter credit fee 5.5%: The fee falls due when credit is topped up and does not sit on the individual call. Other routes, the Vercel Gateway among them, charge nothing for it.

## Zone 5: when an error must not happen

The workhorse class carries the proposing role, bound to a citation duty pointing at the source passage. The frontier class joins it as soon as the anchor of this case reports a measured solve rate.

### What the process has to deliver

In this zone a rule you set in advance makes the decision, and the language model only supplies the proposal. Have every release confirmed by a fixed check in code or by a second person before it takes effect. Arithmetic, posting and reconciliation belong in tools that return the same result for the same input. Record input, model version and output together so the decision stays traceable later. Budget the effort for this review layer from the start, because it carries the reliability of the result.

### The 13 architecture patterns

- Compile once, then run fixed (Recurring, regulated high-volume processes). Determines: The entire runtime path. No model call happens in operation any more. Does not determine: The quality of the artefact produced once. Runtime flexibility is gone. Source: Compiled AI (arXiv 2604.05150, 6 April 2026): "execute deterministically without further model invocation", 57 times fewer tokens across 1,000 transactions — https://arxiv.org/abs/2604.05150. As of 19 August 2026.
- Formal verification with a solver (Safety-critical code, rule derivations). Determines: Correctness against the formal specification, hard. Does not determine: Whether the specification asks for the right thing. The translation step runs in the model itself. Source: Dantas et al., The 4/δ Bound (arXiv 2512.02080, 30 November 2025): over 90,000 runs, every run reached verification; AWS Bedrock Automated Reasoning: "Because this step uses LLMs, it may contain errors" for the translation, "mathematically sound" only for the check — https://arxiv.org/abs/2512.02080. As of 19 August 2026.
- Proposal plus deterministic checker (Payment release, code release, policy compliance). Determines: The decision. The checker is a program and returns the same verdict for the same input. Does not determine: The proposal itself, and everything the checking rule does not cover. Source: Sun et al., Agentic Model Checking (arXiv 2605.21434, 20 May 2026): "agents propose, solvers verify"; every output that could change a verdict passes through the checking chain — https://arxiv.org/abs/2605.21434. As of 19 August 2026.
- Policy as code before execution (Approval gates ahead of irreversible steps). Determines: Whether an action complies with the rules, repeatable across runs. Does not determine: Whether the rule set is complete. Model-written rules reached 70.96 percent recall. Source: SOCpilot (arXiv 2605.05501, 6 May 2026), financial-sector SOC with 200 real incidents: the verifier removed 466 non-compliant, approval-bound actions, "Aggregate rates remain stable across 3 reruns"; AgentSpec (arXiv 2503.18666, ICSE 2026) — https://arxiv.org/abs/2605.05501. As of 19 August 2026.
- The tool computes, the model only writes the call (Amounts, deadlines, posting logic). Determines: The computed result. Does not determine: Whether the right program was chosen. Source: Gao et al., PAL (arXiv 2211.10435): "offloads the solution step to a runtime such as a Python interpreter", because models "often make logical and arithmetic mistakes in the solution part" — https://arxiv.org/abs/2211.10435. As of 19 August 2026.
- Capability interpreter over the data flow (Agents with write access, payout paths). Determines: Which actions are possible at all, regardless of what the model wants. Does not determine: Task completion. Coverage fell from 84 to 77 percent. Source: Debenedetti et al., CaMeL (arXiv 2503.18813, DeepMind and ETH): "solving 77% of tasks with provable security (compared to 84% with an undefended system)" — https://arxiv.org/abs/2503.18813. As of 19 August 2026.
- Fixed blueprint, model only for sub-steps (Case handling with a fixed process). Determines: The process path. The model does not choose the direction of the case. Does not determine: The sub-steps themselves. Those stay probabilistic. Source: Blueprint First, Model Second (arXiv 2508.02721): violations of the rules cut by 96.0 percent, 11 instead of 275 — https://arxiv.org/abs/2508.02721. As of 19 August 2026.
- Schema-bound output (Machine-readable handover to the next stage). Determines: The form: syntax, schema, permitted values of an enumeration. Does not determine: The content. Values inside a valid schema can be invented. Source: OpenAI Structured Outputs: "Structured Outputs can still contain mistakes."; Azure AI Foundry with strict schema does not support minLength, pattern, minimum and maximum — https://developers.openai.com/api/docs/guides/structured-outputs. As of 19 August 2026.
- Citation duty pointing at the passage (Answers that must cite a source). Determines: That the pointer into the source is valid. Does not determine: Whether the statement is actually supported by that passage. Source: Anthropic Citations: "citations are guaranteed to contain valid pointers to the provided documents"; citations and schema-bound output are mutually exclusive — https://platform.claude.com/docs/en/build-with-claude/citations. As of 19 August 2026.
- Freeze and version the answer (Auditable documents that have to stay stable). Determines: Repeatability from the second retrieval onwards. Does not determine: The correctness of the frozen answer. Source: No primary source found (research run 19 August 2026). Common practice, not documented here. (no address). As of 19 August 2026.
- Human release, by two people where required (Payment release, diagnosis, sanction decisions). Determines: The release as a legal act, and accountability for it. Does not determine: The attention of that person. Automation bias takes effect precisely here. Source: AI Act Regulation (EU) 2024/1689 Art. 14(5) requires "at least two natural persons" only for remote biometric identification; BSI Generative AI Models v2.0 R8 on automation bias — https://artificialintelligenceact.eu/article/14/. As of 19 August 2026.
- Sample several times and vote (Raise accuracy, never report it as determinism). Determines: Nothing. Variance drops, repeatability does not arise from it. Does not determine: Everything. The method deliberately samples several different reasoning paths. Source: Wang et al., Self-Consistency (arXiv 2203.11171): "It first samples a diverse set of reasoning paths instead of only taking the greedy one", plus 17.9 points on GSM8K — https://arxiv.org/abs/2203.11171. As of 19 August 2026.
- Filters and classifiers as a guardrail (Coarse filter ahead of the actual check). Determines: Only while the block is a rule. If the block is itself a model, it varies too. Does not determine: Completeness. Filters can be bypassed and do not work as the last authority. Source: BSI M13: filters can be bypassed through encoded outputs; AWS Well-Architected Agentic AI Lens (10 June 2026), AGENTSEC04-BP02: the risk classifier in front of the human gate must not itself be a language model — https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentic-ai-lens.html. As of 19 August 2026.

### What the supervisors require

- Scope: Art. 12, 14 and 15 bind high-risk systems only, under Annex III from 2 December 2027 and under Annex I from 2 August 2028 (postponed by the amending Regulation (EU) 2026/1744, Digital Omnibus, in force since 27 July 2026). The transparency duties of Art. 50 have applied since 2 August 2026. In Germany the Federal Network Agency has run market surveillance under the KI-MIG since 29 July 2026. Whoever builds to these articles today builds towards the deadline, not against a duty already in force. Source: Regulation (EU) 2024/1689 Art. 113, consolidated version 27 July 2026 (artificialintelligenceact.eu); BaFin, press release 29 July 2026: transparency duties from 2 August 2026, high-risk requirements from 2 December 2027 — https://artificialintelligenceact.eu/article/113/. As of 19 August 2026.
- AI Act Art. 12(1): the system must technically allow automatic recording of events over its lifetime. Without a log, a decision cannot be evidenced later. Source: Regulation (EU) 2024/1689, consolidated version 27 July 2026 — https://artificialintelligenceact.eu/article/12/. As of 19 August 2026.
- AI Act Art. 14(4): the overseeing person must remain aware of automation bias, be able to disregard, override or reverse the output, and interrupt operation with a stop button. Source: Regulation (EU) 2024/1689, consolidated version 27 July 2026 — https://artificialintelligenceact.eu/article/14/. As of 19 August 2026.
- AI Act Art. 15(3) and (4): accuracy levels are declared in the instructions for use, and robustness may be achieved through technical redundancy with backup or fail-safe plans. The error rate therefore has to be named. Source: Regulation (EU) 2024/1689, consolidated version 27 July 2026 — https://artificialintelligenceact.eu/article/15/. As of 19 August 2026.
- Scope of the two-person rule: Art. 14(5) requires two natural persons only for remote biometric identification under Annex III(1)(a). It does not say that for creditworthiness or payment release. Source: Regulation (EU) 2024/1689, consolidated version 27 July 2026; a full-text search does not find the terms determinism and reproducibility in the regulation — https://artificialintelligenceact.eu/article/14/. As of 19 August 2026.
- BSI risk R7: even for identical input, the generated output can differ. Measure M20 requires review, cross-referencing with further sources and manual post-processing before any further use where implications may be critical. Source: BSI, Generative AI Models: Opportunities and Risks, version 2.0, 17 January 2025, chapter 4.1 and M20 — https://www.bsi.bund.de/SharedDocs/Downloads/DE/BSI/KI/Generative_KI-Modelle.pdf. As of 19 August 2026.
- NIST frames the topic as confabulation: invented content, confidently presented, with the example of a confabulated patient summary. The critical-infrastructure concept note turns the requirement around: such processes need deterministic behaviour, explainability and fail-safe operation. Source: NIST AI 600-1, Generative AI Profile, July 2024, section 2.2; NIST, Concept Note AI RMF Profile for Critical Infrastructure, 7 April 2026 — https://www.nist.gov/programs-projects/concept-note-ai-rmf-profile-trustworthy-ai-critical-infrastructure. As of 19 August 2026.
- BaFin, press release of 29 July 2026: decisions must remain correctable and reversible by humans, and responsibility stays with the supervised firms and their management. Its guidance of 18 December 2025 names limits on model use for critical functions and a human review duty. Source: BaFin, supervision of AI: BaFin receives new competences, 29 July 2026; BaFin, guidance on ICT risks in the use of AI at financial firms, 18 December 2025, p. 31 — https://www.bafin.de/SharedDocs/Veroeffentlichungen/DE/Pressemitteilung/2026/pm_2026_07_29_ki_verordnung.html. As of 19 August 2026.
- Germany: the national act implementing the AI Regulation has applied since 29 July 2026. The Federal Network Agency runs market surveillance, with BaFin taking it over for the financial sector. Source: German AI Market Surveillance and Innovation Promotion Act (KI-MIG), Federal Law Gazette 2026 I No. 223, in force 29 July 2026 — https://www.bafin.de/SharedDocs/Veroeffentlichungen/DE/Pressemitteilung/2026/pm_2026_07_29_ki_verordnung.html. As of 19 August 2026.

### How practice solves it

- Lemonade derives the ban from the system property: "AI is non-deterministic and has been shown to have biases across different communities. That’s why we never let AI perform deterministic actions such as rejecting claims or canceling policies." Rejection and cancellation stay rule-bound acts. Source: Lemonade, Claim automation, blog (no printed date) — https://www.lemonade.com/blog/lemonades-claim-automation/. As of 19 August 2026.
- The US Medicare documentation rule has named AI explicitly since July 2025: if you use a scribe, including AI, you sign the entry yourself. The scribe does not sign, and without a signature the associated claims may be denied. Source: CMS MLN905364, Complying with Medicare Signature Requirements, July 2025 — https://www.cms.gov/files/document/mln905364-complying-medicare-signature-requirements.pdf. As of 19 August 2026.
- California SB 1120, signed 28 September 2024: a software tool must not deny, delay or modify health care services based on medical necessity. That determination is made by a licensed professional. Source: California SB 1120, Physicians Make Decisions Act, Health & Safety Code § 1367.01 — https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240SB1120. As of 19 August 2026.
- Allianz, Zurich and Swiss Re commit publicly in the same way: the model supplies recommendations and drafts, while coverage and payout are decided by the handler. Zurich requires human review for all consequential decisions and names itself where that pattern breaks down for autonomous agents. Source: Allianz, Smarter claims management, 5 February 2025; Zurich, How do we navigate a safe AI transformation (H2 2025); Swiss Re Corporate Solutions, ClaimsGenAI, 20 June 2025 — https://www.allianz.com/en/mediacenter/news/articles/250205-smarter-claims-management-smoother-settlements.html. As of 19 August 2026.
- In payments the architecture separates model and execution: the agent receives a token, while authorisation runs at the payment provider or in the card network and is checked against the user instruction. Source: Agentic Commerce Protocol (Stripe and OpenAI), 29 September 2025: "The use of the token is programmatically controlled, permissioned, and logged"; Visa Intelligent Commerce — https://stripe.com/blog/developing-an-open-standard-for-agentic-commerce. As of 19 August 2026.
- Bloomberg publishes both figures: 99 percent of AI summaries meet the editorial standard, and by the end of March 2025 at least 36 had to be corrected. Journalists decide whether a summary appears, and every point jumps to its passage in the transcript. Source: Bloomberg, AI-Powered Earnings Call Summaries, 22 January 2024; NYT report of 29 March 2025 (documented through two independent secondary accounts) — https://www.bloomberg.com/company/press/bloomberg-launches-ai-powered-earnings-call-summaries/. As of 19 August 2026.
- In a bank security operations centre with 200 real incidents, the same policy given as prompt text moved two vendors in opposite directions. Placed in front as a program, it removed 466 non-compliant, approval-bound actions without reducing baseline-task recall. Source: SOCpilot (arXiv 2605.05501, 6 May 2026), anonymised production SOC in the financial sector — https://arxiv.org/abs/2605.05501. As of 19 August 2026.
- AWS writes the split as a rule: deterministic controls for everything expressible deterministically, probabilistic controls for the rest. The risk classifier that decides on the human gate must not itself be a language model. Source: AWS Well-Architected Agentic AI Lens, 10 June 2026, AGENTSEC04-BP02 and AGENTSEC05-BP01 — https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentic-ai-lens.html. As of 19 August 2026.

## What the literature recommends

### Answering a support ticket

Recommendation: Workhorse class as the default, compact class as the floor (Claude Sonnet 5, Gemini 3.7 Flash, gpt-5.5, Claude Haiku 4.5).

The floor is the fast class of the strong vendors. Once the answer reaches the customer, the workhorse class is the default. τ²-bench spreads the models from 36 to 79 percent solve rate, and a wrong answer binds the company: Moffatt v. Air Canada ended at 812 CAD for one sentence. For the classes below Haiku and Flash no customer service measurement exists, in either direction. Caching the policies cuts the bill further than a step down in class: a cache read at Anthropic costs one tenth of the input price (pricing documentation, as of 18 August 2026).

Sources:

- Anthropic, Choosing the right model — https://platform.claude.com/docs/en/about-claude/models/choosing-a-model (19 August 2026)
- OpenAI, Model selection guide — https://developers.openai.com/api/docs/guides/model-selection (19 August 2026)
- tau2-bench leaderboard, pass rate — https://www.codesota.com/benchmark/tau2-bench (19 August 2026)
- Moffatt v. Air Canada, 2024 BCCRT 149 — https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html (14 February 2024)

As of 19 August 2026.

### Reading an email and drafting the reply

Recommendation: Compact class (Claude Haiku 4.5, Gemini 3.5 Flash-Lite, gpt-5.4-mini).

The text is supplied, the task is single-step and checkable, and that is exactly what the vendors position their cheapest models for. Google describes Flash-Lite as a model for high throughput, Anthropic describes Haiku 4.5 for volume work with checkable output. The HHEM hallucination measure puts Flash-Lite, nano and mini at 3 to 6 percent on this task; Claude Haiku 4.5 sits at 9.8 percent, gpt-5.4 at 7.0 (as of 11 May 2026). Price may decide here, as long as a sample is read back.

Sources:

- Google, Gemini API models — https://ai.google.dev/gemini-api/docs/models (19 August 2026)
- Anthropic, Optimizing for cost and intelligence — https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence (19 August 2026)
- Vectara, Hallucination Leaderboard — https://github.com/vectara/hallucination-leaderboard (11 May 2026)

As of 19 August 2026.

### Writing or revising a text

Recommendation: Workhorse class (Claude Sonnet 5, Gemini 3.7 Flash, gpt-5.6-terra).

The telemetry shows what people reach for when writing: on Claude.ai 10 percent of conversations run on the frontier class, in the coding environment 54 percent. Whoever produces text works in the middle class. Anthropic assigns writing to the Sonnet class explicitly. The expensive error here is the silent omission and the break in register, and both are caught by the person who reads the text anyway before using it.

Sources:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- Anthropic, Choosing the right Claude model — https://claude.com/resources/tutorials/choosing-the-right-claude-model (19 August 2026)

As of 19 August 2026.

### Sales email and proposal draft

Recommendation: Workhorse class, with proof of belonging to the top writing group (Gemini 3.7 Flash, Claude Sonnet 5, Claude Fable 5).

For a proposal, writing quality is the buying argument, and for that an anchor is retrievable. On the arena.ai creative writing board Gemini 3.7 Flash holds rank 3 with 1493 points, 13 points behind the leader Claude Fable 5 at one thirteenth of the input price. Confidence intervals of ±7 to ±18 points in the top group (±5 to ±28 across the whole field) make the leading ranks indistinguishable, so the value works as a threshold and price decides inside the group. Claude Haiku 4.5 sits at rank 112 with 1389 points and therefore has no place in outward-facing text.

Sources:

- arena.ai, Creative Writing Leaderboard — https://arena.ai/leaderboard/text/creative-writing (12 August 2026)
- The Leaderboard Illusion, arXiv 2504.20879 — https://arxiv.org/abs/2504.20879 (29 April 2025)

As of 19 August 2026.

### Classification and routing

Recommendation: Smallest class that still carries (gpt-5.4-nano, Gemini 3.1 Flash-Lite, Claude Haiku 4.5).

Classifying and assigning is the cheapest case in the whole catalogue and the case with the clearest production evidence for small models. LinkedIn measures 30.6 percentage points more routing accuracy in a production A/B test, Gelato lifts ticket assignment from 60 to 90 percent. The failure mode is calibration, and the answer to it is a confidence value per field with handback to a person. A larger model does not solve that problem.

Sources:

- LinkedIn, production A/B test, arXiv 2608.10224 — https://arxiv.org/abs/2608.10224 (10 August 2026)
- Small Language Models are the Future of Agentic AI, arXiv 2506.02153 — https://arxiv.org/abs/2506.02153 (2 June 2025)

As of 19 August 2026.

### Extracting data from documents

Recommendation: Compact class with targeted retrieval, a specialist model as the alternative outside the price list (Claude Haiku 4.5, Gemini 3.7 Flash).

The measurements run against the marketing: at the data cut of the AI Index 2026, Vals CorpFin v2 reached a best value of 68.26 percent across 200 pages of credit agreement with no model above 70; the vals.ai leaderboard names Claude Opus 5 at 73.19 percent on top as of 12 August 2026, MortgageTax stays at 69.4 percent. A self-hosted specialist model beat five frontier models at contract extraction with 78 to 97 percent lower inference cost and higher precision, and that route appears in no price list. For the API route the rule is targeted retrieval, a window kept as small as the task allows, and a completeness check, because the typical error is the silent omission.

Sources:

- Vals AI, CorpFin v2 — https://www.vals.ai/benchmarks/corp_fin_v2 (12 August 2026)
- Domain-trained MoE model in contract extraction, arXiv 2605.05532 — https://arxiv.org/abs/2605.05532 (7 May 2026)

As of 19 August 2026.

### Knowledge-base question

Recommendation: Middle class with strict retrieval and a citation requirement (Claude Sonnet 5, Gemini 3.7 Flash, gpt-5.6-terra).

What decides here is the citation requirement, and it costs no model upgrade. Open knowledge questions hallucinate at 22 to 94 percent across 26 models, while grounded answers fail invisibly: in a production system 7.8 percent of cases attributed real evidence to the wrong entity, rising to 13.6 percent for newer drugs. No existing hallucination or citation check finds this error type, so the sample against the source passage belongs in the process.

Sources:

- Deceptive Grounding, arXiv 2607.09349 — https://arxiv.org/abs/2607.09349 (10 July 2026)
- Stanford AI Index 2026, section 3.2 (AA-Omniscience) — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)

As of 19 August 2026.

### Research agent over documents

Recommendation: Frontier class as orchestrator, workhorse class as worker, plus a deterministic check (Claude Opus 5, Claude Sonnet 5, gpt-5.6-sol).

The only evidence that measures research specifically places Claude Opus 5 ahead of Claude Fable 5, and Anthropic recommends Opus 5 as the starting point for agents. For web search gpt-5.6-sol leads with 92.2 percent BrowseComp, Claude Sonnet 5 reaches 84.7 percent at half the list price. More important than any model choice is the checking layer: a liveness test on dead links cuts invented sources below 1 percent and costs less than any upgrade. And lower the effort level first, because Anthropic measures that low matched an orchestrator setup at 20 percent lower cost.

Sources:

- Anthropic, Optimizing for cost and intelligence — https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence (19 August 2026)
- BrowseComp, values via BenchLM — https://benchlm.ai/benchmarks/browsecomp (18 August 2026)
- Fabricated and dead citations in deep research agents, arXiv 2604.03173 — https://arxiv.org/abs/2604.03173 (3 April 2026)

As of 19 August 2026.

### Writing or changing code

Recommendation: Workhorse class, compact class in review (gpt-5.6-terra, Claude Sonnet 5, Claude Haiku 4.5).

GitHub Copilot places diff review with the general-purpose models, and the only controlled study on it measures an inverse relation: Claude Haiku 4.5 beats Claude Sonnet 4.6 with higher F1 and 18 percent more recall at 3.2 times lower cost per review. The strongest lever sits outside model choice: diff size determines the outcome, F1 falls from 0.657 below 10 lines to 0.043 above 150 lines. Whoever cuts the diff and has it reviewed by a model other than the author gains more than from any step up in class.

Sources:

- GitHub Copilot, Choosing the right AI model — https://docs.github.com/en/copilot/reference/ai-models/choosing-the-right-ai-model-for-your-task (19 August 2026)
- Bigger Isn't Always Better, arXiv 2606.15689 — https://arxiv.org/abs/2606.15689 (9 April 2026)
- CR-Bench, arXiv 2603.11078 — https://arxiv.org/abs/2603.11078 (10 March 2026)

As of 19 August 2026.

### Working through a long document

Recommendation: Workhorse class with a flat price across the full length (Gemini 3.7 Flash, Claude Sonnet 5, Claude Opus 5).

Two quantities decide, and window size is not one of them. The first is the price tier: OpenAI reprices the entire request above 272,000 tokens, Google doubles above 200,000 tokens in the Pro line, Anthropic dropped the surcharge on 13 March 2026 and Gemini Flash has no tier at all. The second is context fidelity: the same model name falls from 84.9 percent at 128k to 26.3 percent at 1M, and under the same 1M label stand 76 percent against 18.5 percent. For 300,000 tokens no public test measures, and that gap stays open.

Sources:

- Google DeepMind, Gemini 3.1 Pro model card — https://deepmind.google/models/model-cards/gemini-3-1-pro/ (19 February 2026)
- Anthropic, Claude Opus 4.6 — https://www.anthropic.com/news/claude-opus-4-6 (5 February 2026)
- Anthropic, Pricing (long context without surcharge) — https://platform.claude.com/docs/en/about-claude/pricing (19 August 2026)

As of 19 August 2026.

### Meeting notes

Recommendation: Middle class, the bottleneck is transcript quality (Claude Sonnet 5, Gemini 3.7 Flash, Claude Haiku 4.5).

The time saving is best documented here: 83 percent less note-writing effort at Sharp HealthCare, 24 percent less documentation time and 11.3 more patients per month at Northwestern, 20 minutes per half day of clinic in a JAMIA study with 48 physicians. The errors are hallucination, omission and irrelevance, and the participants spot them at once because they were there. A larger model does not lift a poor transcript.

Sources:

- Stanford AI Index 2026, medicine chapter — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf (19 August 2026)
- Failure modes of meeting summarisation, arXiv 2509.15901 — https://arxiv.org/abs/2509.15901 (19 September 2025)

As of 19 August 2026.

### Translation

Recommendation: Compact class (Gemini 3.7 Flash, Claude Haiku 4.5, gpt-5.4-mini).

In the telemetry translation carries the lowest autonomy level of all artefacts, and a person reads it back anyway. So price decides, and specifically the output price, because input and output are about the same size here. The German tokenizer surcharge applies to both sides: the same text produces 1.71 tokens per word on GPT-5 and 3.48 on Claude Opus 4.8. For the quality drop in German only a single-author preprint exists, and that is a thin evidence base.

Sources:

- Anthropic Economic Index, June 2026 — https://www.anthropic.com/research/economic-index-june-2026-report (26 June 2026)
- TextKit, Tokens per word, tokenizer comparison 2026 — https://textkit.tech/blog/tokens-per-word-tokenizer-comparison-2026 (10 June 2026)

As of 19 August 2026.

## Knowledge

### What is a token?

A token is the chunk of text a model works in, usually a short word or a syllable, and also the billing unit at every vendor.

The rule of thumb is one token per four characters of English, roughly 750 words per 1,000 tokens. German packs more densely: compounds and umlauts split into more pieces, so the same content costs a little more.

Every price on this page refers to 1M tokens. That sounds like a lot and fills up fast: a single 500 kB PDF already runs about 125,000 tokens, an eighth of it.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing

As of 18 August 2026.

### Input and output

Input is everything you send along; output is what the model writes back, and per token it usually costs five times as much.

Input covers system instructions, prior conversation and attached documents, and it is billed again on every single call. A typical 10 kB web page runs about 2,500 tokens.

Output includes reasoning steps, even when you never see them. Tasks with long output, drafts for example, shift the largest cost item from input to output.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing

As of 18 August 2026.

### The cache

The cache stores the part of your input that repeats across calls; it is written once at a premium and read cheaply afterwards.

The per-model prices sit in the model table: Anthropic charges 1.25 times the input price to write and one tenth to read, OpenAI charges the same 1.25 times to write since the gpt-5.6 series, Mistral has read for one tenth since August 2026, DeepSeek for about three percent.

This has a consequence that gets overlooked: on a single call, caching is MORE expensive than none, because it is written and never read. The benefit starts with the second call of the same task.

Every model and every effort level keeps its own cache. Switching mid-task forces a fresh write.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · Mistral, prompt caching documentation — https://docs.mistral.ai/studio-api/conversations/advanced/prompt-caching · DeepSeek, pricing page — https://api-docs.deepseek.com/quick_start/pricing

As of 18 August 2026.

### Calls per task

A task is rarely a single call: an agent that searches, reads and plans easily reaches ten or more.

A simple chat is one call. Once tools, search steps or intermediate checks join in, the number multiplies, and with it the input that is billed anew on every call.

This value stays easy to overlook: token prices sit on the pricing page, the number of calls sits nowhere and has to be measured.

Own assessment, no external source.

As of 19 August 2026.

### Attempts until it holds

How often you restart on average until the result holds: every discarded attempt is billed in full.

This is where a cheap model proves whether it is actually cheaper. In research from March 2026, the cheaper-listed model produced the higher total cost in 32 percent of model pairs, mostly through highly variable thinking tokens and more working steps per task.

The break-even in the result works this out for your scenario: the number of attempts at which the cheapest model loses its lead to the priciest.

Where an anchor supplies a solve rate, the calculator derives the attempts from it, with a correction. Yang (arXiv 2605.08563, 8 May 2026) shows that assuming independent tries puts pass@3 17.4 points too high, 98.6 percent against 81.2 percent: a second attempt inherits the failure cause of the first. The calculator therefore applies a factor of 1.2 to the attempts. That factor is a stated assumption drawn from this measurement, not a measurement of our own.

Sources: Chen et al. 2026, arXiv 2603.23971 — https://arxiv.org/abs/2603.23971 · Yang, arXiv 2605.08563, 8 May 2026 — https://arxiv.org/abs/2605.08563

As of 19 August 2026.

### The second price tier

Nine models get more expensive once a single call exceeds a certain amount of text, 272,000 tokens at OpenAI and 200,000 at Google.

The higher rate then applies to ALL tokens of the call, including those below the threshold. That is how the vendors bill, and the thresholds often sit in a footnote below the pricing table.

A calculator that flatly uses the page-one price therefore underestimates long documents by up to half. Anthropic does not tier and offers the full window at the base price.

The "Trait" filter in the model table shows which nine models are affected; in the result, affected rows carry the marker "long-context rate".

Sources: OpenAI, pricing page — https://developers.openai.com/api/docs/pricing · Google, Gemini pricing page — https://ai.google.dev/gemini-api/docs/pricing

As of 18 August 2026.

### The context window

The context window is the largest amount of text a single call can take; a cheap model is no use if the task does not fit.

The range is wide: Claude Haiku 4.5 takes up to 200,000 tokens, the current Opus and gpt-5.6 models around 1M. In the result, models whose window is too small for your input carry the marker "does not fit".

A dash in the model table means the vendor publishes no size for this model on its pricing page. No estimated number stands in for it.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing

As of 18 August 2026.

### Batch processing

If you can wait a few hours for the answer, most vendors take 50 percent off both input and output.

Anthropic, OpenAI, Google and Mistral publish the batch discount as its own price list; the open-weights providers in our set list none. Where it is missing, this calculator drops the discount and says so at the number.

For recurring volume work, overnight document processing for example, that is 50 percent off; quality stays the same and only the wait is added.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing · Google, Gemini pricing page — https://ai.google.dev/gemini-api/docs/pricing

As of 18 August 2026.

### Routers and data residency

Routers such as OpenRouter bundle many models behind one entry point; the fee sits on credit top-ups, and the data question weighs more than the fee.

Routers apply no markup to the tokens themselves; top-ups carry 5.5 percent by card. Requests are forwarded to whichever provider is selected, and that provider may use them for training or improvement. A fixed subprocessor list requires pinning the routing down.

Forced US processing costs 1.1 times the standard rate at Anthropic. The residency parameter currently knows no EU value; EU processing runs through the cloud platforms offering European regions.

Sources: OpenRouter, pricing page — https://openrouter.ai/pricing · Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing

As of 18 August 2026.

### The five model classes

From on-device edge models to the frontier class: the class says what a vendor built a model for, and price alone does not.

The mapping follows vendor naming: Anthropic tiers Opus, Sonnet, Haiku; OpenAI sol, terra, luna plus the pro tiers; Google Pro, Flash, Flash-Lite; Mistral Large, Medium, Small, plus the Ministral line explicitly for on-device use and Codestral and Devstral for code.

The class sorts, the capability column measures. The two can diverge: DeepSeek V4 Pro is its vendor’s frontier class at a compact-class price, and an edge-class model can top list-price rankings while its small context window and missing capability measurement rule it out for many business cases.

Only models bookable at an active endpoint with a published list price enter the set. Circulating headline prices without a bookable endpoint stay out, however tempting the number looks.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing · Google, Gemini pricing page — https://ai.google.dev/gemini-api/docs/pricing · Mistral, API pricing page — https://mistral.ai/pricing/api

As of 18 August 2026.

### Capability, difficulty, solve rate

Whether a model is enough for a task is measurable: capability and difficulty sit on one scale, and their gap sets the solve rate.

The Epoch Capabilities Index (ECI) aggregates over 50 benchmarks into one capability value per model via item response theory, the same tooling that calibrates exams such as the GMAT. The scale is relative: GPT-5 sits at 150 by definition, Claude 3.5 Sonnet at 130.

The economic consequence is cost-of-pass: cost per solved task equals cost per attempt divided by the solve rate. A model just below the task difficulty gets expensive through attempts; far above it you pay for capability the task does not need.

The calculator uses this for business cases with an anchor benchmark, and only with measured solve rates: where no measurement exists, a dash stands instead of an estimated number, and from zone 3 upwards the row cannot win. Where no anchor holds, the report says so openly.

Sources: Epoch AI, Epoch Capabilities Index (CC-BY) — https://epoch.ai/eci · Ho et al. 2025, A Rosetta Stone for AI Benchmarks, arXiv 2512.00193 — https://arxiv.org/abs/2512.00193 · Erol et al. 2025, Cost-of-Pass, arXiv 2504.13359 — https://arxiv.org/abs/2504.13359

As of 18 August 2026.

### Prices and comparability

Every price here is a vendor list price per 1M tokens, and tokens are only approximately comparable across vendors.

Verified 18 August 2026 directly against vendor pricing pages, without discounts. Prices change silently: between two checks of this calculator, one model temporarily sat five times too high in the list, with no notice anywhere.

Anthropic documents a tokeniser for Claude 4.7 and later that produces roughly 30 percent more tokens for the same text. A price comparison across vendors therefore only holds as an order of magnitude.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · OpenAI, pricing page — https://developers.openai.com/api/docs/pricing

As of 18 August 2026.

### What an error costs

Once an error costs more than about one US cent, the error side of the calculation weighs more than the token price.

Zellinger and Thomson put a number on the threshold: “reasoning models offer better accuracy-cost tradeoffs as soon as the economic cost of a mistake exceeds $0.01”. Model cascades lose their advantage from about $0.1. One precondition belongs with it: the threshold holds as long as the stronger model is also the more accurate one. For code review and research the measurements often run the other way; the threshold logic still holds there, the ranking behind it has to be measured.

In customer service the two figures sit two orders of magnitude apart. A ticket of 3,700 tokens costs fractions of a cent in every model class, while the market price of a resolved ticket runs 0.49 to 2.00 USD (Drag and MavenAGI 2026, secondary sources). Saving on the token price while leaving out escalation optimises the smaller of the two numbers.

Two documented cases show the upper end. In Moffatt v. Air Canada (2024 BCCRT 149, 14 February 2024) the tribunal awarded 812.02 CAD and stated: “It should be obvious to Air Canada that it is responsible for all the information on its website.” Klarna announced in February 2024 the work of a calculated 700 full-time agents and $40 million in profit improvement, and corrected course in May 2025: “As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”

The requirements section is where you set this figure as a zone. The presets run from 0 euro through the cost of one more call up to 100 euro per task. The top zone carries no figure, because the error cannot be priced there, and it switches the result to model group and process.

Sources: Zellinger and Thomson, Economic Evaluation of LLMs, arXiv 2507.03834 — https://arxiv.org/abs/2507.03834 · Drag and MavenAGI 2026, market price per resolved ticket (secondary source) — https://www.dragapp.com/blog/state-of-ai-support-pricing/ · Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024 — https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html · Klarna, announcement 27 February 2024 and correction 05/2025 (secondary source) — https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396

As of 19 August 2026.

### Checking costs time and catches part of it

The review layer is a cost item of its own: it costs reading time, and it finds fewer errors than the calculation usually assumes.

A simulation study by MedStar and Georgetown (npj Digital Medicine, 24 April 2025) put 18 patient messages with AI drafts in front of 20 primary care physicians, four of them carrying planted errors. Between 35 and 45 percent of the flawed drafts went out unchanged, every flawed draft was missed by at least 13 of the 20 participants, on average 2.67 of 4 errors went unnoticed, and exactly one participant addressed all four. The authors name the cause as automation bias, “the tendency to over-rely on automation”.

Checking costs measurable time. At UC San Diego reading time per message rose by 21.8 percent (JAMA Network Open, 15 April 2024, p equals 0.008), while reply time fell by 5.9 percent without statistical significance. The calculator therefore applies review minutes times an hourly rate, with 60 euro per hour as a stated assumption.

Attention is a finite resource. Turan (arXiv 2606.08919, 8 June 2026) models reviewers who tire as the escalation load grows: “when the reviewer is modeled as endogenous (fatiguing as escalation load grows), realized safety becomes an inverted-U in the escalation rate: more human oversight can make a system less safe”. Reviewers also agree only moderately on what counts as risky, Fleiss kappa 0.52. The work is a single-author preprint without a human study; it stands here as a caution and does not enter the calculation.

OpenAI writes about GDPval (25 September 2025) that frontier models handle those tasks around 100 times faster and around 100 times cheaper than industry experts, and clarifies in the same passage that these figures cover inference time and API cost only and leave out human oversight, revision and integration. Those are the items the reviewability control brings into the calculation.

Sources: MedStar and Georgetown, npj Digital Medicine, 24 April 2025 — https://www.nature.com/articles/s41746-025-01586-2 · UC San Diego (Tai-Seale et al.), JAMA Network Open, 15 April 2024 — https://doi.org/10.1001/jamanetworkopen.2024.6565 · Turan 2026, Oversight Has a Capacity, arXiv 2606.08919 — https://arxiv.org/abs/2606.08919 · OpenAI, GDPval, 25 September 2025 — https://openai.com/index/gdpval/

As of 19 August 2026.

### The price reversal

In 32 percent of model pairs the cheaper-listed model produces the higher actual cost, because it needs more tokens and more turns.

Chen et al. (arXiv 2603.23971, 25 March 2026, 8 models, 12 tasks) measure: “in 32% of model-pair comparisons, the model with a lower listed price actually incurs a higher total cost, with reversal magnitude reaching up to 28x”. On the same query one model burns up to 900 percent more thinking tokens than another, or ten times as many turns. Repeating the same query on the same model, thinking tokens vary by up to 9.7 times; the paper calls that an irreducible noise floor for any predictor.

OckBench (arXiv 2511.05722) measures the same effect at equal accuracy: more than a 25-fold token difference, about 1,600 against about 42,000 tokens. The paper puts it this way: “cheaper token cost does not always imply cheaper task cost; verbose smaller models can pay an Overthinking Tax”.

The effort level works at the same magnitude. Artificial Analysis measured GPT-5 at high effort using 23 times the tokens of minimal, 82M against 3.5M for the whole index, at 68 against 44 index points; the step from medium to high added one point. The generation counts too: Claude Sonnet 5 produces around 40 percent more output tokens per index task than Sonnet 4.6, needs around three times as many agent turns and costs around twice as much per task, at the same list price.

This version therefore calculates with a volume factor per model and effort level and shows the result as a range. Where no measured spread exists for the upper band, that stands at the figure.

Sources: Chen et al. 2026, arXiv 2603.23971 — https://arxiv.org/abs/2603.23971 · OckBench, arXiv 2511.05722 — https://arxiv.org/abs/2511.05722 · Artificial Analysis, GPT-5 benchmarks and analysis, 7 August 2025 — https://artificialanalysis.ai/articles/gpt-5-benchmarks-and-analysis · Artificial Analysis, Claude Sonnet 5 agentic cost, 30 June 2026 — https://artificialanalysis.ai/articles/claude-sonnet-5-agentic-cost

As of 19 August 2026.

### The tokeniser as a volume factor

The same German text yields a different token count per vendor, and that shifts the calculation before a single price is compared.

Anthropic documents a new tokeniser for models from Claude 4.7 on: “This tokenizer produces approximately 30% more tokens for the same text.” At the same price per token that is a surcharge of roughly 30 percent on the volume side, without any price figure changing.

Across vendors the range is wider. On 10 June 2026 TextKit measured German text at 1.71 tokens per word for GPT-5, 2.18 for GPT-4, 2.64 for Claude Sonnet 4.6 and 3.48 for Claude Opus 4.8; English text sits at 1.17 to 1.88. Between the ends lies a factor of 2.0, and it hits input and output alike.

Measured across 24 EU languages (arXiv 2605.24718) the token count per word spreads by a factor of 2.5, from 1.2 in English to 3.1 in Greek. German averages 1.76 and ranges from 1.55 to 1.98 depending on the vendor, so 1.28 times through the choice of vendor alone.

The calculator applies the factor per model family to input and output and shows it as its own column in the model table. A comparison that treats tokens as a vendor-neutral unit favours the vendors with the coarser tokeniser.

Sources: Anthropic, pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing · TextKit, tokens per word across vendors, 10 June 2026 — https://textkit.tech/blog/tokens-per-word-tokenizer-comparison-2026 · Tokeniser surcharge across 24 EU languages, arXiv 2605.24718 — https://arxiv.org/html/2605.24718

As of 19 August 2026.

### Window size and context fidelity

The advertised context window says how much text fits, and fidelity says how much of it a model still finds.

Two vendors report the same long-context value in their model cards, MRCR v2 with eight needles, and that is where the label falls apart. Gemini 3.1 Pro drops from 84.9 percent at 128,000 tokens to 26.3 percent at 1M (model card, 19 February 2026). Anthropic writes about the 1M variant: “on the 8-needle 1M variant of MRCR v2 … Opus 4.6 scores 76%, whereas Sonnet 4.5 scores just 18.5%”. Same label, fourfold difference.

NoLiMa (ICML 2025) measures the drop where the question does not literally overlap with the passage: at 32,000 tokens eleven of thirteen models fall below half their short-context performance. RULER confirms this across 17 models and reports an effective length below the advertised one: Llama-3.1-70B is listed at 128k and carries 64k.

Chroma shows that the drop hits simple tasks as well: “model performance varies significantly as input length changes, even on simple tasks”. Needle search tests less than real analysis work, which sits above it.

In the model table fidelity has its own column for 128k and 1M. Four models in the set carry a value: three of them are reported by the vendors in their model cards, the fourth, Gemini 3.7 Flash, comes from a secondary source and carries a different metric, GDM-MRCR long context. All remaining models carry a dash, because no figure exists. The context length control in the requirements section works with these values.

Sources: Google DeepMind, Gemini 3.1 Pro model card, 19 February 2026 — https://deepmind.google/models/model-cards/gemini-3-1-pro/ · Anthropic, Claude Opus 4.6, 5 February 2026 — https://www.anthropic.com/news/claude-opus-4-6 · NoLiMa, ICML 2025, arXiv 2502.05167 — https://arxiv.org/abs/2502.05167 · NVIDIA RULER, effective context length — https://github.com/NVIDIA/RULER · Chroma, Context Rot, 14 July 2025 — https://www.trychroma.com/research/context-rot

As of 19 August 2026.

### The same question twice

A language model does not answer twice alike even at temperature 0, because the order of computation on the graphics card depends on server load.

On 10 September 2025 Thinking Machines Lab traced the cause to the varying batch size on the server: with the load, the order in which the kernel reduces changes as well. Measured, 1,000 identical requests at temperature 0 produced eighty different answers. With batch-invariant kernels all 1,000 answers were identical, and runtime rose from 26 to 55 seconds, and to 42 with an improved attention kernel.

The variation does not stay in the wording. Atil et al. (arXiv 2408.04667, five models, eight tasks, ten runs) report: “We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%.” Yuan et al. (arXiv 2506.09501) measure “up to 9% variation in accuracy and 9,000 tokens difference in response length” on a reasoning model and trace it to floating-point arithmetic at limited precision.

Reliability therefore comes from the architecture around the model. The documented base pattern: the model proposes, a fixed rule or a checking program decides, and only the decision counts. Structured outputs secure the form of the answer; about the content they say nothing, and the vendor notes that such outputs can still contain mistakes. Repeated sampling and voting schemes lower the spread and do not create determinism, because they draw several samples on purpose.

The BSI lists non-reproducibility as a risk category of its own, R7 (Generative AI Models, version 2.0, 17 January 2025): “The outputs of many generative AI models are not necessarily reproducible due to the use of random components.” Measure M20 requires, where the impact may be critical, a review of the output, cross-referencing with further sources and manual post-processing where needed. In the result, the top error-cost zone shows the documented patterns with their limits.

Sources: Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, 10 September 2025 — https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ · Atil et al., Non-Determinism of Deterministic LLM Settings, arXiv 2408.04667 — https://arxiv.org/abs/2408.04667 · Yuan et al., Numerical Sources of Nondeterminism in LLM Inference, arXiv 2506.09501 — https://arxiv.org/abs/2506.09501 · BSI, Generative AI Models, version 2.0, 17 January 2025, risk R7 and measure M20 — https://www.bsi.bund.de/SharedDocs/Downloads/DE/BSI/KI/Generative_KI-Modelle.pdf

As of 19 August 2026.

### Felt and measured effect

Between felt and measured productivity, METR measured a gap of around 39 percentage points.

METR (Becker, Rush, Barnes, Rein, arXiv 2507.09089, 12 July 2025) had experienced developers work on mature code, randomised with and without AI. Measured, they were 19 percent slower with AI, while they had estimated themselves faster. Between self-assessment and measurement lie around 39 percentage points.

The counter-figures belong with it. A rollout across tens of thousands of developers at Microsoft shows around 24 percent more merged pull requests (arXiv 2607.01418, 1 July 2026). And the Stanford AI Index 2026 records that METR later could not replicate its own finding, “primarily due to a growing reluctance among developers to work without AI, and that developers in late 2025 were likely sped up by AI relative to the original study period” (page 219).

One rule for this calculator follows: it is never calibrated with self-assessments. Every case study resting on a satisfaction survey can be optimistic by that order of magnitude, and vendor figures on time saved carry the same risk. The numbers here come from price lists, model cards and measurement series with a published setup.

Sources: Becker, Rush, Barnes, Rein (METR), arXiv 2507.09089, 12 July 2025 — https://arxiv.org/abs/2507.09089 · Microsoft rollout across tens of thousands of developers, arXiv 2607.01418, 1 July 2026 — https://arxiv.org/abs/2607.01418 · Stanford AI Index 2026, page 219 (note on the METR replication) — https://hai.stanford.edu/assets/files/ai_index_report_2026.pdf

As of 19 August 2026.

### Which benchmark per business case

An anchor only holds when it measures the task, and two widespread benchmarks measure something else for research and code review.

GPQA Diamond is built to be “Google-proof”, that is, built so searching does not help. A measure that neutralises search ability cannot assess it. On top of that it is saturated: on 15 August 2026 vals.ai counts “24 of the 133 models tested now score 90% or higher”, and the benchmark author, David Rein, says in an interview that it “stopped discriminating between good and great”. For research, BrowseComp and Deep Research Bench hold, the latter being the only one that reports cost per task.

SWE-bench Verified measures producing a patch when the fault location is known. A review has to find the fault first. CR-Bench (arXiv 2603.11078, 10 March 2026) sets out the difference and introduces its own measures, usefulness rate and signal-to-noise: “these agents often face a fundamental trade-off, they either prioritize precision and risk missing critical vulnerabilities, or prioritize recall at the cost of producing noisy and low-actionable feedback.”

For support tickets τ²-bench holds, because it tests the three properties of the case: a policy the system has to follow, tool use, and a simulated customer who follows up on incomplete answers. The spread on the public leaderboard is wide: Claude Opus 4.5 79 percent, GPT-5.2 73, Sonnet 4.5 63, GPT-4o 36. As a blocking criterion for source-bound output the HHEM hallucination leaderboard serves, read as a minimum threshold, because a low rate can also reflect low capability.

For text work there is a retrievable anchor with a price column, arena.ai creative writing (as of 12 August 2026, 1.18M votes). Confidence intervals of ±7 to ±18 points in the top group (±5 to ±28 across the whole field) make the leading ranks statistically indistinguishable. The anchor therefore works as a threshold: it answers whether a model belongs to the top group, and inside that group the price decides.

Sources: MindStudio, interview with GPQA author David Rein, 5 May 2026 — https://www.mindstudio.ai/blog/gpqa-benchmark-graduate-level-google-proof-qa-creator-limits · vals.ai, GPQA evaluation, 15 August 2026 — https://www.vals.ai/benchmarks/gpqa · Deep Research Bench, as of 10 June 2026 — https://evals.futuresearch.ai/ · CR-Bench, arXiv 2603.11078, 10 March 2026 — https://arxiv.org/abs/2603.11078 · τ²-bench, sierra-research/tau2-bench — https://github.com/sierra-research/tau2-bench · CodeSOTA, τ²-bench leaderboard 2026 (secondary source) — https://www.codesota.com/benchmark/tau2-bench · Vectara HHEM-2.3, values reported via CodingFleet 2026 (secondary source) — https://codingfleet.com/blog/ai-model-hallucination-rates-2026/ · arena.ai, creative writing, as of 12 August 2026 — https://arena.ai/leaderboard/text/creative-writing

As of 19 August 2026.

## All models (list prices per 1M tokens, USD)

| Model | Vendor | Status | Input | Output | Cache read | Batch | Context (tokens) | Note |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Claude Fable 5 | Anthropic | current | 10 | 50 | 1 | yes | 1,000k | — |
| Claude Mythos 5 | Anthropic | current | 10 | 50 | 1 | yes | 1,000k | limited availability |
| Claude Opus 5 | Anthropic | current | 5 | 25 | 0.5 | yes | 1,000k | — |
| Claude Opus 4.8 | Anthropic | superseded | 5 | 25 | 0.5 | yes | 1,000k | — |
| Claude Opus 4.7 | Anthropic | superseded | 5 | 25 | 0.5 | yes | 1,000k | — |
| Claude Opus 4.6 | Anthropic | superseded | 5 | 25 | 0.5 | yes | 1,000k | — |
| Claude Opus 4.5 | Anthropic | superseded | 5 | 25 | 0.5 | yes | 200k | — |
| Claude Opus 4.1 | Anthropic | retired | 15 | 75 | 1.5 | yes | 200k | — |
| Claude Opus 4 | Anthropic | retired | 15 | 75 | 1.5 | yes | — | — |
| Claude Sonnet 5 | Anthropic | current | 2 | 10 | 0.2 | yes | 1,000k | — |
| Claude Sonnet 4.6 | Anthropic | superseded | 3 | 15 | 0.3 | yes | 1,000k | — |
| Claude Sonnet 4.5 | Anthropic | superseded | 3 | 15 | 0.3 | yes | 200k | — |
| Claude Sonnet 4 | Anthropic | retired | 3 | 15 | 0.3 | yes | — | — |
| Claude Haiku 4.5 | Anthropic | current | 1 | 5 | 0.1 | yes | 200k | — |
| Claude Haiku 3.5 | Anthropic | retired | 0.8 | 4 | 0.08 | yes | — | — |
| gpt-5.6-sol | OpenAI | current | 5 | 30 | 0.5 | yes | 1,050k | — |
| gpt-5.6-terra | OpenAI | current | 2 | 12 | 0.2 | yes | 1,050k | — |
| gpt-5.6-luna | OpenAI | current | 0.2 | 1.2 | 0.02 | yes | 1,050k | — |
| gpt-5.5 | OpenAI | current | 5 | 30 | 0.5 | yes | 1,050k | — |
| gpt-5.5-pro | OpenAI | current | 30 | 180 | — | yes | 1,050k | — |
| gpt-5.4 | OpenAI | current | 2.5 | 15 | 0.25 | yes | 1,050k | — |
| gpt-5.4-mini | OpenAI | current | 0.75 | 4.5 | 0.075 | yes | 400k | — |
| gpt-5.4-nano | OpenAI | current | 0.2 | 1.25 | 0.02 | yes | 400k | — |
| gpt-5.4-pro | OpenAI | current | 30 | 180 | — | yes | 1,050k | — |
| gpt-5.2 | OpenAI | current | 1.75 | 14 | 0.175 | yes | 400k | — |
| gpt-5.2-pro | OpenAI | current | 21 | 168 | — | yes | 400k | — |
| gpt-5.1 | OpenAI | current | 1.25 | 10 | 0.125 | yes | 400k | — |
| gpt-5 | OpenAI | superseded | 1.25 | 10 | 0.125 | yes | 400k | shutdown announced for 11 Dec 2026 |
| gpt-5-mini | OpenAI | superseded | 0.25 | 2 | 0.025 | yes | 400k | shutdown announced for 11 Dec 2026 |
| gpt-5-nano | OpenAI | superseded | 0.05 | 0.4 | 0.005 | yes | 400k | shutdown announced for 11 Dec 2026 |
| gpt-5-pro | OpenAI | superseded | 15 | 120 | — | yes | 400k | shutdown announced for 11 Dec 2026 |
| gpt-4.1 | OpenAI | superseded | 2 | 8 | 0.5 | yes | 1,048k | — |
| gpt-4.1-mini | OpenAI | superseded | 0.4 | 1.6 | 0.1 | yes | 1,048k | — |
| gpt-4.1-nano | OpenAI | superseded | 0.1 | 0.4 | 0.025 | yes | 1,048k | shutdown announced for 23 Oct 2026 |
| gpt-4o | OpenAI | superseded | 2.5 | 10 | 1.25 | yes | 128k | — |
| gpt-4o-mini | OpenAI | superseded | 0.15 | 0.6 | 0.075 | yes | 128k | — |
| o1 | OpenAI | superseded | 15 | 60 | 7.5 | yes | 200k | shutdown announced for 23 Oct 2026 |
| o1-pro | OpenAI | superseded | 150 | 600 | — | yes | 200k | shutdown announced for 23 Oct 2026 |
| o3 | OpenAI | superseded | 2 | 8 | 0.5 | yes | 200k | shutdown announced for 11 Dec 2026 |
| o3-pro | OpenAI | superseded | 20 | 80 | — | yes | 200k | shutdown announced for 11 Dec 2026 |
| o3-mini | OpenAI | superseded | 1.1 | 4.4 | 0.55 | yes | 200k | shutdown announced for 23 Oct 2026 |
| o4-mini | OpenAI | superseded | 1.1 | 4.4 | 0.275 | yes | 200k | shutdown announced for 23 Oct 2026 |
| Gemini 3.7 Flash | Google | current | 0.75 | 3.75 | 0.075 | yes | 1,049k | Promotional price through 31 Dec 2026, then 1.50 / 7.50 USD |
| Gemini 3.6 Flash | Google | current | 0.75 | 3.75 | 0.075 | yes | 1,049k | Promotional price through 31 Dec 2026, then 1.50 / 7.50 USD |
| Gemini 3.5 Flash | Google | current | 1.5 | 9 | 0.15 | yes | 1,049k | — |
| Gemini 3.5 Flash-Lite | Google | current | 0.3 | 2.5 | 0.03 | yes | 1,049k | — |
| Gemini 3.1 Pro Preview | Google | current | 2 | 12 | 0.2 | yes | 1,049k | — |
| Gemini 3.1 Flash-Lite | Google | current | 0.25 | 1.5 | 0.025 | yes | 1,049k | shutdown announced for 7 May 2027 |
| Gemini 2.5 Pro | Google | superseded | 1.25 | 10 | 0.125 | yes | 1,049k | — |
| Gemini 2.5 Flash | Google | superseded | 0.3 | 2.5 | 0.03 | yes | 1,049k | — |
| Gemini 2.5 Flash-Lite | Google | superseded | 0.1 | 0.4 | 0.01 | yes | 1,049k | — |
| Mistral Medium 3.5 | Mistral | current | 1.5 | 7.5 | 0.15 | yes | 256k | — |
| Mistral Large 3 | Mistral | current | 0.5 | 1.5 | 0.05 | yes | 256k | — |
| Mistral Small 4 | Mistral | current | 0.15 | 0.6 | 0.015 | yes | 256k | — |
| Magistral Medium | Mistral | retired | 2 | 5 | 0.2 | yes | 128k | — |
| Magistral Small | Mistral | retired | 0.5 | 1.5 | 0.05 | yes | 128k | — |
| Ministral 3 3B | Mistral | current | 0.1 | 0.1 | 0.01 | yes | 128k | — |
| Ministral 3 8B | Mistral | current | 0.15 | 0.15 | 0.015 | yes | 128k | — |
| Ministral 3 14B | Mistral | current | 0.2 | 0.2 | 0.02 | yes | 128k | — |
| Devstral 2 | Mistral | retired | 0.4 | 2 | 0.04 | yes | 256k | — |
| Devstral Small 2 | Mistral | retired | 0.1 | 0.3 | 0.01 | yes | 256k | — |
| Codestral | Mistral | current | 0.3 | 0.9 | 0.03 | yes | 256k | — |
| Mistral NeMo | Mistral | retired | 0.15 | 0.15 | 0.015 | yes | 128k | — |
| Mixtral 8x7B | Mistral | retired | 0.7 | 0.7 | 0.07 | yes | 32k | — |
| Mixtral 8x22B | Mistral | retired | 2 | 6 | 0.2 | yes | 64k | — |
| DeepSeek V4 Pro | Open weights, at the model provider | current | 1.32 | 3.96 | 0.044 | no | 1,000k | — |
| DeepSeek V4 Flash | Open weights, at the model provider | current | 0.44 | 1.32 | 0.014 | no | 1,000k | — |
| DeepSeek V4 Pro (Together) | Open weights, at a hoster | current | 1.74 | 3.48 | 0.2 | no | 1,000k | hosted at Together AI |
| Qwen3.7-Max (Together) | Open weights, at a hoster | current | 1.25 | 3.75 | 0.13 | no | 262k | hosted at Together AI |
| Qwen3.7-Plus (Together) | Open weights, at a hoster | current | 0.32 | 1.28 | — | no | 262k | hosted at Together AI |
| Qwen3.6-Plus (Together) | Open weights, at a hoster | current | 0.5 | 3 | — | no | 262k | hosted at Together AI |
| Qwen3.5-397B-A17B (Together) | Open weights, at a hoster | current | 0.6 | 3.6 | 0.35 | no | 262k | hosted at Together AI |
| Qwen3.5 9B (Together) | Open weights, at a hoster | current | 0.17 | 0.25 | — | no | 131k | hosted at Together AI |
| Qwen2.5 7B Instruct Turbo (Together) | Open weights, at a hoster | superseded | 0.3 | 0.3 | — | no | 33k | hosted at Together AI |
| Llama 3.3 70B (Together) | Open weights, at a hoster | superseded | 1.04 | 1.04 | — | no | 131k | hosted at Together AI |

Conversion uses 1.1576 USD/EUR (ECB reference rate of 18 August 2026); that rate is an assumption, not a daily quote. Prices are vendor list prices, verified 18 August 2026 against the pricing pages below, without discounts.

Price sources per vendor:

- Anthropic: pricing documentation — https://platform.claude.com/docs/en/about-claude/pricing (prices, cache rates, context windows, cross-check)
- OpenAI: pricing page — https://developers.openai.com/api/docs/pricing (prices, tiers from 272,000 tokens)
- Google: Gemini pricing page — https://ai.google.dev/gemini-api/docs/pricing (prices, tiers from 200,000 tokens, storage fee)
- Mistral: API pricing — https://mistral.ai/pricing/api (prices, cache discount)
- DeepSeek: pricing page — https://api-docs.deepseek.com/quick_start/pricing
- Together AI: pricing page — https://www.together.ai/pricing (open weights at a hoster)
- OpenRouter: pricing page — https://openrouter.ai/pricing (router fee)

## Sources

- τ²-bench, customer service with policy and tools — https://www.codesota.com/benchmark/tau2-bench (19 August 2026)
- HHEM, hallucination when summarising supplied text — https://github.com/vectara/hallucination-leaderboard (11 May 2026)
- arena.ai creative writing, writing quality by audience vote — https://arena.ai/leaderboard/text/creative-writing (12 August 2026)
- Vals CorpFin v2, reading long credit agreements — https://www.vals.ai/benchmarks/corp_fin_v2 (12 August 2026)
- BrowseComp, research on the live web — https://benchlm.ai/benchmarks/browsecomp (18 August 2026)
- MRCR v2, context fidelity across the full window — https://deepmind.google/models/model-cards/gemini-3-1-pro/ (19 August 2026)
- Volume factors: Reference level of the volume factors (mengen.json, medium equals 1). Anthropic, Optimizing for cost and intelligence: default level high, medium halves cost per task for Opus 5 (as of 19 August 2026). Artificial Analysis states tokens per index task only in prose, not per model and effort level — https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1 (19 August 2026)
- Context fidelity: DeepMind Gemini 3.1 Pro model card 19 February 2026; Anthropic Claude Opus 4.6 5 February 2026 — https://deepmind.google/models/model-cards/gemini-3-1-pro/ (19 August 2026)
- Tokenizer: TextKit, Tokens per word (10 June 2026), Anthropic Pricing (new tokenizer from Claude 4.7, about 30 percent more tokens), arXiv 2605.24718 (German 1.55 to 1.98 per vendor). Reference: GPT-5 tokenizer English = 1.0, rounded down conservatively. — https://textkit.tech/blog/tokens-per-word-tokenizer-comparison-2026 (19 August 2026)
- Epoch AI, Epoch Capabilities Index (data CC-BY) — https://epoch.ai/eci (18 August 2026)
