GitHub
Alexandre Gensse
Data analysis · aggregated from independent measurement sources

Cost, effort and model : the Claude matrix

What each recent Claude model actually costs, at every effort level — reconstructed from public measurements and reduced to a relative cost by chaining same-task comparisons.

Baseline
Opus 4.8 · medium = 1.00
Measurement sources
independent
Generated
08 Sep 2026

Quality vs. Cost

One curve per model, one point per effort (low → max), both axes relative to Opus 4.8 @medium = 1.0; quality is dilated near parity. The faint oval (hover to reveal) is the robust asymmetric uncertainty — width = cost, height = quality. Haiku 4.5 lives only in the matrix and Pareto.

Pareto Frontier

The non-dominated (model, effort) couples : nothing offers more quality at equal-or-lower cost. Solid = frontier, faded = dominated. The faint grey curve is the price curve — what a given quality typically costs, fitted on every measured couple but with each one's influence graded by its distance to the frontier, and the reference for the value index below.

Value index — distance to the price curve

We fit a price curvelog₁₀(cost) = g(quality), what a given quality typically costs across all 35 measured couples — not just the Pareto frontier, but each one weighted by 1 − d/dmax, where d is how many decades dearer it is than the cheapest couple reaching at least its quality. A couple on the frontier counts fully, the single worst one not at all, everything in between on a straight line. Fitting all couples flat would let strictly-dominated models steer the curve; fitting the frontier alone would discard the measurements that populate the middle (R² ≈ , drawn faint grey above). Each couple's distance in log-cost to that curve says how much cheaper or dearer it is than the frontier price for its quality. Measuring in cost (not quality) keeps the weighting linear where the curve flattens. That distance becomes the value index by exponentiating it relative to the anchor : 100 = Opus 4.8 @medium, and because it is a ratio rather than a stretched scale, 350 reads literally as 3.5× the anchor's value for money and 45 as 0.45×. It is unbounded above — the price of an interpretable multiple. Confidence intervals are baked in : both the curve and every index weight the couple's centre (½) and its four CI extremities (⅛ each), averaged in log space, so a wide interval carries the couple toward what the envelope charges across its whole box. The curve is monotone by construction — the price of quality cannot fall as quality rises — which also makes it safe to let it extrapolate freely beyond the observed range: it can flatten or steepen, never fold back. Its shape carries a growing curvature term, because cost climbs gently across most of the quality range then steepens sharply near the ceiling — a plain parabola cannot follow that, and an unconstrained one bends back on itself and flips the sign of every index around it. This grading is also what gives the curve its bend : fitted flat over every couple it comes out almost straight, because the dilated axes already linearise most of the cloud ; weighted toward the frontier it follows the sharp rise near the quality ceiling, where efficient couples actually sit. Only the Pareto-frontier couples are shown.

Model · effortcostqualityvalue index
anchor = 100

Green = above 100, cheaper than the frontier price for its quality (better value than the anchor), red = below 100, dearer; intensity = distance to the anchor in decades, capped at one. A couple on the Pareto frontier can still be red : being non-dominated does not mean being good value for what it delivers.

Best value by task tier — what to actually run

Consolidating the frontier into a decision. Four target complexity levels q* are derived from the data, not fixed: they spread evenly across the frontier's own quality span in the dilated metric, so the bottom tier tracks the weakest couple available and the top tier sits exactly on the best one — as models improve, the windows move with them instead of aiming at a level fixed at some past release. The width follows the spacing on one rule : adjacent windows cross at half weight midway between their centres. For each tier we then pick the Pareto-frontier couple maximising that Gaussian times its value index, with cost raised to a per-tier exponent γ1.20 · 1.05 · 0.95 · 0.80 from throwaway work up to research-grade. Cost does not weigh the same at every complexity : on grunt work you want the cheapest thing that clears the bar, on a hard task a few extra points of capability are worth paying for. γ > 1 punishes cost more than proportionally, γ < 1 less. Only the selection is tilted : the index shown on a card stays the neutral γ = 1 one, so the four cards remain comparable with each other and with the anchor. Cost enters through the index, i.e. in log-cost like everywhere else in this report : a raw quality ÷ cost ratio would be driven almost entirely by cost, which spans a factor 22 across the cloud against 2 for quality. Sliders below override both. The big figure is the pick's value index, the same one used throughout the report : 100 = the anchor, Opus 4.8 @medium, and the value reads as a multiple of it — 350 means 3.5× the anchor's value for money. The anchor is read from every couple, dominated ones included, not just the frontier : a new model can push it off the frontier — Opus 5 does — but never out of the full set, so the reference survives a release. The crown is chosen on a different criterion, local prominence along the frontier, and shows the same index.

Tune the target complexity windows (q*, σ)
Normalized Cost Matrix

Relative cost per model × effort (baseline Opus 4.8 @medium = 1.00), sorted by decreasing relative quality. These values are computed (not hand-entered) from the measured same-task ratios : within each benchmark we take the couple ratios, normalise to Opus 4.8 @medium, then each cell = the source-weighted median of the per-benchmark estimates (weight = number of sources; bridge ×0.5). Intensity = cost; the second figure is the robust CI (per-side Huber ±1.5·MAD, asymmetric ; wide = benchmarks disagree).

ModellowmediumhighxhighmaxRel. quality
@max

"Rel. quality @max" = relative quality at top effort (Opus 4.8 @medium = 1.0), from the same consolidated ratios as the landscape. Haiku 4.5 = a single merged cell (single operating point : no discrete effort parameter, see the measurement sources). n/a : effort not exposed (Sonnet 4.6 : no xhigh).

What moves everything — four factors behind the numbers

The tier picks above are an order of magnitude. Four factors, each measured elsewhere, modulate them strongly — keep them in mind when reading the tiers and the matrix.

Task type

Effort and model choice don't weigh the same depending on the nature of the work — review, reasoning and pure coding reward effort very differently.

Complexity

The harder the task, the more effort costs — up to the point where extra thinking starts to hurt (the inverted-U).

Generation

Each release buys the same quality for less — Fable 5.1 beats Fable 5's max quality at its own low effort for a fifth of the cost, and Opus 5 reaches Fable-5-class quality for roughly half the cost per task; Sonnet 5 is the verbose exception, cheap per token but talkative.

Cache & harness

Two invisible factors — cache-read rate and the surrounding agent harness — that can dominate both cost and score.

Data Sources

Every ratio comes from a source that measured ≥2 (model, effort) couples on the same task1206 same-task measurement points in total. Each row below is one such source, with the actual configuration we verified (harness · effort), its type, and the couples it links.

SourceActual config (harness · effort)TypeLinked couples (same task)
How the numbers are built — the method

Every value in this report is computed from the raw measurements by one procedure — nothing is hand-entered.

In short : we never compare two raw dollar figures across sources — task sizes are incomparable — only ratios measured on the same task, which cancel that variance. Each (model, effort) couple is atomic (no model×effort separability assumed), normalised against the anchor Opus 4.8 @medium, then consolidated into a source-weighted median across benchmarks. The uncertainty band is the robust (Huber) spread between those benchmarks : tight = agreement, wide = they disagree.

The full procedure — six steps

  1. Same-task ratios only. We never compare two raw dollar figures from different sources — task sizes are incomparable. We keep only sources that measure ≥2 (model, effort) couples on the same task; their ratio cancels the task-size variance.
  2. The (model, effort) couple is atomic. Each node is one indivisible couple. We never assume "effort" has the same effect across models — no model×effort decomposition, because separability would be a false assumption.
  3. Per-benchmark normalisation → weighted median. Inside each benchmark every couple is divided by the anchor (Opus 4.8 @medium) when present, else bridged through its shared couples (down-weighted). Each benchmark yields one normalised estimate per couple; the couple's value is the source-weighted median across benchmarks. Single-couple benchmarks are dropped — they would only echo the anchor.
  4. "Default" is verified, not assumed. Each source's actual configuration is read from its harness rather than guessed; a ratio is used only when a source applies the same configuration to all its models.
  5. Unlinked couples are flagged, never invented. A couple with no measured edge is left n/a or given an explicitly wide band — one model's effort shape is never used as a template for another.
  6. The band = agreement. Uncertainty is a per-side Huber spread (deviations clipped to ±1.5·MAD): robust to an outlier benchmark yet still widened by it, and asymmetric. Tight agreement → narrow band; disagreement → wide band.

Alexandre Gensse · github.com/alexandregensse-blip

All figures are indicative and for informational purposes only — derived from public third-party measurements, not an official benchmark, and not affiliated with or endorsed by Anthropic. Verify before relying on any value.