Back to Blog

The Model Benchmark Numbers Vendors Publish Are Measured on Their Own Harness

September 24, 2026

There is a claim circulating right now, and you have probably seen a version of it: "DeepSeek beats Claude on terminal work now."

The number behind it is real. The conclusion is wrong — and not because the model changed, but because nobody asked which test produced the number.

This is not a story about one bad screenshot. It is the single most common way intelligent teams pick the wrong model, and it costs real money.

Start with the claim, then change one thing

"DeepSeek beats Opus on terminal work" comes from Terminal-Bench, a test that gives an agent a real command line and a real task. On that benchmark DeepSeek V4.1 Flash scores 90.6 and Claude Opus 5 scores 89.1. DeepSeek wins.

Terminal-Bench has versions. Here is the same comparison across three of them:

VersionDeepSeek V4.1 FlashClaude Opus 5Who wins
2.190.689.1DeepSeek
3.030.043.3Opus 5
4.031.251.8Opus 5

Same two models. The verdict reverses, and on 4.0 the gap is 20 points.

The circulating number was version 2.1 — the oldest of the three.

Why this keeps happening

A benchmark score is not a measurement of a model. It is a measurement of a model plus a harness — the scaffolding around it: which tools it may call, how its output is parsed, how many attempts it gets, what counts as success.

Think of it like timing two drivers. One is timed on a closed track, in a car the team set up, on tyres they chose. The other is timed in city traffic. Comparing the two lap times tells you about the traffic, not the drivers.

Nobody is lying. Every vendor publishes their number honestly. They just publish the number from their own track, because it is theirs.

Four cases where this is documented and public:

Terminal-Bench, above. The claim reverses across versions.

DeepSWE. DeepSeek's own published table shows its score moving 8.7 points across eight frameworks — 74.2 on one, 65.5 on another. If you run Claude Code 2.1.251, which is what most people have installed, you get 69.8, not the 74.2 headline. Anthropic shows the same effect from the other side: GPT-5.5 scores 78.2 on one harness and 83.4 on its own.

ARC-AGI-3. OpenAI's headline for GPT-6 Astra is 99.9%. That number requires OpenAI's own provider adapter with reasoning state carried between steps. ARC Prize — the organisation that runs the benchmark — tested the same model on a stateless harness and got somewhere between 17% and 63%. Both numbers are real. They are measuring different things.

OSWorld. Astra scores 72.6% on OSWorld 2.0. Opus 4.8 scores 83.40 on OSWorld-Verified. Those are different tests with similar names, and quoting them side by side as a comparison is simply a mistake.

The second trap: scores that reward being free

There is a class of "score" that looks like a benchmark and is not one.

When Union Alpha appeared — an anonymous model, no weights, no published evaluations — the aggregate listing gave it 40 out of 100, rank 234. That reads like a capability measurement. Here is the actual breakdown:

  • Pricing: 100/100 (25% of the total)
  • Capabilities: 67/100 (30%)
  • Recency: 100/100 (15%)
  • Context window: 86/100 (15%)
  • Output length: 83/100 (15%)

Being free scores full marks. Being new scores full marks. Those two categories are 40% of the composite between them. A brand-new free model cannot score below 40 no matter how it performs, and a mature paid model is penalised for charging money.

If you see a single composite score for an AI model, find the breakdown before you believe it.

The third trap: comparisons that were never run

You want to know how DeepSeek compares to Claude Opus 4.8. So did we.

No such comparison exists. DeepSeek published against Opus 5 and GPT-5.6 Sol. Anthropic's Opus 4.8 card benchmarked against Opus 4.7, GPT-5.5 and Gemini 3.1 Pro. Neither vendor tested the other's model.

Any DeepSeek-versus-4.8 table you find circulating online was assembled by pasting two different scorecards next to each other. The rows were never in the same room.

What you can actually verify

Here is the part that does not depend on anybody's harness: the bill.

For a workload of 10 million input tokens and 1 million output tokens, using official published tariffs and nothing else:

UncachedWith batchBatch + 90% cache
Opus (either 4.8 or 5)$75$37.50$17.25
DeepSeek V4.1 Flash, off-peak$2.10—$0.78

Per 1,000 drafted cold emails, that is roughly $37.50 against $0.98.

Two things fall straight out of that table, and neither requires a benchmark:

  1. Caching matters more than model choice. On Opus, batch plus caching cuts the bill by 77%. Teams argue about which model to use while leaving the bigger saving on the floor.
  2. Timing is free money. DeepSeek charges double during peak hours — 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. If you are in Nigeria, that is 02:00 to 05:00 and 07:00 to 11:00. Run your batch outside it and the same work costs half.

One more, because it is a genuine trap: if you are paying for Opus 4.8, you are paying Opus 5's price. Identical — $5 and $25 per million. 4.8 is currently flagged Legacy by Anthropic. The upgrade costs nothing.

How to read a benchmark from now on

Four questions. They take a minute and they have never failed me.

  1. Which version? If a paper does not say, assume the oldest one.
  2. Whose harness? Vendor's own, or an independent one? If it is the vendor's, treat it as a best case.
  3. What is being compared to what? If the two models never appeared in the same table, the comparison was assembled by someone else.
  4. Can I check it myself for under $50? If yes, do that instead. One afternoon of your own workload beats every published number, because your workload is the only one you actually care about.

The honest caveat

Everything above is a reading of published material: Anthropic's system cards, DeepSeek's own benchmark tables, OpenAI's launch posts, official pricing pages. We did not run these models against each other. We read every scorecard we could find and looked for where they disagreed with each other.

That is a weaker claim than "we benchmarked five models", and it is the one we can stand behind. If you want the stronger claim, run the test on your own data — and we would genuinely like to see your numbers.

If you are choosing a model for a real workload and want a second read on the numbers, get in touch.