A chip company claims 3.7x the throughput of its own last product. A startup claims 193x faster than frontier models. A lab claims its new accelerator beats Nvidia on three flagship models. At least three of those four sentences will be true as written and misleading anyway, and the reader has no way to tell which from the headline. This article is the field guide we use internally when we write about benchmarks. It is not cynical: verified benchmarking exists, it works, and it produced most of the good numbers on this site. The point is that every benchmark number is the output of five human decisions, and the decisions, not the number, are where truth hides.
Each section below uses a case we have already covered, with the same caveats we attached the first time. The examples are recent because recent examples are checkable; the method outlives all of them.
Decision one: which model gets measured
The generational claim in our Vera Rubin analysis illustrates the first fork. NVIDIA's MLPerf v6.1 submission reported up to 3.7x throughput over GB300 on Qwen3-VL (a vision-language model) and up to 2.5x on DeepSeek-R1 (a reasoning model). Both real. Neither is "the" number. Working through the raw submissions, the like-for-like DeepSeek server-scenario ratio is 1,175,890 tokens/s versus 603,023, a 1.95x generational step, and Nebius's independent per-GPU framing lands at 1.96x. The honest generational gain is about 2x on frontier reasoning and up to 3.7x on the one model where the architecture happened to align.
The rule: a multiple is never model-independent, because different workloads stress different bottlenecks. Reasoning models punish memory bandwidth; vision-language models punish preprocessing pipelines; agentic workloads punish inter-request latency. When a release says "up to Nx" without pinning the model, the model choice is the claim. When a vendor picks one model from a suite of five, ask which two they did not submit. MLPerf at least forces the model onto the label; vendor blog posts routinely do not.
Decision two: which scenario, and which point on the curve
Every serious benchmark measures a curve, not a point. MLPerf runs offline (maximum throughput, no latency guarantee), server (throughput under a latency cap), and interactive scenarios, and the ratios move between them: Vera Rubin's DeepSeek lead was 1.95x in server but 1.72x in offline. The peak and the floor describe different products, and buyers buy specific points: a chat app lives at interactive latency, a batch pipeline lives at offline throughput.
The sharper version of this trap is the normalization choice. OpenAI's Jalapeño benchmarks were normalized to each accelerator's published package TDP (700W versus the GB300's 1,400W), which is fair and disclosed, and it is exactly why tokens-per-kilowatt beat tokens-per-dollar as the metric of choice: it flatters low-power ASICs by construction on workloads where absolute throughput still decides who wins a datacenter contract. Note the same release's most flattering number, 53.7x throughput "at the previous best time between tokens": a point on the latency curve chosen because the ratio peaks there. That sentence was true, disclosed, and nearly meaningless without the curve underneath it. When you only get one point, ask what denominator it rides on and where the ratio is worst.
Decision three: who ran it, and who paid for the wrapper
Verification has layers, and press releases blur them deliberately.
| Verification layer | Example from our coverage | What it actually guarantees |
|---|---|---|
| Third-party suite, public submission | MLPerf v6.1 (30 orgs, 486 results, MLCommons-audited) | Reproducible methodology; not that your workload is one of theirs |
| Independent observer, vendor-run | SemiAnalysis running InferenceX on Jalapeño in person | Numbers were not fabricated; suite selection still vendor-adjacent |
| Vendor-submitted, self-published | NVIDIA's AgentX claim of 30x over GB300 | Nothing; preview, not in MLCommons results |
| Vendor benchmark, self-authored, self-graded | TypeSafe's four Jev workflows, graded by agreement with GPT-6 Astra and Claude Fable 5.1 averages | Mimicry of two models, on tasks the vendor wrote, with the competitor baselines run through the vendor's own wrapper |
The Jev case deserves a second because it stacks every layer of suspicion at once, and TypeSafe deserves credit for admitting all of it in their own nuance sections: self-authored workflows, reference answers averaged from frontier models, competitor latency measured through their own wrapper. The one independent datapoint (Every: roughly 25x faster, 580x cheaper on a single extraction task) corroborates direction, not magnitude. A benchmark graded by agreement with existing models also has a ceiling problem: a method that outperforms the graders is scored as wrong.
The rule has no exceptions we have found: vendor numbers are marketing until a third party reproduces them, and third-party reproduction has a quality hierarchy of its own.
Decision four: what is in the denominator, and what is not in the price
Benchmark numbers omit cost with near-religious consistency, and the omission quietly changes the claim. In the Vera Rubin piece we computed that a 1.95x rack-level gain implies cost-per-token falls to roughly 51 percent of GB300 if the racks cost the same, and flagged that NVIDIA publishes no Vera Rubin price, so the interesting number literally does not exist yet. Meanwhile in aggregate, AMD's 512-GPU MI355X cluster posted 2.9 million tokens/s on DeepSeek-R1 offline, more than double a single Vera Rubin rack: the efficiency crown and the throughput crown belong to different players because they answer different questions, and per-GPU framing is how a small strong chip and a big cheap cluster each claim victory.
The same denominator problem runs through pricing claims. Gemini 3.8 Live's $0.005-per-minute versus OpenAI's $0.05 looked like 10x, but the rates price different scopes (Google's covers input minutes; OpenAI's is a voice layer billed on full duration, with the reasoning backend billed separately), and the honest balanced-hour figure was $0.69 versus $3.00, a 4x on the voice layer before either backend. And Meta's Arke savings claims have the purest form of the problem: MTIA is internal, so there is no market price, so "saves money" is measured against a counterfactual (what Meta would have paid Nvidia) that will never be public. When the denominator is a counterfactual, no benchmark can close the gap.
Decision five: which metric, and who still uses it
Metrics are chosen for flattering shape, then the industry converges on the flattering shape and it stops being a tell. That is happening now: the cost chain we mapped last week runs on tokens-per-kilowatt because inference is memory-bound and power is the binding constraint, but the moment every lab reports tokens-per-kilowatt, the game moves to which model (decision one), which scenario (decision two), and which verifier (decision three), all over again. One structural note for calibration: MLCommons reports the best per-accelerator DeepSeek-R1 server result in v6.1 was 5.7x the best result of the same test a year earlier. When an individual claim is 1.5x to 3.7x against a field improving ~2-6x a year, the claims are not inflated; they are normal. Skepticism calibrated against a moving baseline matters, because a reader who assumes all numbers are inflated will miss the genuinely large ones as readily as the fake ones.
Our Read
Five checks, in priority order, distilled from every case above. One, pin the model: any multiple without a named model and named scenario is a mood, not a measurement. Two, hunt the worst point on the curve, not the peak; if the ratio at the floor still wins your workload, the claim survives. Three, ask who ran it and whether the baseline traveled through the winner's wrapper; self-authored, self-graded, competitor-wrapped is the weakest tier even when honestly disclosed, and it deserves credit for disclosure and zero credit for magnitude. Four, insert the price: an efficiency claim whose cost denominator is unpublished or counterfactual is a physics claim, not an economics claim, and they are different claims. Five, normalize against the field's drift rate (for frontier inference, roughly multiples per year in 2026), so "up to 2x" reads as routine and "up to 30x on an unverified benchmark" reads as a preview that belongs in a footnote.
Our own bias, stated plainly so you can discount it: we treat the MLPerf-style public submission as the only tier that supports a headline number, everything else supports a story about intent. On this site that means vendor benchmarks appear labeled as vendor numbers in every article, and the strongest claims we publish are the ones where two independent submitters agreed, like the 1.95x and 1.96x Vera Rubin numbers converging from different framings. Convergence between parties with opposite incentives is the closest thing benchmarking has to proof.
Outlook
Expect three pressure points on the method itself over the next year. Agentic workloads (multi-step, tool-using, session-long) do not fit the throughput-at-latency-cap template, and the fight over their standardization, which NVIDIA is pushing toward with preview agentic numbers while MLCommons prepares endpoint-style benchmarks, will produce the decade's most abused metrics. Expect grading-by-agreement (Jev-style benchmarking against frontier-model averages) to spread with distillation, bringing its mimicry ceiling along with it. And expect cost-per-token to keep displacing raw throughput as the headline, which is progress, provided the price denominator is a real invoice rather than Meta-style counterfactual. The five decisions will survive every metric change; that is what makes them worth memorizing instead of the numbers.
References
- StorageReview, MLPerf Inference v6.1 results coverage, storagereview.com, participation record and per-accelerator field improvement
- NVIDIA blog, Vera Rubin NVL72 MLPerf results, blogs.nvidia.com, vendor-reported ratios and AgentX preview claim
- Nebius blog, MLPerf v6.1 writeup, nebius.com, independent per-GPU framing and scaling measurements
- OpenAI, Jalapeño inference chip announcement, openai.com, TDP-normalized benchmark claims
- SemiAnalysis, InferenceX verification of Jalapeño, semianalysis.com, independent benchmark run
- TypeSafe, Jev announcement, typesafe.ai, self-authored workflow evals and disclosed caveats
- Google developer blog, Gemini audio pricing, blog.google, per-minute voice rates
- Meta AI blog, MTIA scale chips, ai.meta.com, vendor-reported efficiency claims