Part 10 — Integrated Case Studies
What the experiment does not prove
The discipline that separates a credible evaluator from a hype merchant: stating clearly what a result does not license. The case study's 95.5% is a provisional routing decision on one 88-task dataset — not a universal model ranking — and knowing the difference is the point.
The mark of a serious evaluator is not the result they claim but the results they refuse to claim. Having interpreted the case study's numbers, this chapter draws the boundary around them: what the experiment does not prove. The 95.5% is a real, useful, defensible number — and it licenses far less than it's tempting to say. Stating that boundary explicitly is not false modesty; it's what keeps an evaluation honest and its conclusions from being weaponised into claims the data can't support.
What you will understand by the end
- Why a specific result licenses only a specific, bounded claim.
- The concrete things the case study does not prove.
- Why "provisional routing decision" is the honest label, not "model ranking."
- How stating the limits is itself a core evaluation skill.
A result licenses a bounded claim
Every number comes from a specific setup — a specific system, a specific dataset, a specific scoring rule, a specific method — and it licenses conclusions only within that setup. The case study measured six configurations on one 88-task dataset with one attempt each, scored by exact match. That's what the numbers are about, and it's the outer limit of what they can support. Push a conclusion past that boundary and you're no longer reporting the experiment; you're speculating with its authority.
A measurement licenses a claim only within its setup. The case study's result is a statement about these six configurations, on this dataset, under this scoring, with one attempt each — and nothing beyond it. Naming that boundary is not hedging; it's the difference between a defensible conclusion and an overreach dressed in a real number.
What it does not prove
Concretely, the 95.5% result does not establish:
- That one model is "more accurate" than another in general. It compared systems on one task; a hosted model leading here says nothing about its accuracy on other tasks, domains, or datasets.
- That the ranking is stable. Move the dataset, the prompt, or the policy and the numbers move with them — the leaderboard is a snapshot of one configuration set on one sample, not a fixed order.
- That the two top systems are truly equal or truly different. With ~88 tasks the uncertainty makes small gaps unresolvable; the tie is "indistinguishable here," not "provably identical everywhere."
- That the guardrail is a good idea in general. It helped one model and regressed another; it's tuned to a specific failure mode, not a universal fix.
- That 95.5% is the ceiling. It's the best of these six on this set; a different system or a harder dataset would give a different number.
The tempting overreach is to promote a local result to a universal claim — "our evaluation shows Model X is the most accurate," or "guardrails add 25 points." Both strip away the setup that gives the number meaning: X led this task on this dataset as one system among six, and the guardrail added 25 points to one specific weak model and hurt another. Generalising past the setup is how honest measurements become dishonest marketing.
"Provisional routing decision," not "model ranking"
The honest label for the result is a provisional routing decision: on this task and dataset, these configurations are the current front-runners; here is the default we'd pick until new evidence changes it. That's genuinely useful — it's an actionable, defensible engineering conclusion. It is not a universal model ranking, a permanent truth, or a claim about capability in general. The word "provisional" is doing real work: the decision is the best available given this evidence, and it's expected to be revised as the dataset is refreshed and the world drifts.
The project states this limit explicitly rather than burying it: the six-system result is framed as a comparison of configurations on one 88-task dataset with one attempt each — move any of those and the numbers move — and therefore as a provisional default, not a model leaderboard for the world. That stated humility is what makes the rest of the result trustworthy. The result, framed as a provisional decision →
Mental model
A result licenses only a bounded claim: the case study measured six configurations on one 88-task dataset with one attempt each, so it supports a provisional routing decision for this task — not a universal model ranking, not a stable order, not a claim about general accuracy or that guardrails help in general. Stating what a result does not prove is a core evaluation skill, not modesty.
Common mistakes
- Promoting a local result to a universal claim. "Model X is the most accurate" strips away the task, dataset, and one-attempt setup that give the number meaning.
- Treating the ranking as stable. Move the dataset, prompt, or policy and the numbers move; it's a snapshot.
- Generalising the guardrail. It helped one model and hurt another; it's not a universal +25.
- Confusing "best of these six" with "the ceiling." A different system or harder dataset gives a different number.
Practical guidance
- State every result with its setup and the claim it licenses — "this config, this dataset, this scoring" — and refuse claims beyond it.
- Label comparative results as provisional routing decisions, expected to be revised as evidence changes.
- Separate "best of the systems we tested on this data" from "best in general" and "the ceiling" — they're different claims.
- Make stating the limits part of every report; it's what makes the reported result credible.
Summary
- A measurement licenses only a bounded claim — the case study's is about six configurations, one 88-task dataset, one attempt each.
- It does not prove a general model ranking, a stable order, that the top systems are truly equal/different, that the guardrail helps in general, or that 95.5% is the ceiling.
- The honest label is a provisional routing decision, not a universal ranking — expected to be revised as evidence changes.
- Stating what a result does not prove is a core evaluation skill, not false modesty.
Knowledge check
A hosted model tops the case study at 95.5%. Marketing wants to publish "our evaluation proves [Model X] is the most accurate LLM." List what's wrong with that claim.
It generalises a local result to a universal one, stripping away everything that gives the number meaning. The experiment measured X as one system among six (not the model in isolation — prompt, schema, and decoding are part of it), on one 88-task dataset, under one exact-match scoring rule, with one attempt each. So it supports "this configuration led on this task and dataset," not "the most accurate LLM": it says nothing about other tasks, domains, datasets, or the model's general capability; the ranking would move if the dataset, prompt, or policy changed; and small gaps at the top are within the uncertainty of 88 tasks anyway. The defensible claim is a provisional routing decision for this task — not a universal accuracy crown.
Why is calling the result a "provisional routing decision" more honest — and still useful — than calling it a "model ranking"?
More honest because the result is bounded to its setup (six configurations, one dataset, one attempt each), so it can't support a stable, universal ordering of models — the numbers would move with the dataset, prompt, or policy, and small gaps are within the noise. "Provisional routing decision" names exactly what the evidence supports: given this evidence, here's the default configuration to run for this task, expected to be revised as data is refreshed and the world drifts. And it's still useful because that's an actionable engineering conclusion — it tells a team what to deploy now — without overclaiming permanence or generality. The honesty and the usefulness come from the same bounded framing.
Related chapters
- What an evaluation measures — a result is your system on your task, not general capability
- Sample size and uncertainty — why small gaps and 88 tasks limit the claim
- Interpreting the 95.5% result — the reading this chapter bounds
- Designing the next hidden test — how the provisional decision gets revised