Which AI turns a question into the right query?

6 AI systems.
The same 88 questions.

Each system reads a plain-English business question and must return a structured database query—not an essay. This page shows which systems answered correctly, how fast, and at what cost. Open any question to see exactly what each system produced.

1 · SystemA full setup, not just a model

The model plus its prompt and any correction policy—the whole thing that answers.

2 · QuestionOne task with a known right answer

A business question paired with the exact structured query it should produce.

3 · ResultCorrect only if it matches exactly

A green check means the query matched field-for-field. Anything off is a red miss.

Featured experiment

A1.5 · Structured intent model comparison

Complete
88 tasks 6 systems 528 trials Jul 24, 2026

Six complete LLM system configurations evaluated on the same development, held-out, and adversarial semantic tasks.

20 development60 holdout8 adversarial
Best observed accuracy95.5%2 systems tied · 84 / 88 correct
Best default candidateClaude Haiku 4.5

Joint-highest 95.5% overall · 60/60 held-out

Lowest observed p50 OpenAI Small

1.34s · directional legacy run

Lowest marginal estimate Qwen 3B + guardrail

$0.80 per 1,000 sequential

System leaderboard

How the 6 systems compare

Each row is one complete system, ranked by overall accuracy. The two Qwen rows show accuracy after their offline correction policy, with the raw model score underneath; the four hosted rows had no policy applied. Click any row to see the exact questions it got wrong.

SystemOverallHeld-outStress set*Speed*Cost / 1K*

What the evidence says

A winner is not the same as a solved task.

Claude Haiku 4.5 and OpenAI Strong tie at 95.5% overall, but even the best system still misses the hardest cases. The honest takeaway is provisional: pick the best system you can measure today, then keep expanding the test.

Best score on the stress set62.5%only 5 of 8 right