Ask Cat › AI Tool Summary › OpenRouter
AI2 Took Apart 34,000 Benchmark Questions: The Social-Bias Benchmark Was Mostly Measuring Reasoning
- Free tier:There is
- Cheapest paid plan:US$0/mo and up
- Free quota:More than 25 free models, 4 free providers, daily limit of 50 …
- Last checked:2026-09-21
Article last updated:2026-09-08
On 2026-09-01 the Allen Institute for AI published BenchMIRT, which does something rarely done: it evaluates the benchmarks instead of the models. Verified 2026-09-08.
Scale: 100 LLMs, 16 benchmarks, more than 34,000 questions, analysed with multidimensional Item Response Theory (MIRT) borrowed from psychometrics, question by question, to see what each item actually discriminates.
1. The finding worth keeping
The analysis was not told how many dimensions to look for. Two emerged on their own: safety and general reasoning.
Which produced the sharp conclusion: benchmarks do not always measure what they were built to measure.
The example given is BBQ, designed to test social bias. It aligned much more strongly with general reasoning than with safety.
Plainly: a high BBQ score more likely means the model reads the question well and reasons well, not necessarily that it is less biased.
2. Two numbers that matter when picking a model
| Number | Meaning |
|---|---|
| 10% of questions | Retain nearly the same capability rankings as the full benchmark |
| 79% accuracy | Predicting model performance on unseen questions |
The first one says something useful: about nine in ten benchmark questions contribute little to telling models apart. Nearly every model gets them right, or nearly every model gets them wrong. Neither discriminates.
That explains something you may already have suspected: a one- or two-point gap between two models usually means nothing. The gap likely comes from a handful of genuinely discriminating items — or from noise.
3. Which benchmarks were covered
Six reasoning benchmarks (including MMLU-Pro, GPQA, MATH, BBH) and ten safety evaluations (including HarmBench, StrongReject, WildJailbreak, BBQ, WMDP, XSTest).
The list is itself useful: the scores on vendor marketing charts usually come from exactly these.
4. How to read a leaderboard now
The paper is not an argument for ignoring benchmarks. It suggests three smarter readings:
- Ask which dimension a benchmark loads on, not just who scored higher. A “safety benchmark” may not be measuring safety.
- Treat a narrow lead as no lead. Unless the gap is large, rank order probably says nothing about your workload.
- Test your own task. The most practical rule: ten real examples from your own work tell you more than a 14,000-question public benchmark. Public items were not designed for your use case.
5. What this changes for readers
If you choose between models — especially through a service like OpenRouter that fronts many providers — the value here is recalibrating how much you trust leaderboards. Not distrust: just knowing the resolution is lower than it looks.
The cheapest useful move: pick two or three candidates, run ten real tasks of your own through each, and see which one flows. That costs under an hour and settles the question better than ten leaderboard analyses.
Comparing many models at once: OpenRouter tool page. A related look at third-party benchmark methodology: four caveats on the GPT-6 Astra robot-arm evaluation.
Source read directly on 2026-09-08: BenchMIRT: What are LLM benchmarks actually measuring? by the Allen Institute for AI (2026-09-01). The 100 models, 16 benchmarks, 34,000+ questions, the two emergent dimensions, the BBQ finding, the 10% retention result and the 79% prediction accuracy are all as published there. We have not reproduced the analysis or independently verified those figures. Sections 4 and 5 are our own view.
Let's take a look at these
- OpenRouter Comprehensive Introduction: Pricing, Features, and Actual Limitations
- OpenRouter Is the free quota enough?
- OpenRouter Alternatives
- Comprehensive Free Quota List for All Tools
More verified articles on this tool
- Five Brakes to Set Before an AI Agent Touches Money: A Checklist Derived From the US$12,431 Invoice Incident
- GPT-6 Astra Lands on OpenRouter: Five Providers for One Model, and the Priciest Costs 4x the Cheapest (Checked Sept 2026)
- How to Pin OpenRouter to the Cheap Provider: order, only, sort and max_price, with JSON You Can Paste (2026 Guide)

