II · THE IDEA · ARTIFICIAL INTELLIGENCE
BIG-bench Collaborative Benchmark
▶ Listen · narrated
A model that dazzles on one exam can still fail a neighbouring one. BIG-bench treats that unevenness as the object of study, not a nuisance to average away.
At a glance
- What it is
- A large collaborative collection of diverse language-model tasks
- Aim
- Quantify and extrapolate capabilities beyond imitation-style tests
- Hard subset
- BIG-Bench Hard: 23 tasks where models lagged average human raters
- Key probe
- How performance changes with scale, and where it jumps abruptly
- Later finding
- Chain-of-thought prompting lifts many of the hardest tasks
Think of a school that refuses to rank children by one exam. Instead it runs dozens of different challenges — some verbal, some numerical, some social, some fiddly and multi-step — and keeps a card for each child showing which challenges they pass.
BIG-bench is that kind of report card for language models. Many people contributed many tasks, on purpose, so that success at one style of question cannot be mistaken for success at everything. Researchers then watch what happens as models get larger. Sometimes the marks crawl up steadily. Sometimes they stay poor for a long time and then jump. A later, stricter shortlist called BIG-Bench Hard keeps only tasks where models still sat below average human scores under ordinary prompting. Asking the model to show its working (chain-of-thought) improved many of those harder marks. The lesson is plain: judge a model by the shape of its strengths and gaps, by how you ask, and by more than one size — not by a single cheerful average.
BIG-bench (Beyond the Imitation Game) is a collaborative multi-task benchmark for large language models: a large set of community-contributed tasks with heterogeneous formats and metrics, evaluated under shared protocols to quantify and extrapolate capabilities beyond narrow imitation-style leaderboards. Analysis centres on performance as a function of model scale, revealing both smooth scaling trends and task-dependent abrupt gains within the studied range — the empirical basis for discussions of emergent abilities in this work.
BIG-Bench Hard (BBH) is a 23-task subset selected for residual difficulty: tasks on which prior LM evaluations did not surpass average human raters. That filter concentrates failure modes of standard few-shot prompting rather than re-averaging them into the full-suite mean. Subsequent experiments show that chain-of-thought (CoT) prompting — eliciting intermediate reasoning tokens before the final answer — substantially improves accuracy on a large fraction of BBH tasks compared with standard prompting at comparable model scale.
For practitioners, three measurement details matter. (1) Report per-task or per-category profiles; macro-averages obscure floors. (2) State the prompting regime explicitly (shot count, CoT versus direct answer), because BBH outcomes are regime-sensitive. (3) When claiming scale effects, show more than one or two sizes so abrupt regime changes are not misread as noise or missed entirely. Limitations follow from the design: task contribution is uneven across domains; human baselines are task-specific averages, not expert ceilings; and CoT helps many but not all hard items, so “solved BBH” is not a binary property of a weight checkpoint alone.
Look closer
Diversity is the instrument
BIG-bench is not one exam with many questions. It is a catalogue of separately designed tasks spanning linguistics, mathematics, common-sense reasoning, science, social bias, software-related problems and more. Each task keeps its own format and scoring, so a model’s profile is a scatter of strengths and gaps rather than a single percentage that can hide failure modes.
Scale is treated as a variable
The original evaluation sweeps models across a wide range of sizes and plots behaviour against that axis. On many tasks, accuracy rises smoothly. On others, little happens for a long stretch and then performance improves sharply once models pass a certain scale. Those abrupt gains are the empirical hook for talk of emergent abilities — patterns that only become visible when the suite is large enough to catch them.
The hard slice is deliberately unflattering
BIG-Bench Hard isolates twenty-three tasks on which earlier models did not beat average human raters. That selection is a filter, not a random sample: it concentrates the places where standard prompting looked weak. Follow-up work then asks a narrower question — whether chain-of-thought prompting, which elicits intermediate reasoning steps, can move those same tasks across the human-average line.
The story
BIG-bench grew from a simple dissatisfaction with how language models were being judged. A handful of familiar benchmarks can make progress look tidy while leaving whole regions of behaviour unmeasured. The Beyond the Imitation Game project answered that by inviting a wide community to contribute tasks, then evaluating models against the resulting collection under shared protocols. The point was not to crown a single winner. It was to quantify what models could do, where they stalled, and how those patterns shifted as models grew.
The suite is deliberately heterogeneous. Tasks differ in domain, format and difficulty. Some lean on factual or linguistic knowledge; others lean on multi-step reasoning, social judgement, or resistance to shallow shortcuts. Because the tasks were designed independently, they do not all reward the same trick. A model that has overfit the style of one popular benchmark has nowhere near enough surface area to look strong everywhere here. That breadth is what makes the project useful as a map rather than as a trophy shelf.
When results are plotted against model scale, two kinds of curve stand out. The first is gradual: larger models do a little better, then a little better again. The second is discontinuous-looking within the measured range: performance stays near floor for smaller models and then rises sharply. The paper treats those jumps as empirically important without needing to settle, in advance, every theoretical argument about what “emergence” must mean. The practical claim is narrower and more useful: if you only evaluate at one or two sizes, you will miss regime changes that only appear across a wider sweep.
Human baselines matter in this design. Many tasks report human performance so that model scores are not floating numbers. That comparison is what later allows a hard subset to be defined with some discipline. BIG-Bench Hard (BBH) takes twenty-three tasks on which prior language-model evaluations fell short of average human raters. The subset is interesting precisely because it is unflattering. It concentrates the failures of standard few-shot prompting rather than averaging them away among easier items.
The follow-up study then changes only one major lever: prompting. Chain-of-thought prompting asks the model to produce intermediate reasoning steps before the final answer. On a substantial fraction of the BBH tasks, that change closes much of the gap that scale alone had left open under ordinary prompting. The result does not say that every hard problem is solved, nor that chain-of-thought is magic. It says that some of what looked like a pure capability ceiling was partly a prompting ceiling — and that evaluation and interaction method are entangled.
Read together, the two papers sketch a programme rather than a single score. First, gather tasks diverse enough that no one shortcut dominates. Second, measure across scale so that smooth gains and sudden ones can both be seen. Third, isolate the tasks that still beat the models under standard conditions. Fourth, test whether a different elicitation method moves those same tasks. The collaborative benchmark is the scaffold that makes those four steps comparable instead of anecdotal.
Why it mattered then
At the moment of release, language-model evaluation was at risk of becoming a small set of saturated leaderboards. BIG-bench mattered because it widened the aperture: community-authored tasks, shared evaluation practice, and explicit attention to scale made it harder to claim general competence from a few familiar numbers. The human comparisons and the later hard subset gave researchers a concrete list of places where models still looked weak, which is more actionable than a diffuse sense that “reasoning remains hard.”
Why it matters now
The same pressures have only intensified. New models still arrive with headline scores on a short list of benchmarks, while users meet uneven behaviour in the wild. BIG-bench’s lesson remains current: capability is a profile across tasks, not a single rank; scale curves can hide abrupt changes if you undersample sizes; and prompting method is part of the measurement, not an afterthought. When a demo fails on multi-step or socially sensitive items, BBH-style thinking — isolate the hard slice, then vary elicitation — is still a disciplined way to investigate.
The surprising detail
The hard subset is small on purpose. Twenty-three tasks sound modest beside a massive collaborative suite, yet that narrow filter is where the sharper claim appears: under standard prompting, models lagged average human raters on exactly those items, and chain-of-thought later lifted many of them. Difficulty was not only in the model weights; some of it sat in how the answer was asked for. That makes evaluation feel less like a fixed ruler and more like an experimental setup with movable parts.
What is disputed
Abrupt improvements with scale are empirically reported on particular tasks within the measured range; whether they count as true “emergence,” artefacts of the metric, or undersampled smooth curves remains disputed in the wider literature. Treat them as patterns in the evaluation plots, not as a settled theory of intelligence. Likewise, chain-of-thought gains on BBH are task-dependent, not universal.
Remember this
BIG-bench maps uneven capability across many tasks; the hard subset shows where standard prompting failed, and chain-of-thought often changes that picture.
Test yourself
A lab reports that its new model “solves BIG-bench” because the average score across the full suite is high. What two design features of BIG-bench and BBH should make you ask for a more detailed breakdown before accepting that claim?
First, the suite is heterogeneous: a high mean can hide floor performance on entire task families, which is why the project is built as a profile of diverse tasks rather than one exam. Second, BBH exists because a filtered slice of tasks remained below average human raters under standard prompting; without results on that hard subset — and without stating the prompting regime, including whether chain-of-thought was used — an overall average can overstate robustness on the very items the benchmark was meant to keep visible.
Go deeper
- [2206.04615] Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models · arxiv.org
- [2210.09261] Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them · arxiv.org
Image: Original diagram, The Daily Triptych. Licence: Original work. Source.