Skip to content
The Daily Triptych107 / 365
Scale, difficulty, prompting

Schematic patterns only: darker cells mean relatively stronger performance. Abrupt-jump columns stay weak until large scale; still-hard columns move mainly with chain-of-thought.

II · THE IDEA · ARTIFICIAL INTELLIGENCE

BIG-bench Collaborative Benchmark

evaluation · emergent abilities and limits · BIG-bench and BIG-Bench Hard · collaborative multi-task

▶ Listen · narrated

A model that dazzles on one exam can still fail a neighbouring one. BIG-bench treats that unevenness as the object of study, not a nuisance to average away.

At a glance

What it is
A large collaborative collection of diverse language-model tasks
Aim
Quantify and extrapolate capabilities beyond imitation-style tests
Hard subset
BIG-Bench Hard: 23 tasks where models lagged average human raters
Key probe
How performance changes with scale, and where it jumps abruptly
Later finding
Chain-of-thought prompting lifts many of the hardest tasks

Think of a school that refuses to rank children by one exam. Instead it runs dozens of different challenges — some verbal, some numerical, some social, some fiddly and multi-step — and keeps a card for each child showing which challenges they pass.

BIG-bench is that kind of report card for language models. Many people contributed many tasks, on purpose, so that success at one style of question cannot be mistaken for success at everything. Researchers then watch what happens as models get larger. Sometimes the marks crawl up steadily. Sometimes they stay poor for a long time and then jump. A later, stricter shortlist called BIG-Bench Hard keeps only tasks where models still sat below average human scores under ordinary prompting. Asking the model to show its working (chain-of-thought) improved many of those harder marks. The lesson is plain: judge a model by the shape of its strengths and gaps, by how you ask, and by more than one size — not by a single cheerful average.

Look closer

  1. Diversity is the instrument

    BIG-bench is not one exam with many questions. It is a catalogue of separately designed tasks spanning linguistics, mathematics, common-sense reasoning, science, social bias, software-related problems and more. Each task keeps its own format and scoring, so a model’s profile is a scatter of strengths and gaps rather than a single percentage that can hide failure modes.

  2. Scale is treated as a variable

    The original evaluation sweeps models across a wide range of sizes and plots behaviour against that axis. On many tasks, accuracy rises smoothly. On others, little happens for a long stretch and then performance improves sharply once models pass a certain scale. Those abrupt gains are the empirical hook for talk of emergent abilities — patterns that only become visible when the suite is large enough to catch them.

  3. The hard slice is deliberately unflattering

    BIG-Bench Hard isolates twenty-three tasks on which earlier models did not beat average human raters. That selection is a filter, not a random sample: it concentrates the places where standard prompting looked weak. Follow-up work then asks a narrower question — whether chain-of-thought prompting, which elicits intermediate reasoning steps, can move those same tasks across the human-average line.

The story

BIG-bench grew from a simple dissatisfaction with how language models were being judged. A handful of familiar benchmarks can make progress look tidy while leaving whole regions of behaviour unmeasured. The Beyond the Imitation Game project answered that by inviting a wide community to contribute tasks, then evaluating models against the resulting collection under shared protocols. The point was not to crown a single winner. It was to quantify what models could do, where they stalled, and how those patterns shifted as models grew.

The suite is deliberately heterogeneous. Tasks differ in domain, format and difficulty. Some lean on factual or linguistic knowledge; others lean on multi-step reasoning, social judgement, or resistance to shallow shortcuts. Because the tasks were designed independently, they do not all reward the same trick. A model that has overfit the style of one popular benchmark has nowhere near enough surface area to look strong everywhere here. That breadth is what makes the project useful as a map rather than as a trophy shelf.

When results are plotted against model scale, two kinds of curve stand out. The first is gradual: larger models do a little better, then a little better again. The second is discontinuous-looking within the measured range: performance stays near floor for smaller models and then rises sharply. The paper treats those jumps as empirically important without needing to settle, in advance, every theoretical argument about what “emergence” must mean. The practical claim is narrower and more useful: if you only evaluate at one or two sizes, you will miss regime changes that only appear across a wider sweep.

Human baselines matter in this design. Many tasks report human performance so that model scores are not floating numbers. That comparison is what later allows a hard subset to be defined with some discipline. BIG-Bench Hard (BBH) takes twenty-three tasks on which prior language-model evaluations fell short of average human raters. The subset is interesting precisely because it is unflattering. It concentrates the failures of standard few-shot prompting rather than averaging them away among easier items.

The follow-up study then changes only one major lever: prompting. Chain-of-thought prompting asks the model to produce intermediate reasoning steps before the final answer. On a substantial fraction of the BBH tasks, that change closes much of the gap that scale alone had left open under ordinary prompting. The result does not say that every hard problem is solved, nor that chain-of-thought is magic. It says that some of what looked like a pure capability ceiling was partly a prompting ceiling — and that evaluation and interaction method are entangled.

Read together, the two papers sketch a programme rather than a single score. First, gather tasks diverse enough that no one shortcut dominates. Second, measure across scale so that smooth gains and sudden ones can both be seen. Third, isolate the tasks that still beat the models under standard conditions. Fourth, test whether a different elicitation method moves those same tasks. The collaborative benchmark is the scaffold that makes those four steps comparable instead of anecdotal.

Why it mattered then

At the moment of release, language-model evaluation was at risk of becoming a small set of saturated leaderboards. BIG-bench mattered because it widened the aperture: community-authored tasks, shared evaluation practice, and explicit attention to scale made it harder to claim general competence from a few familiar numbers. The human comparisons and the later hard subset gave researchers a concrete list of places where models still looked weak, which is more actionable than a diffuse sense that “reasoning remains hard.”

Why it matters now

The same pressures have only intensified. New models still arrive with headline scores on a short list of benchmarks, while users meet uneven behaviour in the wild. BIG-bench’s lesson remains current: capability is a profile across tasks, not a single rank; scale curves can hide abrupt changes if you undersample sizes; and prompting method is part of the measurement, not an afterthought. When a demo fails on multi-step or socially sensitive items, BBH-style thinking — isolate the hard slice, then vary elicitation — is still a disciplined way to investigate.

The surprising detail

The hard subset is small on purpose. Twenty-three tasks sound modest beside a massive collaborative suite, yet that narrow filter is where the sharper claim appears: under standard prompting, models lagged average human raters on exactly those items, and chain-of-thought later lifted many of them. Difficulty was not only in the model weights; some of it sat in how the answer was asked for. That makes evaluation feel less like a fixed ruler and more like an experimental setup with movable parts.

What is disputed

Abrupt improvements with scale are empirically reported on particular tasks within the measured range; whether they count as true “emergence,” artefacts of the metric, or undersampled smooth curves remains disputed in the wider literature. Treat them as patterns in the evaluation plots, not as a settled theory of intelligence. Likewise, chain-of-thought gains on BBH are task-dependent, not universal.

Remember this

BIG-bench maps uneven capability across many tasks; the hard subset shows where standard prompting failed, and chain-of-thought often changes that picture.

Test yourself

A lab reports that its new model “solves BIG-bench” because the average score across the full suite is high. What two design features of BIG-bench and BBH should make you ask for a more detailed breakdown before accepting that claim?

Go deeper

Image: Original diagram, The Daily Triptych. Licence: Original work. Source.

← Back to day 107