“In a moment of crisis for the field—where we struggle to devise tests today’s systems cannot ace—this book places benchmarks in the spotlight where they belong and asks the almost heretical question: why do they work as well as they do? Engrossing, page-turning, lucid, and crucial. Moritz Hardt has written a masterpiece.”
—Brian Christian, author of The Alignment Problem: Machine Learning and Human Values
“From the calculated moves of grandmaster board games to the frontiers of statistical physics and protein chemistry, Moritz Hardt reveals the story of the benchmark, the intelligence yardstick that transformed AI from a hope and a prayer into a trillion-dollar arms race. Written with the deep insight of a leading expert, this is the definitive look at the benchmarks that define our future—and the secrets they keep.”
—David L. Donoho, Stanford University
“For decades, machine learning has progressed with a predictable cycle of training and testing statistical machines on curated datasets. Moritz Hardt explains why this simple approach worked better than most of us expected, and why it faces daunting challenges in the age of pervasive, internet-scale AI. This is an essential read for anyone seeking to understand what all these benchmark scores really mean.”
—Léon Bottou, Flatiron Institute
“Excellently written, Moritz Hardt provides a fascinating read on a topic of central importance to AI: the use, science, and implications of benchmarks in machine learning. Whether you are an expert, an interested student, or a member of the general public curious about AI—including how we got here and where things are going—you will learn something valuable and new from this book.”
—Avrim Blum, Toyota Technological Institute at Chicago
From its roots, machine learning embraces the anything goes
principle of scientific discovery. Machine learning
benchmarks become the iron rule to tame the anything goes.
But after decades of service, a crisis grips the
benchmarking enterprise.
The mathematical foundations of machine learning follow the
astronomical conception of society: Populations are
probability distributions. Optimal predictors minimize loss
functions on a probability distribution.
A single statistical problem illuminates much of the
mathematical tools necessary for benchmarking. The key
lesson is that sample requirements grow quadratically in
the inverse of the difference we try to detect.
The holdout method separates training and testing data,
anything goes on the training data, iron rule on the
testing data. Not all uses of the holdout method are alike.
Statistics prescribes the iron vault for test data. But the
empirical reality of machine learning benchmarks couldn’t
be further from the prescription. Repeated adaptive testing
brings theoretical risks and practical power.
A replication crisis has long gripped the empirical
sciences. Statistical practice is vulnerable for
fundamental reasons. Under competition, researcher degrees
of freedom outwit statistical measurement.
The preconditions for crisis exist in machine learning,
too. And yet, the situation in machine learning is
different. While accuracy numbers don’t replicate, model
rankings replicate to a significant degree.
If machine learning thwarted scientific crisis, the
question is why. Some powerful explanations emerge. Key are
the social norms and practices of the community rather than
statistical methodology.
9
Labeling and
annotationprint only
If the holdout method is the greatest unsung hero, data
annotation is not far behind. But conventional wisdom
clouds the subtle role that annotation plays for
benchmarking.
The ImageNet era ends as attention shifts to powerful
generative models trained on the internet. The new era also
marks a turning point for machine learning benchmarks.
After training, alignment fits models to human preferences.
Part of the post-training pipeline, alignment transforms
evaluation results. How post-training makes such a
difference brings new challenges for benchmarking.
Multi-task benchmarks promise a holistic evaluation of
complex models. An analogy with voting systems reveals
limitations in multi-task benchmarks. Greater diversity
comes at the cost of greater sensitivity to artifacts.
Models deployed at scale always influence future data, a
phenomenon called performativity. Performativity breaks
evaluation and creates the problem of data feedback loops.
Dynamic benchmarks try to make a virtue out of it.
As models gain in capabilities, human supervision
increasingly becomes a bottleneck. The hope is that models
will supervise and evaluate each other, but there are
limits to automatic evaluation.