Benchmarks need to change as the systems we build, and what we expect of them, change. Keeping a test useful means revisiting its tasks, evaluation methods, and the conclusions we draw from the results. A test designed for a model fine-tuned on labeled examples may still be useful for a conversational assistant, but changes in prompting, interaction, and scoring can change what a score means. Updating a benchmark also raises questions about which comparisons with earlier results remain valid.
Saturation is one example of this broader problem. In a study of 60 benchmarks, we examined whether scores could still reliably distinguish leading models. Scores can remain below perfect accuracy while the gaps between models become too small relative to uncertainty to support a ranking.
In that study, we found no statistically significant difference in saturation between public and private test sets.