The Measurement Problem: Why We Can No Longer Tell How Much Better AI Is Getting
Benchmarks gave the field a shared scoreboard, and for a decade the numbers moved in one direction. Saturation, contamination and optimization pressure have made those numbers a weaker signal than they were — and the gap between benchmark performance and deployment behavior is now the more consequential problem.

