stego-tech
today at 8:11 PM
As a PC gamer who grew up in the 00s, this has been something Iāve tried to warn ardent LLM and model enthusiasts about for quite some time.
Benchmarks are handy when theyāre new, novel, and constantly changing. The second you let even a single aspect of it stagnate, it becomes a gameable score rather than a useful metric. In PC Gaming, we saw vendors optimize for specific titles, benchmark tools, and scenarios at the expense of general performance, and eventually the industry had a ācome to Jesusā moment where we had to collectively decide how to move forward from an industry built on thoroughly gamed benchmarks, with entities like Gamersā Nexus and Digital Foundry being the end results of that falling out.
LLMs were always going to end up the same way, because the people building the benchmarks - well-intentioned as they were - ultimately fell into the exact same traps with fixed scoring rubrics, known test questions, and believing in some form of ācompletenessā that could be attained or achieved. The net result are models consistently scoring better on benchmarks but also seeing diminishing returns and rising vulnerabilities, because actual improvement or utility isnāt what theyāre being optimized for so much as bragging rights. Itās why thereās so much growing interest in things like MoE execution on unified memory platforms as a means of porting larger models to consumer kit, or ternary models (shoutout to Bonsai) as a means of reducing overall size: both take leading edge, benchmark-saturating models and show that with minimal score loss, they function about as well as frontier models might.
Building a new benchmark wonāt solve the problem, either. To move forward, we must evaluate LLMs objectively and with continuously evolving workloads. More āpelican on a bicycleā stuff, but from varying perspectives and use cases. Radiologists putting models through their paces with usable sample data they donāt share with AI labs, or IT folks tasking agents with bootstrapping specific, real-world workloads. To prove general intelligence, we need more specialists evaluating them specifically and generally in ways that are transparent to consumers but difficult or impossible for AI companies to prepare against.
Only then will scoring values matter.
throw10920
today at 8:14 PM
> Building a new benchmark wonāt solve the problem, either.
It will if the benchmark is proprietary. If you can't train on it, then it's extremely difficult to game, and if it's hard enough, then it's economically more efficient to just...make the model smarter