179 questions
No questions match those filters.
A leaderboard number like Chatbot Arena ELO looks objec...
This is one of the questions in the full AI/ML interview bank. Pro unlocks all 1789 questions; Premium includes the same bank plus the highest daily Practice limit.
See plansThe moment a metric becomes the thing labs are judged and marketed on, Goodhart’s law kicks in: it stops measuring what it was designed to measure and starts measuring whoever is best at optimizing for it specifically. Arena-style ELO is vulnerable to protocol-level gaming that has nothing to do with model quality — documented cases include providers getting privileged access to the comparison pool, submitting multiple variants under one label, and having influence over which models their entries get compared against, all of which can move a leaderboard number without any underlying capability change.
There are also structural biases baked into the measurement itself, independent of manipulation: voters skew toward technically literate, English-speaking users, and both human and AI judges show a systematic preference for longer, more confidently-formatted answers over accurate but terse ones, so the score partly measures style. The practical response isn’t to discard leaderboards, it’s to treat any single benchmark as one noisy signal: cross-validate a claimed gain against a benchmark built on a different methodology, treat score differences under some threshold as noise rather than a real ranking change, and specifically check whether an improvement on one benchmark is concentrated on a slice of behavior that the benchmark you actually care about doesn’t sample.