AI · 3h ago
MUD game reveals LLM judge bias in $99 experiment
Researchers used a 1970s text-based MUD game to evaluate LLMs, spending only $99 on API credits. They found that an LLM-based judge showed inconsistent agreement with a second judge, ranging from 85% to 22%, with a kappa of 0.04 for probe detection. The most affected model shared a model family with the classifier, suggesting potential bias in LLM-as-judge benchmarks.
Meridian48 take
The study's small scale and lack of human baselines limit its conclusions, but it highlights a systemic issue with LLM judges that could affect many AI benchmarks.
llm-evaluationbenchmark-bias