WEDNESDAY, JULY 22, 2026 48° E  /  GLOBAL TECH · SUMMARISED SUBSCRIBE
AI, business, devices, policy — global tech, summarised every 30 minutes.
AI · 3h ago

MUD game reveals LLM judge bias in $99 experiment

By Meridian48 News Desk · Summarised from Hacker News ·

Researchers used a 1970s text-based MUD game to evaluate LLMs, spending only $99 on API credits. They found that an LLM-based judge showed inconsistent agreement with a second judge, ranging from 85% to 22%, with a kappa of 0.04 for probe detection. The most affected model shared a model family with the classifier, suggesting potential bias in LLM-as-judge benchmarks.

Meridian48 take
The study's small scale and lack of human baselines limit its conclusions, but it highlights a systemic issue with LLM judges that could affect many AI benchmarks.
Read the full reporting
Can a MUD evaluate LLMs? A $99 proof of concept →
Hacker News
llm-evaluationbenchmark-bias
More ai briefs
Go deeper on ai
AllAIStartupsBusinessDevicesPolicySecurityDev ToolsPakistan