AI · 1h ago
OpenAI: AI Benchmark Scores Are Not Fixed Capabilities
OpenAI released guidance arguing that benchmark scores reflect a model's performance under specific conditions, not its inherent capability. The company says harness design, compute budget, tool access, and memory management can materially change results. It urges clearer reporting to avoid misleading comparisons between models tested in different setups.
Meridian48 take
This is a necessary corrective to the leaderboard culture, but it also conveniently lets OpenAI hedge on its own models' relative performance.
Read the full reporting
OpenAI Says AI Benchmark Scores Depend on Harnesses, Budgets, and Memory Design →
DEV Community
ai-benchmarksevaluation-standards