AI · 10h ago
New quiz pits humans against frontier AI on expert-level benchmark
A developer built "Humans vs. HLE," a web-based quiz using Humanity's Last Exam, a benchmark of expert-level questions designed to stump AI models. Players answer multiple-choice questions and see their scores compared to frontier model accuracy and a persistent human leaderboard. The app runs on Cloudflare Workers with answers encrypted server-side to prevent cheating.
Meridian48 take
The quiz is a clever interactive demo, but its real value is in highlighting how quickly AI is closing the gap on expert-level reasoning—a trend that deserves more scrutiny than a single leaderboard can provide.
Read the full reporting
Can You Beat an LLM? Building Humans vs. Humanity's Last Exam →
DEV Community
ai-benchmarksdeveloper-tools