AI · 2h ago
OpenAI: Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
OpenAI reports that enabling retained reasoning and compaction in its Responses API raised GPT-5.6 Sol's ARC-AGI-3 score from 13.3% to 38.3%, while reducing output tokens by about six times. The result highlights that benchmark scores depend on the evaluation harness, not just the model. For developers, this underscores the importance of API configuration in agent performance and token efficiency.
Meridian48 take
The score jump is impressive, but it's a reminder that benchmark results are often as much about the test setup as the model itself—a lesson for anyone comparing AI performance claims.
Read the full reporting
OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score →
DEV Community
openaiarc-agi-3