AI · 9h ago
Two API settings triple GPT-5.6 scores on ARC-AGI-3 benchmark
OpenAI found that enabling two API settings—retaining reasoning traces and enabling compaction—tripled GPT-5.6's scores on the ARC-AGI-3 benchmark. The improvements boosted both accuracy and efficiency without additional training. This suggests simple configuration changes can unlock significant performance gains in frontier models.
Meridian48 take
The finding is notable but the benchmark is narrow; real-world gains may be less dramatic.
Read the full reporting
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark →
OpenAI
openaibenchmark-performance