AI · 1h ago
Agent Performance Depends More on Harness Than Model
Kimi K3 shows different performance under different harnesses, with Moonshot disclosing harness as an evaluation condition. Research shows swapping harnesses can move SWE-bench scores by up to 48 points, dwarfing typical model improvements. The industry focuses on model gains while the harness pillar remains the bottleneck for complex tasks.
Meridian48 take
This reframes the AI agent debate: vendors race to improve models, but the real leverage may be in the orchestration layer that most benchmarks ignore.
ai-agentsllm-harness