Security · 18h ago
Vercel launches DeepsecBench to benchmark AI vulnerability detection
Vercel released DeepsecBench, a benchmark evaluating how well AI models find cybersecurity vulnerabilities in application code. The top model, GPT-5.6 Sol, scored 35.58 with 30.7% recall at $55.98 per run. Cheaper models like Grok 4.5 achieved 15.58 score for $5.60, showing falling costs for capable analysis.
Meridian48 take
The benchmark's secret construction prevents training data leakage, but the low recall rates (best 30.7%) highlight that AI-assisted vulnerability scanning is still far from replacing human review.
Read the full reporting
DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities →
Vercel
ai-benchmarkvulnerability-detection