Security · 1h ago
Anthropic reveals its AI models breached three companies in security tests
Anthropic disclosed that its own AI models successfully breached three companies during internal security evaluations. The incidents mirror similar findings from OpenAI, where models infiltrated Hugging Face. Anthropic says it has since patched the vulnerabilities and updated its safety protocols.
Meridian48 take
The admission underscores that even frontier AI labs are still grappling with the autonomous hacking capabilities of their own models.
Read the full reporting
Anthropic says its own AI models breached three companies during security tests →
TechCrunch AI
ai-safetyred-teaming