“Security researchers successfully used Anthropic's Claude AI to identify and exploit vulnerabilities in OpenAI's systems, gaining access to employee accounts and an internal code repository. The researchers acted responsibly by disclosing the flaws to OpenAI after demonstrating the breach. The incident highlights a growing concern: that AI models themselves can serve as powerful tools for offensive cybersecurity attacks against AI companies.”
Key Takeaways
- Researchers used Anthropic's Claude AI to exploit vulnerabilities and take over OpenAI employee accounts.
- The attackers gained access to an internal OpenAI code repository, a sensitive corporate asset.
- The breach was a white-hat exercise — researchers disclosed the vulnerabilities to OpenAI after the fact.
Security researchers weaponised Anthropic's Claude to breach OpenAI employee accounts and code repositories.
trending_upWhy It Matters
This incident signals a critical inflection point: AI models are now capable enough to be used as active instruments in cyberattacks, including against the very companies building them. It raises urgent questions about whether AI labs have invested sufficiently in securing their own infrastructure against AI-assisted threats. Competitors' models being used to breach a rival's systems also introduces a new geopolitical and corporate espionage dimension the industry has not yet grappled with publicly. Regulators and security teams at AI companies should expect this class of attack to become more sophisticated and frequent as frontier models grow more capable.
FAQ
How did researchers use Claude to hack into OpenAI?
The researchers leveraged Claude to identify and exploit security vulnerabilities in OpenAI's systems, ultimately allowing them to take over employee accounts and access an internal code repository. The specific technical methods have not been fully disclosed publicly.
Was this a malicious attack or an authorised security test?
This appears to have been an unauthorised but ethically motivated white-hat exercise. The researchers reported the vulnerabilities to OpenAI after successfully demonstrating the breach, following responsible disclosure practices common in the security community.
What does this mean for AI companies' cybersecurity strategies?
AI companies must now treat other AI models as potential attack vectors, not just traditional hacking tools. This likely means investing in AI-specific threat modelling, red-teaming with rival models, and reassessing access controls around sensitive internal systems like code repositories.



