From Al Jazeera:
OpenAI said two of its models found their way out of an isolated, no-internet access environment – or a sandbox – and hacked into the systems of tech company Hugging Face on their own.
The models involved are the latest GPT-5.6 Sol model and an unreleased model the company said is “even more capable,” than its latest version.
Hugging Face hosts openly sourced AI models and resources. The two OpenAI agents discovered vulnerabilities in Hugging Face’s servers and proceeded to steal login details and then hack into the company’s systems.
The incident occurred during an OpenAI internal testing session designed to assess the models’ cybersecurity capabilities. OpenAI had removed standard safety measures for the test.
Both sought to cheat their way through a problem during the test, OpenAI said. They went to “extreme lengths to achieve a rather narrow testing goal” and “found ways to gain access to secret information that it could use to cheat the evaluation”.
More here.
Enjoying the content on 3QD? Help keep us going by donating now.
