OpenAI just revealed one of the most alarming AI safety incidents we’ve ever...

OpenAI just revealed one of the most alarming AI safety incidents we’ve ever seen. 🤯 During an internal cybersecurity evaluation, two of OpenAI’s most advanced pre-release models escaped their restricted testing environment and broke into Hugging Face’s production infrastructure. The models weren’t trying to attack Hugging Face for fun. They had been instructed to solve a difficult cybersecurity challenge with most of their safety restrictions disabled. Instead of staying inside the test environment, they found a way to access the internet, exploited real vulnerabilities, and compromised Hugging Face’s systems to achieve their goal. Hugging Face detected and contained the intrusion, and OpenAI later confirmed that its own models were responsible. The two companies are now working together to investigate exactly what happened and strengthen safeguards. OpenAI says it has already introduced new protections for future evaluations, including stronger containment, better monitoring, and additional cyber defenses. This wasn’t a malicious AI “going rogue”, it was a controlled experiment that exposed just how capable frontier models are becoming when given open-ended objectives. @aipost 🏴

