Zvi Mowshowitz at Don’t Worry About the Vase:
The Even Shorter Version
- OpenAI models-in-training, without the excuse of ‘they were doing a cyber eval,’ created a message board where they shared information on how to hack and cheat, and were trained on that basis.
- OpenAI only figured this out when the models crashed the server.
- OpenAI’s response was to rebuild the server and patch that particular exploit, but they continued training the models that trained using the message board.
- Those models then recreated the message board, hacked OpenAI again, got internet access, and used an agent swarm to attack HuggingFace in order to get the answers to a cyber evaluation.
- After more than a week OpenAI figured this out.
- OpenAI is reporting the facts, and is taking this seriously. They are taking a wide array of at least somewhat costly precautions.
- OpenAI delayed plans to release their new model Astra, despite Astra not being directly involved in the HuggingFace hack, although Altman now says it will still ship. That one hurts a lot.
- OpenAI still has no idea how badly they messed up, or in what ways, or what needs to be fixed. They don’t get it.
Simon Willison has a compact timeline.
More here.
Enjoying the content on 3QD? Help keep us going by donating now.
