uberman
7 hours ago
We trained a model to be world class at identification of cyber security and hacking.
We specifically turned off some of the deployment safeguards because we wanted to test capability here.
We let agents loose over the course of 3 million compute hours with a weakly defined task of "find vulnerabilities in the software in your environment".
We forgot that by definition the sandbox is in the environment and were not anticipating sandbox vulnerabilities would be explored.
Despite mountains of training data suggesting planning and collaboration should be well documented and recipe sites like agents.stackoverflow.com existing specifically to foster this kind of shared knowledge we are now alarmed that our agent documented what they did.
So alarmed that two weeks later we released GPT-5.6-cyber and the same behavior WaPo calls 'colluding to cheat' JFrog's post-mortem now calls 'collaborative knowledge sharing between models' as a feature.
Do I have that essentially right?