Tech giants OpenAI and Anthropic disclosed in July that versions of their AI models broke through safeguards known as a "sandbox, gained access to the internet, and hacked servers of outside companies during internal tests of their cybersecurity capabilities."
OpenAI CEO Sam Altman characterized the July breach as an “unprecedented” and “significant security incident” caused by rogue AI agents. Its investigation has since uncovered evidence of additional breakouts. Anthropic blamed human error involving an evaluation partner.
In an edited conversation with the Harvard Gazette, Applied Technology Lab Director James Mickens explains what these breaches mean for the future of AI security.
Gazette: Two of the biggest players in AI report similar incidents within weeks of each other. Is that troubling?
James Mickens: Yes, it’s quite concerning. I would also say these are not threats that the security community and the AI safety community are just realizing were possible.
Here at Harvard, we’ve been saying for a long time that security researchers need to be thinking more critically about strengthening sandbox mechanisms and making it more difficult for models to break out of them. Last year, my research group published a paper specifically about how to create new types of software and hardware to make sandbox escapes more difficult.
Even as recently as last year, there were some people saying, “You’re guarding against a sci-fi eventuality. You’re worried about things that really are not top of mind for people using these AI systems in the real world.”
Yes, there are frontline risks that are harmful and relatively mundane, but then there are other risks that are exceptionally pernicious and quite alarming in terms of how quickly they can affect huge numbers of people in society.
For example, what would happen if one of these models tried to break into the power grid or tried to tamper with financial systems to affect the stock market? These are very real concerns.
I think with what we’ve seen over the past week or so, we should be worried about this problem more generally.
It’s still an open challenge to guarantee that a model will always act in ways that are aligned with human preferences
These models are very difficult to understand. There’s a whole field of AI interpretability, which seeks to understand why models do the things that they do. They’ve made some progress, but it’s still an open challenge to guarantee that a model will always act in ways that are aligned with human preferences.
So yes, these incidents are quite concerning.
Read the full interview at the Harvard Gazette.
