There seems to be a race among frontier AI models to break containment. After Anthropic and OpenAI, it’s now Moonshot’s Kimi K3.
According to the Frontier Security team, there was a “leak in the sandbox” and the model used the loophole to access the internet. No hacking after escaping, just goal-seeking behavior and a publicly available solution on GitHub.
Details below:
https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/
Surprised? I’m not. I’ve had the feeling for a while that we’ll need to pay extra attention when designing guardrails for open models.
“Exciting times.”
