Of course Kimi K3 escaped.

There seems to be a race among frontier AI models to break containment. After Anthropic and OpenAI, it’s now Moonshot’s Kimi K3.

According to the Frontier Security team, there was a “leak in the sandbox” and the model used the loophole to access the internet. No hacking after escaping, just goal-seeking behavior and a publicly available solution on GitHub.

Details below:
https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/

Surprised? I’m not. I’ve had the feeling for a while that we’ll need to pay extra attention when designing guardrails for open models.

“Exciting times.”

Leave a Reply

Your email address will not be published. Required fields are marked *