One thought about “risk and control” of frontier AI models.

As these models become more powerful, they can perform complex reasoning and solve complicated tasks without detailing the step-by-step process of their Chain-of-Thought (CoT).

When it comes to assessing risk and control, I understand that CoT logs and analysis are key components. Therefore, if we had limited visibility on those CoT….

I have not seen a lot of rumor about this matter, so to make your weekend more “relaxing”, here is an excerpt from OpenAI on the topic:

“We have found that GPT‑6 Astra is more capable of controlling its own CoT than GPT‑5.6 Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks?”

Source: OpenAI (https://openai.com/index/safety-overview-gpt-6-astra/)

Nice!

CategoriesAI