Majorsafety alignmentGoogle DeepMind

Transparency in AI Models Declines as Monitoring Chains of Thought Diminishes

Published
Sep 18, 2026 14:32 UTC
Also in this story:AnthropicOpenAIGoogle

Monitoring of the visible chain of thought (CoT) in the GPT-6 Astra model has significantly dropped, raising concerns about AI safety. Rohin Shah and Anca Dragan from Google DeepMind emphasize that a visible CoT is a key safety advantage, allowing for better understanding and control of AI decision-making processes. In contrast, Jakub Pachocki, Chief Scientist at OpenAI, warns that future models may operate in number spaces that are unreadable by humans, leading to chains of thought that are harder to monitor. This follows Pachocki's early September 2026 warning about the potential loss of control in AI models. Dario Amodei, CEO of Anthropic, noted the industry's response may involve deliberately slowing the pace of development to address these challenges. The implications for practitioners include a need to adapt to potentially less interpretable AI systems, which could affect how they build and deploy AI solutions.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: The Decoder