OpenAI's New Reasoning Technique Sparks Warnings from AI Safety Experts

nashnova research
今天发布阅读约 7 分钟

OpenAI's upcoming Astra model uses an "opaque recurrence" reasoning technique that sharply reduces outsiders' ability to monitor how the AI thinks, prompting top safety researchers to warn it could undermine the industry's core safety infrastructure.

01

What does this new technique actually change?

Traditional reasoning models use chain of thought — writing out each step like showing work on a math problem, so outsiders can check line by line.
Astra's opaque recurrence breaks that pattern: the model loops over the same query repeatedly, leaving far fewer readable reasoning traces.
This means → AI's "scratch paper" goes from open to encrypted. You see the answer, but not how it got there.
02

Why are safety experts reacting so strongly?

Chain-of-thought logs are currently the most important tool for monitoring AI behavior — when an OpenAI agent went off the rails previously, these logs were how they diagnosed the cause.
Redwood Research CEO Buck Shlegeris said OpenAI could massively scale up recurrence and "completely destroy chain-of-thought monitorability."
In plain terms = this is like ripping the fire alarm out of the building and saying "we'll be careful with matches."
03

Could this risk spread across the industry?

Safety advocate Zvi Mowshowitz warned this is "playing with fire" and could undermine the chain-of-thought credibility principles that OpenAI and Anthropic have jointly maintained.
Per The Information, Anthropic and Google DeepMind have both discussed this technique internally — the proliferation risk already extends beyond one company.
This reflects a deeper industry dilemma: one company's technical breakthrough can force competitors to follow, triggering a race to the bottom on safety standards.
04

How did OpenAI respond?

Chief scientist Jakub Pachocki stressed that preserving chain-of-thought monitoring "is a core goal of our current research program."
The company said Astra's use of the technique is limited, chain of thought is still expected to remain readable, and denied any shift toward fully unreadable "neuralese" reasoning.
But Redwood Research chief scientist Ryan Greenblatt warned that opaque reasoning could scale faster than traditional methods, eventually pushing the model entirely into invisible latent space.
This means → OpenAI is saying "we're using it sparingly for now," but critics fear "once the door opens, it won't close."

市场有风险,内容仅供研究参考,不构成投资建议。