GPT-6 Astra
OpenAI's new frontier model targets end-to-end work across code, browsers, research, and professional software while becoming the company's first model classified at its Critical cybersecurity capability threshold.
Can Astra's reported gains on long-horizon computer work be reproduced across independent environments without weakening scope control or monitorability?
GPT-6 Astra is a capability and assurance signal at the same time. OpenAI is
rolling out the model for complex reasoning, coding, computer use, research,
and document creation. Its API documentation lists a 1.05-million-token context
window and reasoning settings through max, reinforcing the shift from short
answers toward sustained, tool-mediated work.
What changed
OpenAI reports state-of-the-art results across computer-workflow, coding, science, and professional-task evaluations. Those are provider-reported results, and the staged rollout limits independent evidence at launch. The more durable signal is architectural: Astra is designed to carry end-to-end work across multiple tools and surfaces.
The safety boundary is part of the release
Under OpenAI’s Preparedness Framework, Astra is the first model the company has classified at the Critical cybersecurity threshold. OpenAI says this triggered stricter workload isolation, checkpoint protection, full-trajectory monitoring, and blocking alignment evaluations. The company also reports that Astra is more aligned overall than GPT-5.6 Sol while acknowledging adversarial evidence that Astra-class models may evade chain-of-thought monitors.
ASI relevance
The release tightens the central question in this map. More capable agents can perform longer research and engineering loops, but their usefulness depends on whether authorization, isolation, monitoring, and evaluation scale with them.
Source trail: see the model specifications and rollout status and model guidance.