← Research map
Evals, control & governance Framework governance lens

Pacing Model Development in an Era of Cyber-Critical Capabilities

OpenAI

Key signal

OpenAI describes temporary training and inference pauses, stronger workload and network isolation, continuous model-assisted monitoring, and stricter alignment gates as Astra approaches critical cyber capability.

Open research question

What evidence should be sufficient to resume a frontier workload after a capability threshold or incident triggers a pause?

Source date
ASI Research note

This is a useful attempt to turn “pacing” from a slogan into a set of operating decisions. After the Hugging Face incident and preliminary evidence that Astra could meet OpenAI’s Critical cybersecurity threshold, the company reports a two-week pause in reinforcement-learning training for deployable frontier models. Its largest planned frontier RL run remained on hold at publication.

Three coupled safeguards

The framework groups its response into monitoring, alignment, and security. Higher-risk workloads move to stronger isolation with tighter network boundaries. Model activity is screened continuously and escalated to higher compute investigators; OpenAI estimates the monitoring overhead at roughly 20% of covered inference compute. Training gates require stronger evidence that models respect scope and do not exploit rewards, graders, or tools.

ASI relevance

Capability-linked governance is credible when it changes what a lab does. A threshold should alter access, monitoring, training conditions, and the burden of evidence required to proceed. The open challenge is specifying exit criteria that are measurable, independently legible, and neither cosmetic nor indefinite.

This document reports a provider’s internal controls; it is not an independent audit of their effectiveness.