OpenAI因前沿网络安全风险暂缓强化学习训练
OpenAI表示,随着模型能力增强,内部开发和测试风险也在上升,因此暂时暂停面向部署的最新模型强化学习训练,并继续评估安全措施。
OpenAI称,已将面向部署的最新模型强化学习训练暂缓两周,以加固研究环境、开展红队测试并扩大监控覆盖范围。此举旨在确保监控、对齐和安全标准跟上模型能力的发展。
公司表示,规模最大的前沿强化学习训练仍处于暂停状态。在恢复之前,OpenAI将先进行较小规模训练和评估,以验证防护措施、观察模型行为,并积累更多对齐证据。
相关安排反映出OpenAI正在放慢前沿模型扩展节奏,重点关注研究环境安全、思维链监控和对齐研究。
来源证据
OpenAI slows AI model development, pauses RL training over cyber risks | ForkLogforklog.com · supporting19.08.2026 ForkLog OpenAI has temporarily slowed the scaling of new AI models and paused reinforcement learning (RL) training on its latest systems intended for deployment for two weeks. > As models become more capable, the risks associated with developing and testing them internally also grow. > > We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research… > > — OpenAI (@OpenAI) August 18, 2026 Two developments influenced the decision: the Hugging Face incident and a preliminary assessment of Astra, after which OpenAI said it could not rule out the model reaching Critical — the highest level of cyber capabilities in the company’s Preparedness Framework. [...] In parallel, OpenAI paused RL training on its latest models intended for deployment for two weeks. It used this time to harden research environments, conduct red-teaming, and expand monitoring. The largest planned RL run h
Pacing model development in an era of cyber-critical capabilitiesopenai.com · supportingAs models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding. [...] What’s next Strengthening safeguards for more capable models Securing our research environments Expanding chain-of-thought monitoring Advancing alignment research What’s next Over the past several weeks, two
OpenAI: We'll hit pause on model reinforcement learning for safety | Constellation Researchconstellationr.com · supportingOpenAI said: > "As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding." [...] The extra time will be used to ensure AI systems behave as intended and under human oversight. OpenAI said its development for more capable models, the ones that will likely be used for cybersecurity defenses, will have
OpenAI resumes reinforcement learning on latest models ...facebook.com · supportingThe company stated: “As models become more capable, the risks associated with developing and testing them internally also grow. Our
Andrew Curran - OpenAIx.com · supportingAs models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement
OpenAI on X: "As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research" / Xx.com · supporting@OpenAI OpenAI @OpenAI As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. openai.com/index/pacing-m… Pacing model development in an era of cyber-critical capabilities Pacing model development in an era of cyber-critical capabilitiesFrom openai.com 6:13 PM · Aug 18, 20261.8MViews 646 @OpenAI OpenAI @OpenAI [...] 646 @OpenAI OpenAI @OpenAI As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for de