Anthropic发布第二份负责任扩展政策风险报告
Anthropic表示,其第二份公开风险报告已发布,内容涵盖模型风险、现有防护措施以及后续安全计划。
Anthropic于2026年8月14日发布第二份公开风险报告。报告依据公司的“负责任扩展政策”编制,旨在说明其人工智能系统可能带来的风险,以及公司目前应对这些风险的准备程度。
官方信息显示,报告覆盖此前报告发布至2026年7月15日期间的模型风险、相关缓解措施和各风险类别的后续计划。该政策要求Anthropic定期公开风险评估,以提高先进模型开发过程的透明度。
此次发布延续了Anthropic定期披露安全信息的做法,但官方材料并未表示报告意味着某一项具体风险已经发生。
来源证据
Anthropic publishes second Risk Report under Responsible Scaling Policycryptobriefing.com · supportingAnthropic publishes second Risk Report under Responsible Scaling Policy The AI company is trying to build a safety playbook for models that keep getting more powerful, complete with third-party audits and public scorecards. by Editorial Team Share Add us on Google Via abc7news.com Anthropic has released its second public Risk Report, a detailed assessment of what could go wrong with its most advanced AI systems and what the company plans to do about it. The report falls under the company’s Responsible Scaling Policy, a framework that essentially says: before we make these models more capable, we need to understand what new dangers come with that capability. ## What the Responsible Scaling Policy actually requires [...] Share Add us on Google by Editorial Team Anthropic has released its second public Risk Report, a detailed assessment of what could go wrong with its most advanced AI systems and what the company plans to do about it. The report falls under the company’s Res
Announcing Anthropic's Responsible Scaling Policyanthropic.com · supporting# Introducing Anthropic's Responsible Scaling Policy Today, we’re publishing our Responsible Scaling Policy (RSP) – a series of technical and organizational protocols that we’re adopting to help us manage the risks of developing increasingly capable AI systems. As AI models become more capable, we believe that they will create major economic and social value, but will also present increasingly severe risks. Our RSP focuses on catastrophic risks – those where an AI model directly causes large scale devastation. Such risks can come from deliberate misuse of models (for example use by terrorists or state actors to create bioweapons) or from models that cause destruction by acting autonomously in ways contrary to the intent of their designers.
Anthropic’s Responsible Scaling Policy \ Anthropicanthropic.com · supporting## August 14, 2026 We shared our August 2026 Risk Report. Our Risk Reports aim to provide a direct, candid, and informative description of how we see the risks of our systems and our state of preparedness for them (particularly the catastrophic risks addressed in our Responsible Scaling Policy). It covers the risks of Anthropic's models and actions between our previous February 2026 risk report and the report's coverage date of July 15, as well as our mitigations for those risks and our forward-looking plans in each category. ## July 8, 2026 [...] See the PDF Version 3.4 and redline (effective July 8, 2026) Version 3.3 and redline (effective May 26, 2026) Version 3.2 and redline (effective April 29, 2026) Version 3.1 and redline (effective April 2, 2026) Version 3.0 (effective February 24, 2026) Version 2.2 and redline (effective May 14, 2025) Version 2.1 (effective March 31, 2025) Version 2.0 (effective October 15, 2024) Version 1.0 (effective September 19, 2023) ## Ris
Anthropic's Responsible Scaling Policy (version 3.0)anthropic.com · supportingdirection simply because we are unable to achieve them. By establishing this expectation, we hope to create a forcing function for work that would otherwise be challenging to appropriately prioritize and resource, as it requires collaboration (and in some cases sacrifices) from multiple parts of the company and can be at cross-purposes with immediate competitive and commercial priorities. Publishing our Frontier Safety Roadmap may also help inform broader industry and policy discussions on AI safety. Our current Frontier Safety Roadmap is available at anthropic.com/responsible-scaling-policy/roadmap. We will also keep past Roadmaps available at that link. 3. Risk Reports We will publish Risk Reports discussing the risks of our systems and how we have made determinations about whether to [...] 30 days of determining that we have an internally deployed model that is in-scope (per the description above), we will publish a discussion (in a System Card or elsewhere) of how that model’s cap
Anthropic's Transparency Hubanthropic.com · supportingOur Responsible Scaling Policy (RSP) evaluation process is designed to systematically assess our models' capabilities in domains of potential catastrophic risk before releasing them. Under our Responsible Scaling Policy, we regularly publish comprehensive Risk Reports addressing the safety profile of our models. And if we release a model that is “significantly more capable” than those discussed in the prior Risk Report, we must “publish a discussion (in our System Card or elsewhere) of how that model’s capabilities and propensities affect or change analysis in the Risk Report.” Claude Opus 4.7 is significantly more capable than Claude Opus 4.6, the most capable model discussed in our most recent Risk Report. Despite these improved capabilities, our overall conclusion is that catastrophic
Anthropic on X: "As part of our Responsible Scaling Policy, we ...x.com · supportingLog inSign up ## Post # Anthropic on X: "As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: @AnthropicAI Anthropic @AnthropicAI As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: anthropic.com/aug-2026-risk-… 6:00 PM · Aug 14, 2026958.8KViews 305 @AnthropicAI Anthropic @AnthropicAI [...] 305 @AnthropicAI Anthropic @AnthropicAI As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: anthropic.com/aug-2026-risk-… 6:00 PM · Aug 14, 2026958.8KViews 305 @oraphaelmour