Frontier labs still publish little about how they would contain a rogue model
Guidelight AI Standards graded public containment plans from OpenAI, Google, Anthropic, Meta, and xAI and found that few specify what they would shut off. OpenAI scored highest at 3 out of 5; Anthropic and Meta scored lowest on published plans.

On August 22 TechCrunch reported a Guidelight AI Standards review of public containment plans from OpenAI, Google, Anthropic, Meta, and xAI. Few of those plans specify what the lab would shut off if a model went rogue.
What we know
- Guidelight AI Standards graded public containment plans from OpenAI, Google, Anthropic, Meta, and xAI.
- Few plans specify what the lab would shut off.
- OpenAI scored highest at 3 out of 5; Anthropic and Meta scored lowest on published plans.
- Steven Adler, a former OpenAI safety researcher, is involved in the work.
- OpenAI said it has paused workloads or taken a model offline.
Takeaways
- Public containment write-ups still rarely name the kill switch.
- On Guidelight’s published-plan score, OpenAI led at 3/5 and Anthropic and Meta sat at the bottom.
- OpenAI says it has already paused workloads or taken a model offline.
Source: TechCrunch


