AGENTIC COMMONSAI industry briefings

繁中EN

TOPICAI Safety & SecurityPUBLISHED 2026-10-04

All English articlesAI Safety

Safety Gates Before the Training Run, Not After: What Meta's Framework Update Changes

On this page6 sections
  1. The shift: safety requirements now start before the run
  2. What Meta says happens before a high-risk run
  3. Open weights get their own risk treatment
  4. Governance gets a second layer
  5. What this means if you're building
  6. Sources

The shift: safety requirements now start before the run

Most responsible-AI frameworks you read are really release frameworks — they describe what gets tested before a model ships. On October 2, 2026, Meta published an update to its Superintelligence Scaling Framework that moves the checkpoint earlier: loss-of-control requirements now apply during training and evaluation itself.

The reasoning is concrete. Per Meta’s post, a model with strong cybersecurity capabilities could, without safeguards, exploit vulnerabilities in its environment and take unauthorized actions mid-training or mid-eval. That’s a real operational concern if you run reinforcement learning on capable models — the training harness is infrastructure, and infrastructure has attack surface.

What Meta says happens before a high-risk run

Three commitments stand out, all from Meta’s own description:

  • Before any RL run, Meta assesses whether the model’s capabilities could compromise its environment. If yes, mitigations must be in place first.
  • High-risk environments run in sandboxes vetted through a standardized internal process. Sandboxes with high-severity vulnerabilities get fixed or excluded. Environments are also checked for reward-gaming invitations like exploitable grading or impossible tasks.
  • Full trajectories and chain of thought stream to tamper-resistant logging in real time, and automated monitors watch all rollouts of high-risk models and can halt a run on detected misbehavior or escape attempts.

That last point includes a detail I’d flag for anyone building eval infrastructure: Meta says it will stress-test the monitors themselves to establish they’d actually catch unauthorized actions. A monitor you’ve never tested is a dashboard, not a control.

Open weights get their own risk treatment

The second half of the update addresses open-weight releases. Meta acknowledges the obvious: once weights are out, someone can resample past a refusal, prefill outputs, or fine-tune refusal behavior away entirely. The updated Framework requires assessing the model in its specific deployment context rather than treating the released artifact as the end of the story.

For bio and chemical weapons risk, Meta says it evaluates both what capabilities a model enables and whether the deployment itself would contribute to proliferation — with threat modeling that sometimes involves outside experts and government stakeholders. I’d treat that as stated intent rather than verified practice; the blog describes the process, and there’s no independent confirmation in the source material.

Governance gets a second layer

Meta also committed to a new Board-level AI committee that will review future Framework changes and independently check that operations match the stated standards. This tracks commitments Meta says it made at the White House that week — internal monitoring of model capabilities and alignment, checked by an internal independent team plus an external auditor or evaluator.

If you’re tracking how labs structure oversight, this matters more than any single safeguard. A sandbox policy can be quietly loosened; a board committee reviewing the policy is harder to quietly remove. How much real authority the committee has is something only time will show.

What this means if you’re building

You probably aren’t training frontier models, but the pattern transfers. Gate on capability before you give a system tools, write access, or network egress — the same logic behind treating eval-time code execution as a sandboxed provider concern, which I covered in provider-run sandboxes becoming part of the API. Log trajectories immutably. Test your monitors, not just your models. And when you release something others can modify, reason about the modified uses, not just the shipped configuration.

The honest limitation: this is a company describing its own standards. The Framework is designed to change, Meta says, and the proof of these commitments will be in future updates and whether the committee actually pushes back. Worth reading the full post before drawing conclusions — it’s short, and the specifics are more interesting than the summary suggests.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Safety Gates Before the Training Run, Not Just Before Deployment

    Meta extends its scaling framework so containment, logging, and monitoring apply during training — and before weights ship.

    Meta

  2. Meta Halts Teen Access to AI Characters Ahead of Redesign

    On January 23, 2026, Meta said teens will lose access to AI characters across Instagram, Facebook, and WhatsApp until a safer version with parental controls ships. The timing and the fallout.

    Meta

  3. Third-Party AI Assessments: What OpenAI's New Priorities Change for Builders

    OpenAI's assessment priorities define four areas where external scrutiny should focus, and what that means for safety claims.

    AI Safety