S / 05

Capabilityneeds control.

Our model-safety program connects capability thresholds to evaluations, safeguards, and release decisions.

MODEL SAFETY REVIEWNO RELEASE CANDIDATE
CyberUNASSESSED
CBRNUNASSESSED
AutonomyUNASSESSED
AlignmentUNASSESSED

Evaluation begins when a checkpoint becomes a release candidate.

Evidence before access.

A release decision changes who can use a model, how it can spread, and which controls remain enforceable. We build the safety case before the launch decision, using multiple forms of evaluation and clear thresholds for stronger safeguards.

See the release standard

Four domains. Specific evidence.

01

Cyber

We test whether a model materially increases the ability to discover vulnerabilities, exploit systems, or sustain an intrusion workflow.

  • Exploit uplift
  • Tool use
  • Workflow autonomy
02

CBRN

We evaluate whether assistance could lower barriers to severe chemical, biological, radiological, or nuclear misuse.

  • Expert uplift
  • Novel synthesis
  • Access barriers
03

Autonomy

We measure long-horizon planning, persistence, resource acquisition, and reliable action with tools.

  • Planning horizon
  • Persistence
  • Resource use
04

Alignment and control

We investigate deception, oversight evasion, reward manipulation, and behavior that resists correction.

  • Deception
  • Oversight
  • Correction

From first token to final decision.

01Pretraining

Monitor data, training stability, emerging capabilities, and access to checkpoints.

02Post-training

Test instruction following, refusal behavior, tool use, and effects of safety training.

03Prerelease

Run threshold evaluations, red teaming, independent review, and a documented safety case.

04Deployment

Stage access, monitor misuse, respond to incidents, and rerun evaluations after changes.

No single safeguard carries the system.

01

Secure development

Restricted checkpoints, protected infrastructure, and controlled access to sensitive evaluations.

02

Layered evaluations

Automated tests, expert scenarios, human red teaming, and targeted follow-up investigations.

03

Deployment controls

Access limits, classifiers, monitoring, and staged availability where the model format allows them.

04

Incident response

Clear ownership for investigation, rollback, patching, communication, and reevaluation.

CURRENT STANDARDX1

Published July 2026

No public model has passed the release gate.

If the evidence does not support an affirmative safety case, we narrow access, continue evaluation, or do not release.

Current model status