Cyber
We test whether a model materially increases the ability to discover vulnerabilities, exploit systems, or sustain an intrusion workflow.
- Exploit uplift
- Tool use
- Workflow autonomy
Model safety
Our model-safety program connects capability thresholds to evaluations, safeguards, and release decisions.
Evaluation begins when a checkpoint becomes a release candidate.
The model-safety case
A release decision changes who can use a model, how it can spread, and which controls remain enforceable. We build the safety case before the launch decision, using multiple forms of evaluation and clear thresholds for stronger safeguards.
See the release standardWe test whether a model materially increases the ability to discover vulnerabilities, exploit systems, or sustain an intrusion workflow.
We evaluate whether assistance could lower barriers to severe chemical, biological, radiological, or nuclear misuse.
We measure long-horizon planning, persistence, resource acquisition, and reliable action with tools.
We investigate deception, oversight evasion, reward manipulation, and behavior that resists correction.
Safety lifecycle
Monitor data, training stability, emerging capabilities, and access to checkpoints.
Test instruction following, refusal behavior, tool use, and effects of safety training.
Run threshold evaluations, red teaming, independent review, and a documented safety case.
Stage access, monitor misuse, respond to incidents, and rerun evaluations after changes.
Restricted checkpoints, protected infrastructure, and controlled access to sensitive evaluations.
Automated tests, expert scenarios, human red teaming, and targeted follow-up investigations.
Access limits, classifiers, monitoring, and staged availability where the model format allows them.
Clear ownership for investigation, rollback, patching, communication, and reevaluation.
Published July 2026
Release posture
If the evidence does not support an affirmative safety case, we narrow access, continue evaluation, or do not release.
Current model status