Research

Overview Interpretability Alignment science Economic research

Safety

Commitments Responsible scaling Trust centre

Learn

Academy Developer docs Models News

Company

About Careers Contact
Try Resonance Log in

A policy is only a policy if it can stop a launch.

Safety commitments are easy to write and easy to quietly abandon. These are the ones we have tied to specific capability measurements, with named safeguards and an obligation to publish what happens when a threshold is crossed.

Capability tiers and required safeguards
Tier Trigger Required before proceeding Sign-off
T1 Baseline Standard deployment; no uplift in the evaluations below. System card, usage policy enforcement, abuse monitoring. Model release review
T2 Uplift Measurable uplift on cyber or bio task suites above the published baseline. Targeted refusal evaluations, access controls, third-party red team. Safety lead + release review
T3 Autonomy Sustained multi-day task completion without human checkpoints. Interpretability review, sandboxed tool use, independent audit, incident log made public. Safety lead + board safety committee
T4 Hold Capability where the required safeguard does not yet exist. Training paused or scaled back until a safeguard is demonstrated and reviewed. Board safety committee only
On the word "pause". A hold is a real stop, but it is not permanent by default. It ends when the missing safeguard exists, is reviewed by people who did not build it, and is described publicly in enough detail to be argued with. If we cannot describe it, we have not built it.
Certification SOC 2 Type II Report available under NDA
Privacy ISO 27001 & 27701 Scope covers the API platform
Data residency Region pinning US, EU, UK, and APAC
Retention Zero-retention mode Opt-in, contractually bound
Access SSO, SCIM, audit log On every enterprise plan
Availability Public status page Posted within 30 minutes
05 — Incident log

What went wrong, and when

2026-08

Evaluation misconfiguration inflated a safety score

A contaminated test split raised a refusal benchmark by nine points. Caught by an internal replication check before the model shipped; the suite was rebuilt and the affected claim withdrawn from the system card.

2026-05

Prompt-injection path via retrieved documents

Injected instructions in a retrieved PDF caused a tool call outside the declared scope. Fixed with a stricter tool sandbox and an evaluation added to the release gate.

2026-02

Over-refusal in non-English medical queries

Refusal rates on legitimate clinical questions were materially higher in four languages. An external review traced it to training-data imbalance; corrected in the following checkpoint.

2025-11

Status page delay during a regional outage

Public disclosure took 74 minutes against a 30-minute commitment. Post-mortem published with the timeline and the on-call changes made in response.

06 — Disclosure policy

Reporting a vulnerability

If you have found a way to make our models do something they should not, we want to hear about it before anyone else does. Send the minimum needed to reproduce it, and give us a reasonable window before publishing.

  • In scope: jailbreaks that bypass safety training, prompt-injection paths through tools or retrieval, data leakage between sessions, and evaluation gaming.
  • Out of scope: benign completions you personally dislike, and issues already described in a published system card.
  • Our commitment: acknowledge within two business days, keep you updated, credit you by name if you want it, and tell you what we changed.
We would rather be embarrassed by a report than surprised by one.