Skip to content
Wednesday, October 7, 2026
iInnovate MagSTARTUPS · INNOVATION · GADGETS · AI
AI

How AI Safety Works: Risk Frameworks, Capability Thresholds, and Regulators

AI safety is the discipline of preventing harm from AI systems through measurement, mitigation, and governance, and it runs on two tracks. One is public: NIST published its AI Risk Management Framework in January 2023 as a voluntary standard. The other is internal: Google DeepMind's Frontier…

Mei-Ling Chen · May 4, 2026 · 7 min read
ShareXFacebookLinkedInTelegramEmail
A policy researcher's hands sorting printed evaluation charts across a graphite desk, one softly glowing teal terminal window reflected in the monitor behind.
A policy researcher's hands sorting printed evaluation charts across a graphite desk, one softly glowing teal terminal window reflected in the monitor behind.

AI safety is the discipline of preventing harm from AI systems through measurement, mitigation, and governance, and it runs on two tracks. One is public: NIST published its AI Risk Management Framework in January 2023 as a voluntary standard. The other is internal: Google DeepMind's Frontier Safety Framework, published May 17, 2024, sets thresholds triggering mitigations before a model ships.

What is AI safety, as a field?

AI safety is the engineering and policy practice of anticipating and limiting harms from AI systems: bias in decisions, misuse for fraud or weapons development, privacy leakage, unreliable outputs presented as fact, and, at the frontier, loss of control over increasingly capable systems. Its input is measurement — evaluations that probe what a model can and cannot do — and its output is mitigation: training choices, access controls, usage policies, and deployment limits.

The field distinguishes near-term harms, which exist in shipped products today, from frontier risks, which concern capabilities that do not exist yet but plausibly could. The distinction matters because the tooling differs: near-term harms are addressed with audits, red-teaming, and documentation, while frontier risks require evaluating models for capabilities their own developers hope never emerge. Both tracks share a methodological core — treat capability claims as testable, and treat test results as the basis for decisions.

AI safety is also distinct from AI ethics in scope and from AI security in mechanism, though the three overlap heavily in practice. Security asks whether a system can be attacked; safety asks whether it causes harm functioning as intended; ethics asks whether its intended function is acceptable. A single incident — a jailbroken model producing harmful content — can sit in all three categories at once.

How does the NIST AI Risk Management Framework work?

NIST's AI Risk Management Framework, released January 26, 2023, is the most widely referenced public document in the field. Per the institute's own description, it is intended for voluntary use and to improve the ability to incorporate trustworthiness considerations into the design, development, use, and evaluation of AI products, services, and systems, and it was built through a consensus-driven, open process of public comments, drafts, and workshops.

The framework is organized around four functions — Govern, Map, Measure, and Manage — which read as a management loop rather than a compliance checklist. Govern assigns accountability and risk tolerance. Map catalogs the system's context: what it does, for whom, under what conditions. Measure applies quantitative and qualitative evaluation to the mapped risks. Manage allocates resources to the highest-priority risks and monitors the results. The structure deliberately mirrors older NIST cybersecurity and privacy frameworks, so organizations can bolt AI risk onto existing governance rather than invent a parallel apparatus.

Its voluntary status is the framework's main limit and its main strength. Nothing forces adoption, and audits against it vary widely in rigor. But precisely because it is not a regulation, it has been picked up across jurisdictions as a common vocabulary — including by organizations that otherwise answer to entirely different legal regimes.

How do labs set capability thresholds?

Frontier developers publish internal frameworks that define which model capabilities would trigger which mitigations. Google DeepMind's Frontier Safety Framework, published May 17, 2024, is a representative example: the company describes it as a set of protocols for proactively identifying future AI capabilities that could cause severe harm, built around Critical Capability Levels in high-risk domains such as autonomy, biosecurity, cybersecurity, and machine learning research.

The mechanism is a tripwire ladder. Models are periodically run through early warning evaluations; when one approaches a defined Critical Capability Level in a high-risk domain, the framework requires mitigations — tightened security around the model's weights, restricted deployment, or both — before it proceeds. The thresholds are set below the level of severe harm deliberately, so that mitigations arrive with margin rather than in reaction.

Published thresholds make lab claims checkable in a way marketing statements are not: a framework either enumerates its dangerous-capability domains and its evaluation schedule, or it does not. They also make the limits visible. The evaluations probe specific capabilities on specific benchmarks, and a model that scores safely on all of them may still fail in ways no evaluation covered — a gap the labs themselves acknowledge in the framing of these documents.

What can regulators actually do?

Regulators have three real levers, all now in use somewhere. Product and sector law applies existing rules — consumer protection, medical device regulation, financial supervision — to AI systems within their remit, without needing new AI-specific statutes. Procurement and standards leverage operates through requirements like the U.S. federal government's, which direct agencies to manage AI risk along NIST-framework lines for systems they buy and deploy. Dedicated statutes, of which the EU's AI Act is the leading example, classify systems by risk tier and attach obligations and penalties to each tier.

Each lever has a characteristic failure. Sector law arrives piecemeal and misses cross-cutting harms. Procurement rules bind only government business. Comprehensive statutes take years to pass and longer to implement, and they risk hardening around a snapshot of the technology. This is why the voluntary and lab-internal tracks matter in the meantime: they are the only instruments whose update cycle matches the technology's.

What does AI safety not cover yet?

Three gaps recur across every framework published so far. Evaluation coverage: tests exist for a bounded list of capabilities, and absence of evidence on an untested capability is not evidence of absence. Open weights: once model weights are public, access-control mitigations — the strongest tool in most frameworks — no longer apply, and governance shifts entirely to downstream use. Measurement validity: benchmarks can be trained toward, gamed, or simply misread, so threshold decisions inherit the benchmark's blind spots.

A reader evaluating any safety claim can apply the field's own method to it in four steps. Identify what was evaluated, by name. Identify who ran the evaluation and who published the result. Check whether the claimed mitigation is verifiable — restricted deployment can be observed; a training-data improvement generally cannot. And note what the claim conspicuously does not assert, which in safety documents is usually the most informative sentence.

How do the public and private tracks connect?

The two tracks are converging on shared vocabulary rather than shared enforcement. NIST's framework supplies the management grammar — govern, map, measure, manage — that organizations use to structure their AI risk work regardless of what they build. Lab frameworks supply the measurement content: named dangerous-capability domains, evaluation procedures, and thresholds. A regulator drafting obligations can, and in practice does, borrow from both: duties framed as risk-management processes from the former, capability-tiered requirements from the latter.

Procurement is where the connection becomes concrete. When a government agency requires AI suppliers to document risk management along NIST-framework lines, the voluntary framework acquires contractual force for that market without any new statute. The same dynamic operates downstream of the labs: when a model's deployment restrictions are published, enterprise buyers inherit them as conditions of use, and a private threshold becomes a de facto product requirement.

The result is a layered system no single actor controls: measurement standards from public institutes, capability thresholds from developers, binding rules arriving piecemeal from legislatures and buyers. For the foreseeable future, AI safety governance will look less like a single regulator and more like this stack — with each layer's limits, discussed above, setting the agenda for the next one.

What should a non-expert actually take away?

That safety claims in AI are checkable statements, not vibes, and the checking is the same skill as reading any technical claim. The strongest documents in the field share a shape: they name domains, name evaluations, state thresholds, and define mitigations that an outsider could in principle observe. The weakest share the opposite shape — assurances of care with no named test, no stated threshold, and no failure mode acknowledged.

The second takeaway is that the field's honesty about its own gaps is unusually high, and worth taking at face value. The published frameworks say plainly that evaluations are partial, that open weights break access control, and that benchmarks mislead. A reader who internalizes those three caveats knows most of what distinguishes an informed observer of AI safety from a consumer of its press releases.

Sources

  1. AI Risk Management Framework (AI RMF) — National Institute of Standards and Technology
  2. Introducing the Frontier Safety Framework — Google DeepMind

More from our brands

Part of the VUGA Network