Nobody Gets Above a C+: The AI Safety Index, Graded
Table of contents
The Report Card Nobody Wanted
The Future of Life Institute just published the Summer 2026 AI Safety Index. It grades nine frontier AI companies across 37 indicators, reviewed by an independent panel of seven experts including Stuart Russell and David Krueger.
The headline: nobody scores above a C+. Three companies got an F. And the weakest domain across the entire industry is existential safety, where no company manages better than a C-.
This is not a fringe activist report. It is a structured, methodological assessment with domain-specific grading, company surveys, and external benchmarks. Let me walk through what stands out.
The Grades
| Company | Overall | Trend |
|---|---|---|
| Anthropic | C+ (2.66) | Stable |
| OpenAI | C (2.28) | Down from C+ |
| Google DeepMind | C (2.01) | Stable |
| Meta | D+ (1.32) | Up from D |
| Z.ai | D- (0.88) | Down from D |
| Alibaba Cloud | D- (0.87) | Stable |
| xAI | F (0.65) | Down from D |
| DeepSeek | F (0.47) | Down from D |
| Mistral | F (0.33) | New entry |
Anthropic leads in five of six domains. OpenAI now leads in Risk Assessment specifically, with broader external testing partnerships. Meta climbed from 6th to 4th. xAI dropped from 4th to 7th.
But the real story is not who is first. It is that first place is a C+.
Three F's Across Three Continents
The three failing companies come from three different regions: xAI (US), DeepSeek (China), and Mistral (Europe). This matters because it kills the "safety is a Western concern" narrative. Inadequate safety is a global problem.
The European dissonance is particularly striking. The EU leads the world in AI regulation. Its top AI company, Mistral, scored dead last on safety. Mistral's leadership has consistently downplayed frontier risk rather than articulating any control or alignment strategy. The company did not even submit a survey response.
xAI's drop is dramatic. It went from 4th place to 7th in one cycle. The panel found "gaping holes" in its evaluations, including no AI R&D data and weak elicitation, against what they described as a "regressing Grok 4 model with concerning propensities." xAI also failed to submit a survey for the first time.
DeepSeek has no published safety framework, no whistleblower policy, and no governance structure. Its rating largely reflects the Chinese regulatory environment rather than any independent safety leadership.
The Military AI Pivot
Reviewers flagged what they call an industry-wide pivot to military AI. Between 2024 and 2026, companies that previously banned military applications gradually reversed course.
Anthropic, OpenAI, Google DeepMind, and Meta all moved from explicit military bans to active defense partnerships. They join xAI and Mistral, who were already seeking defense contracts. The panel criticized Anthropic specifically for "questionable military engagements," citing a reported link to the Minab school strike that caused mass civilian deaths.
Meanwhile, Chinese firms face US allegations of military ties that Alibaba Cloud and Z.ai deny.
The shift is significant because military applications represent some of the highest-stakes deployments of AI systems. And the companies making this pivot are the same ones publishing safety frameworks.
Safety Rhetoric vs. Revealed Behavior
This is the thread that runs through the entire report. Companies say one thing and do another.
Across Google DeepMind, OpenAI, and xAI, the panel found that "leadership's reassuring public messaging diverges from commercial conduct and legislative stance, making stated commitments an unreliable proxy for actual safety practice."
The most concrete example: pause commitments. Anthropic, OpenAI, Google DeepMind, and Meta have all weakened or voided pledges to pause development if red lines are approached. Some now cite competitor-contingent conditions, meaning they will only pause if others pause first. The panel calls this "moving goalposts" and argues it has "undermined safety frameworks across the board."
Companies are publishing and updating safety frameworks as US and EU compliance deadlines approach. But these frameworks lack quantitative thresholds, genuinely independent audits, and clear decision authority. They have weak teeth.
Existential Safety: The Industry's Weakest Domain
| Company | Existential Safety Grade |
|---|---|
| Anthropic | D+ |
| OpenAI | D+ |
| Google DeepMind | D |
| Meta | F |
| Z.ai | F |
| Alibaba Cloud | F |
| xAI | F |
| DeepSeek | F |
| Mistral | F |
No company exceeds C- in existential safety. Six of nine companies received an F.
Constructive attempts exist. Anthropic's constitutional classifiers, OpenAI's call for governance institutions, Google DeepMind's monitoring commitments, Meta's loss-of-control provisions. But panelists judged these to be "entirely inadequate."
The dominant paradigms, interpretability and chain-of-thought monitoring, are questioned because "detection is not prevention." Finding that a model is misaligned after deployment is not the same as preventing misalignment.
The Domain Breakdown
Looking across all six domains reveals where companies actually invest versus where they just talk.
Information Sharing is the strongest domain. Anthropic scores B+, OpenAI and Google DeepMind score B-. This is the easiest domain to score well on because it mostly requires publishing documents and responding to surveys.
Current Harms benchmarking shows Anthropic at B-, with OpenAI and Google DeepMind at C. The panel notes this domain understates real-world harm because it relies on standardized benchmarks rather than measuring downstream effects like AI-associated self-harm, erosion of cognitive independence, or environmental costs.
Governance and Accountability separates the leaders from the pack. Anthropic scores B, OpenAI C, Google DeepMind C-. Everyone else is D+ or below. Meta's whistleblowing protections are undermined by active enforcement of non-disparagement agreements.
Risk Assessment and Safety Frameworks show the top three clustered at C to C+ range. The panel identified three near-universal gaps: no human uplift trials for flagship models, no genuinely independent audits, and limited disclosure of evaluator independence.
What This Means for Builders
If you are building on top of these models, the safety index tells you something important: do not outsource your safety posture to your model provider.
The best company in the world at AI safety got a C+. That is not a passing grade you want to inherit. The frameworks these companies publish are aspirational documents, not engineering guarantees.
This does not mean you should not use frontier models. It means your own guardrails, evaluation pipelines, and deployment constraints matter more than you might think. The model provider's safety framework is a starting point, not a ceiling.
The full report is available at futureoflife.org.
Methodology Notes
For those who care about rigor: the index uses 37 indicators across six domains, graded by a seven-person independent review panel. Companies were selected based on Arena leaderboard rankings (April 11, 2026), sustained frontier investment, and regional representation. Evidence collection covered information through June 3, 2026. Five of nine companies submitted survey responses (Anthropic, Google DeepMind, Meta, OpenAI, Z.ai). Grading uses the US GPA system (A+ = 4.3 down to F = 0).
The panel acknowledges limitations: they rely primarily on public information, cannot verify company claims, and their benchmarks may understate real-world harm. But the structured methodology and expert review make this one of the more credible comparative assessments available.