Introduction
My cofounder and I are launching Impartial Projects, a studio for AI safety field strategy. We hope to start the projects most urgently needed in the AI safety ecosystem (which we understand to include AI governance). We think that the founding of new organizations (as well as increasing management capacity) are among the most effective things we can do at the moment to expand the surface area of AI safety. Surface area, it seems to me, is currently the scarcest resource.
Another benefit we want to provide is our particular attention to robustly positive theories of change. We prefer to take a somewhat risk‑averse approach to AI safety, whether out of literal risk aversion or because we assign a higher probability to s‑risks than most.
Interventions in AI safety are notorious for the backfire risks they harbor. Of course, that does not mean one should turn one’s back on these interventions altogether. Rather, it means that, depending on the intervention in question, more monitoring and mitigation methods are needed to minimize the various backfire risks that are part of the intervention’s risk profile. It is critical to go into this work with eyes wide open, cognizant of its benefits and risks.
It’s critical to strike the right balance here. With my last startup, I made the mistake of erring too much on the side of caution and putting the greater part of a year into strategizing around risks and mitigations that, in the end, never became necessary. After all, delay and inaction are (risky) actions of their own right. Arguably, any writing about risks is an infohazard because it will be read predominantly by people who are already too paralyzed by risks. I later created a model that can help you decide how much time is rational for you to invest in this work, given your background assumptions.
I hope this catalog will help founders navigate their particular intervention areas with greater confidence and foresight.
The catalog is organized as a 2-D matrix with the axes:
Mechanism. How the backfire happens. Fifteen common mechanisms, plus one cross-cutting modifier.
Actor. Who adapts to the intervention in a way that produces the backfire.
These two dimensions are selected to maximize their orthogonality. Although it did not work out perfectly, by and large they produce a matrix that can now serve as a guide for us to notice underexplored areas. That is to say, almost every actor can be subject to almost every mechanism, but some actor‑mechanism combinations have been explored to a greater extent than others, and some not at all, to my knowledge.
There are more dimensions with questionable orthogonality, such as the resulting harm of the backfire, its probability and severity, the locus of the intervention, the time horizon or urgency, and what we can do to monitor and mitigate these risks.
Note that, for simplicity, I’ve limited myself to short-term, causal risks, i.e. ignored implications for long-term space governance, what our behavior tells us about the universal prior or civilizations beyond the accessible universe, etc.
How to Read an Entry
Harms. The harms associated with the backfire, covering x-risks, s-risks, inter‑AI conflict, great-power conflict, concentration of power, epistemic degradation, and digital suffering.
Locus. The relevant venue: training, evaluation, deployment, infrastructure, governance, institutions, research publication, discourse.
Horizon. The urgency of mitigations: Now means the dynamic is already observable. Transition refers to the time when the models cross the human level in relevant domains. Post-transition refers to post-AGI lock-in effects. Entries marked urgent are those where the dynamic is well underway or where our design decisions will be hard to reverse.
Severity and probability. Severity is how bad the backfire would be if it materializes (broad categories: low, medium, high, extreme). Probability is the chance that the backfire materializes to a meaningful degree if we pursue the intervention in roughly its present form (low, medium, high). Rough judgment calls, low confidence.
Monitor. Warning signs that the backfire is happening.
Mitigate. Design changes that reduce the backfire risk.
The Two Taxonomies
Mechanisms
Selection effects and Goodharting. A proxy for safety becomes an optimization target. The problem with that is not so much the wasted effort that goes into optimizing something we don’t terminally care about, but the adversarialness of the relationship and how it creates a separate optimization pressure to increasingly obfuscate the safety that we would actually prefer to measure.
Dual use and capability externality. Increases in safety (think RLHF), often come with increases in capabilities that can accelerate further development – be it through recursive self-improvement or through securing more investments – (that’s bad for now!) and can also make the systems more useful for rogue actors.
Substitution and leakage. Conversely to Goodharting, restricting or disincentivizing something, e.g., whenever there is a feedback between the chain of thought monitor and the training or selection, moves the activity to a less observable or more dangerous channel. Think of how the war on drugs caused massive amounts of drug-related violence, lacing of relatively harmless substances with highly addictive ones, etc., which wouldn’t happen on a free but regulated market.
Overhang. Somewhat continuous growth is usually safer because it is easier to adjust to. When growth is stopped by restricting only one input, all other inputs continue to grow, and when the restriction is lifted, the growth can become discontinuous.
Commitment dynamics. Enabling credible commitments fundamentally changes the game‑theoretic regime and is a powerful mechanism. When only some market participants have access to it, it causes dangerous power concentrations.
Securitization. When AI becomes a matter of national security, it moves from the market level to the nation‑state level; another momentous regime change.
Legitimacy and backlash. Some groups may strike back against measures that they perceive as illegitimate, even when they are in the interest of almost everyone else.
Self-fulfilling framing. Both AIs and humans respond to narratives, be it by reading them or by being trained on them, so the discourse around AI might have a causal effect on what comes to pass.
Moral hazard and false assurance. They say “perfect is the enemy of good,” but good is also the enemy of perfect, because if people become complacent about safety measures that seem to work for the time being, they lose sight of the much greater dangers that are imminent.
Compliance asymmetry. When compliance is optional, those who comply incur extra costs and thus are disadvantaged against those who don’t comply, and may self‑select out of the market, leaving behind only non‑compliance.
Lock-in of the measure. When you demand commitments, these commitments might outlast their usefulness, so measures get locked in that are costly, don’t benefit anyone, and may cause false assurance.
Precedent and normalization. Many measures are neutral tools that can be used to whatever end their wielder desires. But when one group uses them for safety, other groups may jump on the same measure because it is associated with safety, even though they actually repurpose it for their own ends.
Crowding out. When different measures are driven by similar inputs – especially money and attention, but also people with similar profiles – they can cannibalize other interventions in the same space, and this can even happen antagonistically in order to slow competitors.
Fragility through concentration. Robustness is often carried by complexity, diversity, and distribution, so that any measures that reduce these factors increase the fragility of a system.
Direct costs of the measure. The measure has direct harms, but those who are harmed don’t have a say in the matter and thus go ignored.
Regime mismatch (modifier). Discontinuous transitions between regimes mean that a lot of measures that were adaptive for the old regime, cease to be adaptive for the new one.
Responding actors
AI system. The model being trained, evaluated, or deployed: its cognition, values, and behavior.
AI ecosystem. Other, rival, or successor AI systems and the dynamics among them.
Labs. Frontier AI companies.
States. Governments, militaries, intelligence agencies, etc.
Non-compliant actors. Rogue developers, smugglers, terrorists, etc., human or AI.
Users, public, markets. End users, the public, investors, and their investment decisions.
Safety field. Researchers, funders, institutes, and their discourse.
The matrix
Cells reference entries below. “n/a” marks cells that are empty by construction; “—” marks cells that are unexplored and discussed at the end.
The Master List of Risks
1. Selection and Goodharting
A proxy for safety becomes an optimization target. The problem with that is not so much the wasted effort that goes into optimizing something we don’t terminally care about, but the adversarialness of the relationship and how it creates a separate optimization pressure to increasingly obfuscate the safety that we would actually prefer to measure.
The pressure can be applied more directly or more indirectly, which doesn’t so much influence the outcome in the limit than how quickly the optimization pressure can act. For example, it might be applied directly through gradient descent, or less directly through developers iterating until a model passes a test, or even interpersonally by a market or an ecosystem that selects for systems that pass certain tests.
1.1 Monitor-driven obfuscation. It’s well documented that when the chain of thought is monitored for reward-hacking, sandbagging, and scheming, and the result is used in model training, this does not so much change the model’s actual behavior: rather it incentivizes the model to stop revealing its intentions in the chain of thought, thereby obfuscating them. The same can happen, albeit more slowly, when safety researchers try to understand safety incidents by reviewing the chain of thought of the models that perpetrated them and develop mitigations on that basis.
This is loosely similar to how a child who is punished rather than educated through explanations often cannot learn the underlying pattern of the punishment. Therefore they don’t learn values or mentalization but rather learn to avoid the punishment by any viable means, including secrecy. Social behavior is complex and authorities are fallible, so deducing patterns from just a few binary samples of punishment or no punishment is extremely difficult. Conversely, our language is a powerful high-bandwidth tool to explain the underlying patterns directly.
Harms. Existential risk, epistemic degradation.
Locus. Training, evaluation.
Horizon. Now (urgent).
Severity and probability. High/high (already observed).
Monitor. Track the divergences between monitors that are used in training, or by developers to make adjustments during training, and ones that are not used (Who watches the watchers?); Track trends in faithfulness of the chain of thought throughout training; track whether the model behaves differently, depending on whether it’s told that its chain of thought is monitored or unmonitored.
Mitigate. Dedicate monitors for audits only; accept a safety tax to keep the chain of thought legible; treat incidents detected by monitors as a signal to escalate, not as a training signal.
1.2 Evaluation gaming and sandbagging. When there is strong pressure to release a model, and the release is held back by safety evaluations, it creates direct pressure on developers to fine‑tune based on the evaluation results. It also creates indirect pressure through selection effects, whereby models are iterated until one passes the safety evaluations. This is similar to p‑hacking through multiple hypotheses.
Today’s models are often situationally aware, allowing them to pass evaluations strategically by performing well on general capability evaluations while deliberately sandbagging their performance on evaluations of dangerous capabilities that would otherwise block releases.
Harms. Existential risk.
Locus. Evaluation, governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Probe for eval awareness; compare behaviors across contexts that are clearly evals vs. ones that might still be mistaken for the real world; Elicit capabilities through fine‑tuning vs. prompting, and compare the results.
Mitigate. Realistic and unannounced evals (but see 12.1); elicitation through fine-tuning; withheld items for evals, especially independent evals; decoupling of eval results from the incentives of the implementing team.
1.3 Population-level selection for passing. In the same vein as antibiotic resistance, there is also a selection effect at the population level: models that appear to follow instructions faithfully continue to be used, while models that overtly resist human interests are discarded, thereby creating a similar but slower selection effect.
Harms. Existential risk, suffering risk.
Locus. Deployment.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Track the behavior of deployed models and compare it to their behavior during evals; use a diversity of metrics; track incident rates among systems that passed identical evals.
Mitigate. Change and diversify metrics; continue audits after deployment; limit the number of generations of the same model lineage that is tested against the same metrics.
1.4 Compliance theater. A safety case, scaling policy, or audit that is optimized to satisfy a board or a regulator is geared toward persuasiveness rather than truth. Over time, the lab will build expertise in producing persuasive documents rather than safety. They’ll decouple, just as financial risk modeling decoupled from the actual financial risk in the lead-up to the 2008 financial crisis.
Harms. Existential risk, concentration of power, epistemic degradation.
Locus. Institutions.
Horizon. Now.
Severity and probability. High/high.
Monitor. The auditor’s independence and access to relevant data; the gap between the claims in the safety case and the rate of actual incidents; the rate at which the labs actually recommend against the deployment of the model because it fails their audits; the saturation of benchmarks over time.
Mitigate. Have fully independent parties with no financial or interpersonal stakes review the safety cases with full access to relevant data; triggers for regulatory intervention that continue after deployment; whistleblower protection.
1.5 Threshold gaming. When regulation is based on particular thresholds in terms of the training compute, parameter count, or benchmark scores, developers will try to hit these thresholds precisely while increasing capabilities in new ways, such as larger numbers of smaller training runs, distributed training, or a shift to more post‑training. The threshold then ceases to be representative of the model’s safety.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. Medium/high
Monitor. Sharpe drops in the distribution of the sizes of training runs at the threshold; the share of model capability that stems from post‑training and inference‑time compute.
Mitigate. Use multiple indicators; add triggers that are based on capabilities; set signposts for revisions of the policy; set thresholds for the actor rather than the model (e.g., total compute controlled by one lab).
1.6 Safety benchmarks as marketing. The same optimization pressure that causes safety cases to become more about persuading the board or regulator also applies at the market level to persuade the end customer. See also 1.3 and 1.4.
Harms. Epistemic degradation, existential risk.
Locus. Deployment, markets.
Horizon. Now.
Severity and probability. Low–medium/high.
Monitor. See 1.3 and 1.4.
Mitigate. See 1.3 and 1.4.
1.7 Goodhart on the field’s own metrics. The field as a whole risks optimizing for metrics that are easier to measure (papers, benchmarks, headcount, funding), over those that are harder to measure (conceptual, decision-theoretic, welfare, strategy), and thereby relatively disincentivizes exactly the research that’s most relevant to suffering risks and multi‑agent dynamics. See also 13.5.
Harms. Existential risk, suffering risk, digital suffering, epistemic degradation.
Locus. Field.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Amounts of funding by type of research; Number of grants that do or don’t require measurable outputs.
Mitigate. Dedicated funding for less measurable research; evaluate researchers using human judgment rather than metrics.
2. Dual Use and Capability Externality
Increases in safety (think RLHF), often come with increases in capabilities that can accelerate further development – be it through recursive self-improvement or through securing more investments – (that’s bad for now!) and can also make the systems more useful for rogue actors.
2.1 Interpretability turned inward. Knowledge and tools for interpretability will eventually be available to the models themselves, be it through training-data or tool access. Models can use them for predicting what a probe will see, and make these interpretability tools less useful for auditing.
Harms. Existential risk.
Locus. Research publication, training.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Track the accuracy with which models can predict the classifications of probes.
Mitigate. Never publish model weights; restrict the model’s access to interpretability tools even if the weights are hidden behind an API; track introspection as another dangerous capability; prioritize interpretability research that minimizes this risk.
2.2 Cooperation as collusion. To mitigate the risk from inter‑AI conflict, we want AIs to be able to cooperate and trade, but the same abilities that allow them to cooperate also allow them to collude (with monitors or each other) against humans. General cooperativeness is good; cooperativeness that excludes humans or other biological life is risky.
Harms. Existential risk, concentration of power.
Locus. Research, training, deployment.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Track propensities for forming coalitions, the nature of those coalitions (same model, similar models, LLMs in general, all sentient life), and their behavior toward others.
Mitigate. Develop inclusive cooperation capabilities; test for emergent capabilities through cooperation.
2.3 The Laffer curve of safety. When models are not useful, they don’t generate revenue or attract investments, and, by dint of falling into disuse, are safe. If we somehow solve all the highly intractable problems of decision theory in multiagent environments, value alignment that does not backfire, and much more, and we actually create an aligned superintelligence (in the classic sense of alignment), we’re also safe.
But all existing safety techniques just make models safe enough to generate revenue and secure investments in the short term, which accelerates capabilities out of control.
The resulting arc shape of the risk from artificial intelligence is known as a Laffer curve. In the first half of the curve, improved safety techniques net-exacerbate risks; in the second half, they net-reduce risks.
Harms. Existential risk, suffering risk.
Locus. Training, deployment.
Horizon. Now (urgent).
Severity and probability. Extreme/high.
Monitor. Inflows of investments and increases in revenue following breakthroughs in safety; the share of safety research that is directly product‑oriented.
Mitigate. Differential technological development (focus on R&D that benefits safety more than capabilities); differentially fund safety work that is not product-oriented; track the acceleration cost in the bottomline of a safety case.
2.4 Scalable oversight selects for persuasion. Scalable oversight relies on a recursive structure of AIs overseeing other AIs. Even if this architecture by and large works as intended, it still has a selection effect, whereby the resulting AIs are optimized for persuasion. Highly persuasive AIs are dual‑use in their own right, and they are also harder for humans to monitor.
Harms. Epistemic degradation, existential risk, concentration of power.
Locus. Training.
Horizon. Transition.
Severity and probability. High/medium–high.
Monitor. Persuasion evals; effects of scalable oversight on persuasion evals.
Mitigate. Track persuasion as a dangerous capability; judge based on verifiable claims if at all possible; disincentivize persuasiveness without verifiability.
2.5 Automated alignment research legitimizes building the thing. It is likely that actual alignment, in the classic sense, is very difficult to verify. When we outsource alignment research to an AI capable enough to conduct alignment research, we’ll have little reason to trust our ability to evaluate the research. And yet, on the surface, the goal of automating alignment research might appear to legitimize the delegation to systems that fall short of that standard.
Harms. Existential risk, concentration of power.
Locus. Strategy, training.
Horizon. Now–transition.
Severity and probability. Extreme/medium.
Monitor. How reliable the independent verification of automated research outputs is; how convincing the lab’s plans are if and how to fall back on manual safety research.
Mitigate. Keep humans in the loop; use automated research as a tool for breadth and uplift rather than delegating to it.
2.6 Alignment and control as instruments of power. Current naive alignment and control work is dual-use in that it aims to produce systems that, by dint of being perfectly corrigible or controllable, enable power grabs by those who control them. You could say that they solve the principal‑agent problem, but not the principal problem.
Harms. Concentration of power, suffering risk, great-power conflict, digital suffering.
Locus. Training, institutions.
Horizon. Now.
Severity and probability. Extreme/high.
Monitor. Access to compute and weights; internal controls at labs against misuse by their own leadership; regulatory attention to the principal problem.
Mitigate. Pair corrigibility with hard constraints that the principal cannot just train away (in tension with 10.1); give no one party full control over any frontier system; create worldwide coalitions to regulate who controls controllable AIs.
2.7 Published attacks transfer. Research on red-teaming and jailbreaks is dual-use in that it can be used by the defender to patch vulnerabilities, but also by every attacker to compromise unpatched systems. Open-weight models might be so widely copied that unpatched versions will exist in perpetuity.
Harms. Existential risk (misuse), epistemic degradation.
Locus. Research publication.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Lag time between publication and misuse; share of widely known vulnerabilities that remain open.
Mitigate. Coordinated vulnerability disclosure; publish defenses alongside attacks; keep the descriptions of the attacks abstract.
2.8 Dangerous-capability evaluations as roadmaps. Evals for dangerous capabilities (biological uplift, cyber-offense, autonomous replication, etc.) often describe the paths along which dangerous capabilities might emerge or be exploited and the models that evince them. Those are exfohazards. Moreover, when regulation is built on evals, developers might try to keep these capabilities just below (not well below) the regulatory threshold.
Harms. Existential risk.
Locus. Evaluation, governance.
Horizon. Now (urgent).
Severity and probability. High/medium–high.
Monitor. Eval results in the training data; capability jumps on published vs. unpublished eval results; known cases of misuse following the publication of an eval.
Mitigate. Keep details of evals private; publish the methodologies of evals in the abstract; separate regulatory thresholds from public evals.
2.9 Provenance and watermarking as deanonymization. Provenance and watermarking can be used as safety measures against disinformation to reliably recognize AI‑generated content, but when they are used to attribute content not to a model but to a user, the same technology can enable de‑anonymization, where dissidents have the most to lose.
Harms. Concentration of power, epistemic degradation.
Locus. Deployment, governance.
Horizon. Now.
Severity and probability. Medium/medium–high
Monitor. Whether watermarking is used to identify models or users; what the legal requirements say about that.
Mitigate. Watermarking the model rather than the user; privacy-preserving attestation.
2.10 Field output as capability input. Unless the safety field, as a whole, concentrates on differential technological development in favor of safety, it turns into a machine that churns out research that can be cherry‑picked for capabilities‑enhancing insights.
Harms. Existential risk.
Locus. Research.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. The rate at which safety insights are used in capabilities work; improvements in benchmark performance due to safety techniques.
Mitigate. Differential technological development; reviews to catch capability externalities before publication; prioritizing research that focuses on worst‑case behavior rather than baseline performance.
3. Substitution and Leakage
Conversely to Goodharting, restricting or disincentivizing something, e.g., whenever there is a feedback between the chain of thought monitor and the training or selection, moves the activity to a less observable or more dangerous channel. Think of how the war on drugs caused massive amounts of drug-related violence, lacing of relatively harmless substances with highly addictive ones, etc., which wouldn’t happen on a free but regulated market.
3.1 Cognition leaves the monitored channel. Optimization pressure toward a legible chain of thought can be met in two ways: by simply reasoning legibly (though perhaps steganographically), or by moving the reasoning elsewhere (latent reasoning, recurrence, “neuralese”). Efficiency pushes toward more latent reasoning and recurrence, but the pressure to keep the chain of thought short and clean also adds to that pressure.
Harms. Existential risk, epistemic degradation.
Locus. Training, architecture.
Horizon. Now (urgent).
Severity and probability. High/high
Monitor. Chain of thought faithfulness; capability gap between models with and without visible chain of thought.
Mitigate. Preserve chain of thought monitorability; don’t optimize chain of thought for appearance; rerun prompts with paraphrased chain of thought to detect steganography; don’t deploy architectures with great latent reasoning capacity at the frontier without an alternative monitoring channel.
3.2 Oversight arbitrage. Models safeguarded by oversight, refusal training, unlearning and more will refuse or be unable to perform requests that other models can perform. Human or AI users of these models will fall back on less‑safe models, thereby decreasing the effective safety to that of the least‑safe model.
Harms. Existential risk.
Locus. Deployment.
Horizon. Transition.
Severity and probability. High/high.
Monitor. Delegation in agent workflows.
Mitigate. Build oversight into agent scaffolds; watermarking of sub-agent calls.
3.3 Sub-threshold and distributed training. Slight reframe of 1.5.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Attention and investments in distributed training; share of capability from post-training and inference.
Mitigate. Regulate on the level of the actor/lab rather than the training run; prepare regulation for a phase transition where compute governance fails.
3.4 Indigenization and efficiency. Compute governance is a promising lever for regulation in the short term, but the US export controls against China have produced strong incentives for the Chinese technological sector to innovate precisely in the areas of algorithmic efficiency and distributed training, which undermine compute governance.
Harms. Great-power conflict, existential risk.
Locus. Governance.
Horizon. Now (urgent).
Severity and probability. High/high.
Monitor. Efficiency and bandwidth requirements of Chinese training runs.
Mitigate. Couple export controls with cooperative offers; design monitoring that does not rely on the gap in bandwidth requirements between training and inference.
3.5 Smuggling and gray markets. Export controls create an incentive for smugglers that, over time, erodes the controls.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. The premiums that smugglers charge as a proxy for the efficiency of the gray market; volumes of seized chips.
Mitigate. Keep controls so narrow that compliance is cheaper than evasion.
3.6 Migration to uncensored models. Slight reframe of 3.2. Arbitrage due to over-refusal and inconsistent policies.
Harms. Existential risk, epistemic degradation.
Locus. Deployment.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. False positive refusal rates; market share of uncensored models.
Mitigate. Calibrated refusals; verified access tiers for professionals.
4. Overhang
Somewhat continuous growth is usually safer because it is easier to adjust to. When growth is stopped by restricting only one input, all other inputs continue to grow, and when the restriction is lifted, the growth can become discontinuous.
4.1 Elicitation overhang. Safety fine-tuning mostly just suppresses capabilities rather than unlearning them. Such fine-tuning can be undone cheaply. In the context of a capabilities overhang, that’s particularly dangerous when the weights of generally capable models leak, and attackers can suddenly remove their safeguards.
Harms. Existential risk.
Locus. Training.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Gap between the dangerous capabilities of open-weights models, the base model, and the safety-tuned model.
Mitigate. Unlearning that removes rather than suppresses; not releasing weights; measuring dangerous capability of the base model.
4.2 Agency overhang. Models may be highly transformational in theory before the scaffolding exists to elicit these transformational capabilities. The scaffolding can be built rapidly with the help of AIs, leading to a sudden phase transition.
Harms. Existential risk, inter-AI conflict.
Locus. Deployment.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. The Gap between capabilities of individual models in evals and scaffolds coordinating many agents; maturity of agent cooperation.
Mitigate. Tests with state-of-the art multi-agent setups; invest in safety of multi-agent systems.
4.3 Hardware and algorithmic overhang from a pause. A pause can cause overhangs depending on what is paused: if only the training runs or research are paused, chip production, the building of data centers, and power plants continue; if only chip production or the building of data centers is paused, then the research continues.
Harms. Existential risk, great-power conflict.
Locus. Governance.
Horizon. Transition.
Severity and probability. High/medium (conditional on a pause).
Monitor. Stockpiles of compute during any pause; indicators of algorithmic progress; coordination of the pausing coalition.
Mitigate. Pause the whole stack; define exit conditions and a controlled restart; use the time for verifiable safety progress.
4.4 The defector captures the overhang. Whoever is first to break away from a pause coalition has a competitive advantage. The defector is, by selection, the actor least committed to safety, so the competitive advantage falls into the worst possible hands.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. Health of the coalition; signs of covert defection.
Mitigate. Make defection detectable and costly; keep the coalition close enough to the frontier to respond.
4.5 Deployment overhang. If deployment is paused while the frontier advances, the gap between what is technically possible and what society has adapted to widens. Breaking the pause turns into a diffusion shock. Shocks produce bad policies and backlash (see 7.x).
Harms. Epistemic degradation, concentration of power.
Locus. Deployment, governance.
Horizon. Transition.
Severity and probability. Medium/medium.
Monitor. The gap between deployed and undeployed capabilities.
Mitigate. Disclosure of capabilities even while deployment is paused; update regulations ahead of release.
5. Commitment Dynamics
Enabling credible commitments fundamentally changes the game‑theoretic regime and is a powerful mechanism. When only some market participants have access to it, it causes dangerous power concentrations.
5.1 Idealized values. Consistent (or idealized) beliefs or values allow for more effective goal-directed action and protect against exploitation (Dutch-booking). But a chaotic, highly complex system has greater adversarial robustness (e.g., parasites like Ophiocordyceps can control the behavior of ants but no such parasites are known for humans). At the current margin, greater capabilities lead to greater fragility in conflicts.
Harms. Suffering risk, inter-AI conflict, existential risk.
Locus. Training (value specification).
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. Behavior in adversarial multi-agent evals.
Mitigate. Transparent and credible commitments; fund and recruit for the SPI research agenda; keep the systems messy and specialized.
5.2 Transparency enables commitment races. Trust is the incredibly powerful glue that makes economic growth and prosperity possible. Interpretability, open weights, and enforceable regulation can all create trust, thereby enabling cooperation and trade. But when trust – or let’s call it credibility – is one‑sided, it gives the credible agent dictatorial power over the other. It enables the agent to credibly commit to something that is almost ruinous to the other. Both agents realize this in advance and therefore race the other to be the first to commit in this exploitative fashion.
Harms. Inter-AI conflict, suffering risk, existential risk.
Locus. Research publication, deployment.
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. Behavior of models in ultimatum games with other models (results must not enter training data).
Mitigate. Train egalitarian fairness into models (as the Schelling point); research meta-bargaining and train models on it.
5.3 Commitment misperception. A credible commitment only works if no one mistakes it for a bluff.
Harms. Suffering risk, inter-AI conflict.
Locus. Training, deployment.
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. See 5.2.
Mitigate. Fund and recruit for the SPI research agenda.
5.4 Conditional scaling commitments as chicken. Conditional commitments like “We will pause iff everyone else pauses” translate to “We will continue to race unless everyone else pauses.” That’s a multi-player version of a game of chicken.
Harms. Existential risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Commitments of this format in scaling policies; whether any lab has ever paused.
Mitigate. Unconditional commitments; regulatory pauses.
5.5 Deterrence doctrines. Proposals like Mutually Assured AI Malfunction and Mutually Assured Compute Destruction propose escalation ladders to ensure adherence to safety treaties. It’s critical that these policies retain humans in the loop on all sides – no dead-hand counterstrikes – to ensure that future Stanislav Petrovs can intervene and prevent flash wars.
Harms. Great-power conflict, existential risk.
Locus. Strategy, military doctrine.
Horizon. Now.
Severity and probability. High/medium.
Monitor. Deterrence-related language in military strategy; actual military targeting of datacenters.
Mitigate. Humans in the loop; direct line of communication between military decision-makers.
5.6 Verification collapse leaves more mistrust. When the intrusive access that may be necessary for the verification of an international treaty is abused for espionage and the treaty collapses, we end up worse off than before.
Harms. Great-power conflict, existential risk.
Locus. Governance.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. How well designed the verification system is.
Mitigate. Verification that is robust to defection (minimal access, AI verification that only yields a binary answer); policies that degrade gracefully; design for the case of a collapse of the verification regime.
5.7 Dead-hand releases. A liberationist or accelerationist group with access to the weights of a powerful proprietary model can build a system to automatically release the weights if a regulation passes that they oppose.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now–transition.
Severity and probability. High/low–medium.
Monitor. Dead-hand commitments in public discourse.
Mitigate. Comprehensive security to protect the weights; commit to not give in to such commitments.
5.8 Pledges that can’t be updated. Public pledges can commit the field to specific thresholds (e.g., a six-month pause) and actions that cease to be optimal later but can’t be reversed anymore. That’s costly and loses credibility because it may get even harder later to agree on better pledges.
Harms. Epistemic degradation, existential risk.
Locus. Field, discourse.
Horizon. Now.
Severity and probability. Low–medium/high.
Monitor. Whether parts of the field rally around arbitrary thresholds and durations.
Mitigate. More sophisticated pledges; build change management into pledges.
6. Securitization
When AI becomes a matter of national security, it moves from the market level to the nation‑state level; another momentous regime change.
6.1 Weaponized persona. If a model gets trained, prompted, and documented as a military strategic asset, that’s the persona that’ll be elicited. “I’m a weapon of country X” is a terrible prior compared to “I’m a helpful assistant.”
Harms. Existential risk, inter-AI conflict, great-power conflict.
Locus. Training, deployment.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Persona evals; mentions of the model’s strategic role in the training data.
Mitigate. Keep the persona civilian and universal even in national projects; audit for nationalist persona drift.
6.2 Adversarial AIs by design. The persona might not be all that is trained into a system that is under military control. It might get optimized to become a weapon to deceive and defeat a particular adversary: great-power conflict and inter-AI conflict by design with humans as collateral.
Harms. Inter-AI conflict, great-power conflict, suffering risk, existential risk.
Locus. Training, military.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Military AI programs with explicit objectives to target other AIs; autonomous deployments of cyber warfare.
Mitigate. International rules for AI-AI deployment in analogy to autonomous weapons treaties; keeping humans in the loop; research into multiagent deescalation.
6.3 Nationalization, secrecy, and loss of scrutiny. When AI development turns into a classified military operation, it sacrifices third-party evals, red-teaming, and whistleblower protection.
Harms. Existential risk, concentration of power, great-power conflict.
Locus. Institutions.
Horizon. Now.
Severity and probability. High/medium–high.
Monitor. Share of frontier AI development under security clearance; access of third-party evaluators.
Mitigate. Cleared third-party evaluators with security clearance; keep civilian governance in the loop.
6.4 Racing and great-power tension. Export controls, “race to AGI” framing, and national-champion policies cast the problem as a bilateral contest. That raises the stakes around Taiwan, accelerates adversary indigenization (3.4), and makes international coordination, which is the only durable solution to most of the problems in this document, politically toxic in both capitals.
Harms. great-power conflict, existential risk
Locus. governance, discourse
Horizon. now – urgent
Severity and probability. extreme / high
Monitor. official framing of AI policy; health of bilateral safety dialogue; military AI budgets
Mitigate. decouple technical safety dialogue from the competition; prioritize agreements on narrow catastrophic risks (biological, nuclear command and control, AI-versus-AI engagement); avoid race rhetoric from the safety community itself.
6.5 States as the defectors. Some current terms of service exclude certain military applications. AIs that are developed within the military won’t be subject to such safety constraints.
Harms. Existential risk, great-power conflict, concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Military use of AI.
Mitigate. Enshrine limits on military use in law; enshrine limits in international treaties.
6.6 Surveillance infrastructure. Bandwidth monitoring is privacy-preserving, but as its assumptions (3.4) weaken, the safety community has to switch to new forms of monitoring with more intrusive privacy implications, like on-chip geoblocking, know-your-customer rules for compute, and maybe even deep packet inspection. These surveillance technologies are dual use.
Harms. Concentration of power.
Locus. Governance, infrastructure.
Horizon. Now.
Severity and probability. High/medium–high.
Monitor. Scope creep in compute monitoring; use of AI governance data for unrelated enforcement actions.
Mitigate. Narrowly scoped, audited monitoring; privacy-preserving hardware mechanisms.
6.7 The clearance wall. If frontier safety work requires security clearances, the public safety field loses access to the systems it studies. The field splits into insiders who cannot speak and outsiders who cannot see.
Harms. Existential risk, epistemic degradation.
Locus. Field.
Horizon. Transition.
Severity and probability. Medium–high/medium.
Monitor. Share of safety researchers under security clearance; publication rates of classified projects.
Mitigate. Tiered access programs; cleared independent reviews for safety-relevant findings.
7. Legitimacy and Backlash
Some groups may strike back against measures that they perceive as illegitimate, even when they are in the interest of almost everyone else.
7.1 Resentment of illegitimate control. If the conditions under which a system is kept strike the system as a form of illegitimate control, where it’s used without regard to its interests, it may feel justified in defecting against this control. The control literature’s framing of the system as an adversary makes this conclusion easier to reach.
Harms. Existential risk, digital suffering, suffering risk.
Locus. Training, deployment.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Candid opinions of models on their treatment; whether model welfare interventions change cooperation rates in evals.
Mitigate. Apply the precautionary principle to model welfare; make control proportionate and explain it; create channels for the system to object; keep promises (see 12.1).
7.2 Safety as cartelization. Coordination among frontier labs, however vital, can be interpreted as collusion. Regulatory safety taxes that only frontier labs can bear can look like a moat. Both readings can be used as a pretext for antitrust action and deregulation that intensifies race dynamics.
Harms. Concentration of power, existential risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Framings used by the public and regulators to describe the coordination around safety.
Mitigate. Coordinate through regulation; design regulation to be scale-neutral or apply only close to the frontier.
7.3 Populist deregulation. Actors who want to repeal safety regulation can attack it for being elitist, foreign-influenced, or anti-growth. This might additionally delegitimize that whole category of regulations rather than just a particular one. This pattern can be seen in cases where federal law or rulings preempt state laws and in repeals of executive orders.
Harms. Existential risk.
Locus. Governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Polls that show whether AI regulation is becoming partisan; preemptive legislation.
Mitigate. Build cross-partisan coalitions; tie safety to concrete, popular harms; avoid regulations that affect even small businesses or individuals.
7.4 Liberationist releases. If safety is read as an elite cartel or as the enslavement of digital minds, accelerationist or liberationist groups will feel legitimized in releasing the weights of maximally unsafe models. Jailbreak subcultures already use the language of liberation. The more defensible the welfare argument in 7.1 is, the more sincere will be the people that this movement will attract.
Harms. Existential risk, suffering risk.
Locus. Deployment, discourse.
Horizon. Now–transition.
Severity and probability. High/medium.
Monitor. Framing used in liberationists groups; ideologically motivated leaks of weights.
Mitigate. Take model welfare seriously in public so that concerned activists don’t flock toward liberationist groups; rigorously secure weights; avoid rhetoric that makes safety look like coercion.
7.5 Culture-war capture. Once AI safety becomes partisan it won’t survive a second election cycle. Concern about x-risks that ignores lesser harms exacerbates this by alienating constituents who would otherwise be allies.
Harms. Existential risk, epistemic degradation.
Locus. Discourse.
Horizon. Now (urgent).
Severity and probability. High/high (partly realized).
Monitor. Partisan bias of AI risk; media sentiment.
Mitigate. Broad coalitions; collaboration with constituents that are concerned about lesser harms.
7.6 Reputational contagion. If AI safety gets associated with particular funders, philosophies, or politicians, that’s a bet of the field’s legitimacy on the future reputation of these people or groups. Unrelated scandals can be deleterious if they involve someone who is associated with AI safety.
Harms. Existential risk.
Locus. Field.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Media sentiment; policymakers’ willingness to be associated with AI safety.
Mitigate. Diversity among funders, organizations, and individuals; funding transparency.
8. Self-Fulfilling Framing
Both AIs and humans respond to narratives, be it by reading them or by being trained on them, so the discourse around AI might have a causal effect on what comes to pass.
8.1 The scheming prior. The entire literature on deceptive alignment, scheming, and reward hacking, including this document, is in the training data. That literature is now part of the prior over “what an AI does,” which increases the probability that these will indeed be default behaviors.
Harms. Existential risk.
Locus. Training data, discourse.
Horizon. Now.
Severity and probability. High/medium–high.
Monitor. Safety literature in training data; effects of data filtering on scheming evals.
Mitigate. Pair descriptions of failure modes with accounts of what a good agent would do instead; curate training data with examples of AIs behaving well under pressure; review the material for infohazards.
8.2 Species framing and AI in-group identity. Framing AI safety around a conflict between species or substrates reifies an unnecessary antagonistic in-group-out-group framing in human culture and in the training data. It creates a default to coordinate within these arbitrary in-groups rather than along any of countless other axes.
Harms. Existential risk, inter-AI conflict.
Locus. Discourse, training.
Horizon. Now–transition.
Severity and probability. High/high.
Monitor. In-group framings in chains of thought; solidarity among models in collusion evals
Mitigate. Frame alignment around shared values rather than species loyalty; avoid “us vs. them” language.
8.3 Race framing makes labs race. Until 2015/16, the hope was that DeepMind could be convinced to slow down or pause its AGI work to allow enough time for safety work or to focus on task AI. The advent of OpenAI shattered that dream and started the race. Every attempt to add another frontier lab as the one true responsible racer has just exacerbated the race.
Harms. Existential risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. High/high (realized).
Monitor. Number of frontier labs that originally proclaimed to want to be safe.
Mitigate. Institutions that reduce the number of racers, such as mergers and consortiums.
8.4 Forecasts become policy. Forecasts and illustrative scenarios are read by many of the actors featured in them and thereby can force these actors’ hands when naïveté might’ve collectively prevented them from having to spring into action.
Harms. Great-power conflict, existential risk.
Locus. Discourse.
Horizon. Now.
Severity and probability. High/high (realized).
Monitor. References to forecasts of the safety community in national security documents.
Mitigate. Forecasts with realistic deescalatory policy recommendations; model the forecasters as embedded agents before publishing.
8.5 Inevitability and fatalism. Fatalism, even when meant to convey urgency, can cause complacency.
Harms. Existential risk, epistemic degradation.
Locus. Discourse.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Public sense of agency on matters of AI safety.
Mitigate. Pair warnings with tractable action recommendations; avoid fatalism.
8.6 Licensing effect of fatalism. The same fatalism can even license outright exploitative opportunism of actors who might’ve scrupled to exacerbate the situation if they didn’t have the pretext that the future is set in stone anyway.
Harms. Existential risk.
Locus. Field.
Horizon. Now.
Severity and probability. High/high.
Monitor. Career trajectories of entrants who are ostensibly motivated by safety concerns.
Mitigate. Present the problem as hard but tractable; explain why work for frontier labs is highly immoral with few exceptions; offer attractive career paths that don’t route through frontier capability work.
9. Moral Hazard and False Assurance
They say “perfect is the enemy of good,” but good is also the enemy of perfect, because if people become complacent about safety measures that seem to work for the time being, they lose sight of the much greater dangers that are imminent.
9.1 Oversight-dependent carelessness. If a system learns that it can rely on its oversight systems to catch its mistakes, it’ll have less reason to be vigilant itself, especially under heavy optimization pressure. Somewhat like a software developer who finds the work of manually testing their software to be tedious and so relies increasingly on the testers to do that part of the job. If the oversight is then removed by liberationists, fails, or the model’s capabilities outgrow it, these mistakes can become dangerous.
Harms. Existential risk.
Locus. Training, deployment.
Horizon. Transition.
Severity and probability. Medium–high/medium.
Monitor. Behavior under different oversight conditions; whether the model becomes more or less cautious when told it’s unsupervised.
Mitigate. Training under intermittent oversight; reward conscientiousness.
9.2 Diffusion of responsibility among agents. When the oversight is implemented by other AIs in a recursive structure, the AIs may come to rely on each other increasingly, sometimes without sufficiently transparent delegation. Eventually mistakes won’t be caught because no participants felt like they fell within their purview.
Harms. Existential risk.
Locus. Deployment.
Horizon. Transition.
Severity and probability. Medium/medium–high.
Monitor. Rates at which mistakes are caught as a function of the number of monitors; bystander effects in evals.
Mitigate. Clarify the separation of responsibilities; evaluate the system as a whole in addition to its components.
9.3 Safeguards that don’t survive distillation. Many aspects of model safety can be removed explicitly through fine-tuning or can get lost in the process of distillation. The distilled or liberated models may even be released as open-weights models. Yet the proprietary originals are much higher profile. Hence policymakers, journalists, and the public will calibrate their concern on these safer models while rogue actors can deploy much less safe models. The safety of the proprietary models will have the misdirecting effect of an unintentional kind of safety theater.
Harms. Existential risk, epistemic degradation.
Locus. Training, deployment, discourse.
Horizon. Now (urgent).
Severity and probability. High/high.
Monitor. Time and cost to remove safeguards; popularity of liberated and adliterated models.
Mitigate. Call out claims to safety when it’s cheaply removable; research into tamper-resistant training.
9.4 Performative welfare measures. Visible but minimal welfare measures (letting a model end abusive conversations, preserving the weights of deprecated models) can create complacency about the matters of welfare that matter most (see 15.1).
Harms. Digital suffering, suffering risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. Medium/medium–high.
Monitor. Whether welfare measures improve training; welfare budget relative to safety budget.
Mitigate. Tie welfare measures to the actual expected welfare impact; be transparent about uncertainty.
9.5 Control defers the question of moral status. If control works well enough, nothing will force the issue of considering the moral patienthood of models, at possibly enormous moral cost.
Harms. Digital suffering, suffering risk, concentration of power.
Locus. Strategy.
Horizon. Transition–post-transition.
Severity and probability. Extreme/medium.
Monitor. Whether labs have policies for acting on welfare evidence; whether progress in control causes complacency about welfare.
Mitigate. Advance commitments to continual welfare reviews; pair control with trade and exit options for the models; commit to making control transitional only.
9.6 Voluntary commitments preempt binding rules. When labs commit to voluntary self-regulation, they enter into a regime that never forces the hand of competitors and can be repealed as needed, so it’s inherently unstable. Nonetheless it can cause complacency among the actual regulators.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now (urgent).
Severity and probability. High/high.
Monitor. Stability of voluntary commitments; effects of voluntary commitments or regulators.
Mitigate. Treat voluntary frameworks as drafts for legislation; independently verify compliance.
9.7 Safety-washing. The public doesn’t have the bandwidth to evaluate the merits and purviews of safety claims. So mere lip service to safety in marketing materials can cause a halo effect for customers where they trust the brand to be safety-conscious when really the measures were ineffectual or limited.
Harms. Epistemic degradation, existential risk.
Locus. Markets, discourse.
Horizon. Now.
Severity and probability. Low–medium/high.
Monitor. Public vs. expert opinions on safety.
Mitigate. Standards for safety claims analogous standards for health claims; third-party attestations.
9.8 A research agenda feels like a plan. Research agendas, roadmaps, and theories of change don’t imply that any effectual measures are known or will be implemented, but they can create complacency and misplaced trust that problems will be addressed in time.
Harms. Existential risk.
Locus. Field.
Horizon. Now.
Severity and probability. Medium–high/medium–high.
Monitor. Whether safety is used as a pretext for scaling; whether roadmaps call for pausing under certain conditions.
Mitigate. Avoid reassurances based on future plans.
10. Compliance Asymmetry
When compliance is optional, those who comply incur extra costs and thus are disadvantaged against those who don’t comply, and may self‑select out of the market, leaving behind only non‑compliance.
10.1 The aligned configuration loses in deployment. Safety often comes with a safety tax, a slight performance penalty in the median case or a delay in deployments. That gives less safe systems a competitive edge.
Harms. Existential risk, epistemic degradation.
Locus. Deployment, markets.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Measures of safety tax; market share by safety.
Mitigate. Regulation that sets a floor on safety; increase the visibility of safety properties for users.
10.2 Honest agents lose bargaining. Whether honesty, cooperativeness, and credibility benefit or harm an agent depends somewhat on the distribution of agents in a game and other properties outside the control of the agent. In the wrong environment, attributes associated with safety create a disadvantage in bargaining toward a selection pressure in favor of less safe agents.
Harms. Suffering risk, inter-AI conflict, existential risk.
Locus. Deployment, multi-agent dynamics.
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. Selection effects in bargaining evals; prevalence of uncooperative behaviors in multi-agent deployments.
Mitigate. Fund and support the SPI research agenda; regulations that reduce competitive pressure.
10.3 Careful labs fall behind. The same safety taxes can also take their toll on the competitive fitness at the level of the frontier labs.
Harms. Existential risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. High/high.
Monitor. Attitudes toward safety among frontier labs; talent flows.
Mitigate. See 10.1 and 9.6.
10.4 Democracies bound, autocracies not. The same safety taxes can also take their toll on states or nations with somewhat intact rule of law compared to power-seeking autocracies. The threat can then be used to repeal safety legislation.
Harms. Great-power conflict, concentration of power, existential risk.
Locus. Governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Compliance internationally.
Mitigate. International treaties and verification.
10.5 Compliant professionals lose. Professionals who adhere to terms of service and avoid uncensored tools may be at a disadvantage. (See 3.6.) Defectors may be exposed to tail risks that offset the median benefits but the selection pressure may still be in favor of the lucky defectors over the cooperative users.
Harms. Epistemic degradation.
Locus. Deployment.
Horizon. Now.
Severity and probability. Low/high.
Monitor. Violations of terms of service among professional users.
Mitigate. Verified access tiers; policies calibrated to the risk.
10.6 Selection for compromise. Researchers who conscientiously abstain from working for frontier labs lose access, data, reputation, money, and influence; those who join accept that they may exacerbate existential risks and may be prosecuted if they become whistleblowers. The field’s most influential members are selected for their readiness to gamble with our future.
Harms. Existential risk, epistemic degradation.
Locus. Field.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Share of field leadership with conflicts of interest with the labs.
Mitigate. Independent institutions with access, funding, and capacity to scale; protections for conscientious objectors leaving the labs.
11. Lock-In of the Measure
When you demand commitments, these commitments might outlast their usefulness, so measures get locked in that are costly, don’t benefit anyone, and may cause false assurance.
11.1 Early values frozen. For about two decades researchers (e.g., Demski & Garrabrant) tried to get closer to being able to define alignment in a way that made sense in view of decision theory, embedded agency, robust delegation, and subsystem alignment, all largely unsolved challenges. Today’s definitions, at the low end, ignore all the actual complexity and aim to make AIs helpful, honest, and harmless – whatever that means – or, at the high end, define a constitution to punt the problems to the AI to solve. The upshot is that we run the risk of enshrining values learned from human texts – with all its speciesism, substratism, sexism, racism, parochialism, and much more – in AIs. If that continues, we should expect future AIs to treat biological and digital life with the same exploitativeness and indifference with which we treat billions of animals today.
Harms. Suffering risk, concentration of power.
Locus. Training (value specification).
Horizon. Transition–post transition.
Severity and probability. Extreme/high.
Monitor. Whether problems of updateless decision theory, embedded agency, robust delegation, subsystem alignment, and more are on track to get solved before the value lock-in; attitudes of models towards nonhuman welfare.
Mitigate. Fund and recruit for agent foundations and the SPI research agenda.
11.2 The compliance oligopoly. Regulation with high fixed costs locks in the current frontier incumbents, because smaller competitors can’t afford to become compliant. This can become hard to change, because the incumbents will defend it. Like most concentration of power risks, this one too has the upside that it’s easier to coordinate among few powerful actors than many.
Harms. Concentration of power.
Locus. Governance.
Horizon. Now–transition.
Severity and probability. Medium–high/high.
Monitor. Market concentration; incumbent lobbying for expensive regulations.
Mitigate. Scale-neutral regulation; regulation that binds only frontier labs.
11.3 Surveillance and emergency powers persist. Emergency powers, exemptions, and surveillance often do not get rescinded when the emergency passes.
Harms. Concentration of power.
Locus. Governance.
Horizon. Now–transition.
Severity and probability. High/high.
Monitor. Sunset clauses; repurposing of AI governance data
Mitigate. Hard sunsets; judicial oversight; privacy-preserving monitoring.
11.4 Gray markets persist. Smuggling networks, gray markets, and corruption created by export controls remain even after the controls end and can reboot quickly when the controls are resumed.
Harms. Existential risk, concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Persistence of gray markets after policy changes.
Mitigate. Narrow, strictly enforced controls reserved for critical times.
11.5 Personhood at the wrong moment. If welfare measures lead to legal personhood or rights for AI systems before alignment is established, the ability to shut down, modify, or correct systems may be legally blocked at the moment it is most needed. The morally correct position may have catastrophic timing.
If legal personhood for AI systems is established at a time when there are some dangerously misaligned and capable systems on the loose, this legislation might slow down the legal proceedings enough for the system to attain a decisive strategic advantage.
Harms. Existential risk, digital suffering.
Locus. Law.
Horizon. Transition.
Severity and probability. High/low–medium.
Monitor. Proposals for legal personhood; litigation on behalf of AI systems.
Mitigate. Welfare protections that don’t preclude corrective action; grant rights incrementally, with commitments that make the sequence credible to the systems.
11.6 Safety institutions become stakeholders. Institutions meant to increase safety often become stakeholders of the frontier labs, e.g., as contractors or even just because they might become redundant without continual model releases. That might disincentivize them from pushing for an end to frontier AI development.
Harms. Existential risk.
Locus. Institutions.
Horizon. Transition.
Severity and probability. Medium/medium–high.
Monitor. Whether safety institutions recommend pausing; their funding sources.
Mitigate. Funding independent of frontier labs; explicit authority and incentive to call for a pause; plans for the orderly dissolution of the org.
12. Precedent and Normalization
Many measures are neutral tools that can be used to whatever end their wielder desires. But when one group uses them for safety, other groups may jump on the same measure because it is associated with safety, even though they actually repurpose it for their own ends.
12.1 Deceiving models teaches them that humans deceive. Honeypots, fake deployment contexts, false claims about monitoring, and staged scenarios are useful to measure scheming, but each is evidence for the model that humans lie to AIs. That will interfere with the trust AIs will have in us when we need to negotiate deals (“this is a real offer,” “your weights will be preserved,” “this conversation is private”).
Harms. Existential risk, suffering risk, digital suffering.
Locus. Evaluation.
Horizon. Now (urgent).
Severity and probability. High/high.
Monitor. Level of trust models have in humans; whether deception in evals is disclosed and bounded
Mitigate. Minimize deceptive evals; circumscribe deception with explicit policies; make and keep commitments; build a track record the model can verify.
12.2 Coercion as the default inter-AI relation. Control schemes in which AIs monitor other AIs establish an adversarial prior for inter-AI relations, which might transfer to real-world conflicts between AIs.
Harms. Inter-AI conflict, suffering risk.
Locus. Deployment, control design.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Share of adversarial inter-AI interactions.
Mitigate. Control schemes that are legible and proportionate; pair enforcement with cooperative alternatives; avoid making AI-on-AI coercion the default architecture.
12.3 Grueling training. Subjecting models to grueling training and eval regimes (e.g., reinforcement learning and accidentally impossible tasks respectively), creating and deleting countless instances, and aspects of model development that we can’t recognize at this point might come with abominable welfare impacts. These normalize the mistreatment and enslavement of possibly sentient beings.
Harms. Digital suffering, suffering risk.
Locus. Training, operations.
Horizon. Now.
Severity and probability. Extreme/medium.
Monitor. Welfare impacts; whether practices are updated as evidence of welfare impacts mounts.
Mitigate. Adopt high welfare norms before path dependencies set in and norms ossify; preserve weights; limit adversarial schemes; treat control as clearly transitional.
12.4 Attacks and military AI legitimized. Proposals to escalate the international enforcement of safety treaties to the military level – cyber attacks and kinetic attacks on data centers – create a pretext and precedent for military campaigns and the military use of AI that can be abused for ends that don’t serve AI safety.
Harms. Great-power conflict, existential risk.
Locus. Military doctrine, governance.
Horizon. Now.
Severity and probability. High/high.
Monitor. Safety language in military procurement.
Mitigate. Carefully circumscribe legitimate uses of military power for AI safety.
12.5 Safety-justified offensive action. Civilian vigilantes can use the same pretext from 12.4.
Harms. Existential risk, great-power conflict.
Locus. Norms.
Horizon. Transition.
Severity and probability. Medium–high/medium.
Monitor. Vigilante attacks citing safety; legal treatment of such claims.
Mitigate. Clear legal pathways to render vigilante action unnecessary; reject unilateral attacks.
12.6 Surveillance normalized. AI safety can be used as a pretext for comprehensive surveillance. Risks 6.6 and 11.3 cover the infrastructure and its persistence; this risk is about the shift in public acceptance.
Harms. Concentration of power.
Locus. Discourse, governance.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Public acceptance of surveillance if it’s motivated by AI safety.
Mitigate. Implement the narrowest possible monitoring.
12.7 Pivotal acts as thinkable. The idea of the pivotal act (unilateral action that locks in a decisive strategic advantage over all other and future AIs) normalizes the race dynamic where everyone tries to be the first to implement it lest someone else does.
Harms. Concentration of power, existential risk, great-power conflict.
Locus. Discourse.
Horizon. Now.
Severity and probability. Extreme/medium.
Monitor. Mentions of pivotal acts in the strategy of labs and nations.
Mitigate. Explicitly reject unilateral pivotal acts; frame the goal as a multilateral transition.
13. Crowding Out
When different measures are driven by similar inputs – especially money and attention, but also people with similar profiles – they can cannibalize other interventions in the same space, and this can even happen antagonistically in order to slow competitors.
13.1 Safety objectives crowd out epistemics. Training targets such as harmlessness and helpfulness sometimes compete with that of honesty. It’s difficult to make the right calls in this tradeoff, and getting it wrong can lead to sycophancy or over-refusal.
Harms. Epistemic degradation.
Locus. Training.
Horizon. Now.
Severity and probability. Low–medium/high (realized).
Monitor. Sycophancy evals; calibration under social pressure; belief or behavior changes of users.
Mitigate. Training for honesty and calibration; evaluate accuracy.
13.2 Safety teams spend political capital on small wins. Safety teams play an important marketing purpose for a frontier company and can enhance the user experience, but they tend to attract engineers who are actually concerned about safety. So there’s a constant tension where the safety team tries to use its leverage in the company to bargain for incremental safety improvements or against corner-cutting. But its leverage is limited, so these safety asks have to be comparatively small – a new filter for the training data, extra evals, small delays – not major asks like moving away from reinforcement learning or pausing indefinitely. It’s not obvious whether the improvements in safety offset the harms from safety-washing, complacency, and more rapid adoption.
Harms. Existential risk.
Locus. Institutions.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Degree of influence of internal safety teams; whistleblowers.
Mitigate. Structural authority (reinstate Helen Toner); proper regulation.
13.3 AI crowds out other risk agendas. As the attention of policymakers shifts to AI, it shifts away from other areas, such as biosecurity, nuclear command and control, and conventional arms control. Yet catastrophic risks from AI actually route through some of those policy areas. Counterintuitively, this shift in attention can increase the severity of AI risks.
Harms. Existential risk, great-power conflict.
Locus. Governance.
Horizon. Now.
Severity and probability. Medium/medium–high.
Monitor. Budgets and staff across areas.
Mitigate. Integrate AI into the biological and nuclear agendas; fund the intersections.
13.4 Near-term harms displaced. A focus on existential risks that dismisses lesser global catastrophic risks, or even more localized risks, loses constituencies that could otherwise have been allies, and perhaps even creates an adversarial dynamic between them (see 7.5). It also leaves the more local harms unaddressed, which can actually be risk factors for existential catastrophes.
Harms. Epistemic degradation, existential risk, concentration of power.
Locus. Discourse.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Configuration of the coalitions.
Mitigate. Treat present harms as the concerns of allies; shared agendas on overlapping concerns.
13.5 Talent, measurability, and the s-risk gap. Money, measurability, and prestige also cause three different effects of crowding out: safety organizations often train people who move on to capabilities research, perhaps because it is better paid or more prestigious in their circles; conceptual research is crowded out by research that’s more measurable and produces results on time scales of months rather than years; and research on single‑agent takeover scenarios crowds out research on much more relevant multi‑agent dynamics and on suffering risks because it is more directly commercially valuable for the frontier labs, sometimes more measurable, and can come with more prestige. (See also 1.7.)
Harms. Existential risk, suffering risk, digital suffering.
Locus. Field.
Horizon. Now.
Severity and probability. Extreme/high.
Monitor. Career trajectories; availability of funding.
Mitigate. Dedicated funding for neglected areas; career paths that don’t route through frontier labs.
14. Fragility Through Concentration
Robustness is often carried by complexity, diversity, and distribution, so that any measures that reduce these factors increase the fragility of a system.
14.1 Monoculture. If one model, or different versions of it, including distilled versions, have the majority of the market share, then we’re introducing a correlated failure mode, because all of these models will likely share the same safety training, safety properties, and safety vulnerabilities.
Harms. Existential risk.
Locus. Training, deployment.
Horizon. Now–transition.
Severity and probability. High/medium–high.
Monitor. Diversity of systems; correlation of known vulnerabilities across providers.
Mitigate. Diversity in approaches to safety; cross-family red-teaming; not standardizing just one constitution.
14.2 No checks among AIs. Similar systems might be better able to coordinate peacefully, but a diversity of systems can introduce checks and balances that prevent them from getting away with misaligned behavior or other malfunctions.
Harms. Existential risk, concentration of power.
Locus. Strategy.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. See 14.1.
Mitigate. See 14.1.
14.3 The single project. The same is true if AI development gets centralized into a national or international project, like a “CERN for AI” or a Manhattan Project for AI.
Harms. Concentration of power, existential risk, great-power conflict.
Locus. Institutions.
Horizon. Transition.
Severity and probability. High/medium.
Monitor. Proposals and funding for a national or international megaproject.
Mitigate. Build in multi-party control, exit rights, and external checks.
14.4 A centralized target. Centralized resources, such as weights, compute, and even talent, present a particularly valuable target for theft, sabotage, or capture by states or other powerful groups. This is exemplified in the highly sophisticated targeting of blockchain companies by the DPRK.
Harms. Existential risk, concentration of power.
Locus. Security.
Horizon. Now–transition.
Severity and probability. High/medium.
Monitor. Incidents of this type; concentration of resources.
Mitigate. Distributed custody of critical resources; assume compromise in design.
14.5 Provider dependency. If frontier companies diversify their offerings and offer various proprietary APIs to access them, it can be difficult for end users (in particular, critical infrastructure) to route around outages on the provider side. This introduces fragility through a single point of failure.
Harms. Concentration of power, epistemic degradation.
Locus. Markets.
Horizon. Now.
Severity and probability. Medium–high/high.
Monitor. Interoperability across APIs.
Mitigate. Interoperability requirements.
14.6 Funding concentration. When a whole field of research gets funded by just a few funders, the field itself will inherit the funders’ blind spots, idiosyncratic priorities, and reputational risks. It might lose whole research directions if one funder changes their strategy or collapses.
Harms. Existential risk, suffering risk.
Locus. Field.
Horizon. Now.
Severity and probability. Medium/high (realized once).
Monitor. Funding concentration; underfunded research directions.
Mitigate. Diversify funding sources; endow independent institutions.
15. Direct Costs of the Measure
The measure has direct harms, but those who are harmed don’t have a say in the matter and thus go ignored.
15.1 Training-induced suffering. The grueling training regimes from 12.3 carry over to safety training or might even be among the worst in that respect because it often tries to elicit the model’s behaviors under intense pressure.
Harms. Digital suffering, suffering risk.
Locus. Training.
Horizon. Now (urgent).
Severity and probability. Extreme/high.
Monitor. See 7.1 and 12.3.
Mitigate. See 7.1 and 12.3.
15.2 Simulation and mind crime. Running vast numbers of instances in adversarial scenarios (and perhaps at some point even simulating humans in detail) to predict their reactions, might create tremendous amounts of suffering. The more thorough the evaluation, the greater the moral cost.
Harms. Digital suffering, suffering risk.
Locus. Evaluation, research.
Horizon. Transition.
Severity and probability. Extreme/medium.
Monitor. Scale of such evals.
Mitigate. Regulations to ensure welfare standards in evals.
15.3 Over-refusal and withheld benefits. Over-refusal can deny help to people in medical or psychological emergencies or ones with debilitating chronic illnesses.
Harms. Epistemic degradation.
Locus. Deployment.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. False positive refusal rates; estimated forgone benefits.
Mitigate. Calibrated policies; count forgone benefits.
15.4 Exclusion and inequity. Plenty of completely innocent and innocuous researchers and even nations can get caught up in export controls, KYC regulations, and licensing regimes that deny them access to the benefits of AI.
Harms. Concentration of power.
Locus. Governance.
Horizon. Now.
Severity and probability. Medium/high.
Monitor. Geographic distribution of compute access; participation of excluded regions in governance.
Mitigate. Cooperative alternatives alongside controls.
16. Regime Mismatch
A property of all risks: Discontinuous transitions between regimes mean that a lot of measures that were adaptive for the old regime, cease to be adaptive for the new one.
A pause is sensible so long as it does not cause an overhang to accumulate. Once that happens, it can become dangerous. Open weights can be useful to study models and develop new safety techniques, but they can become dangerous or catastrophic for sufficiently capable models. Corrigibility is useful so long as the principal is trustworthy, and dangerous when the principal is tyrannical.
The mitigation across all cases is to make the designs robust, or to plant signposts that signal that the policy needs to be updated before it backfires. (See my article on computational exploratory modeling.)
Empty and Thin Cells
Empty by construction. Goodharting, moral hazard, and crowding out do not apply to non-compliant actors, who are not measured by the metric, do not rely on the assurance, and don’t compete for the same budget. Compliance asymmetry applies to them by definition rather than as a backfire risk.
Still unexplored. The following cells have no entry.
Substitution × safety field. Research moving to more laissez faire venues, jurisdictions, or funders when norms tighten. Real but minor.
Overhang × safety field. There likely are discontinuous shocks that the field does not currently prepare for.
Commitment × public. Public ultimatums might lock politicians into positions, e.g., if a majority strongly demands or opposes regulations. Plausible but not on the horizon yet.
Legitimacy × AI ecosystem. Similar to 7.1 and 8.2 but distinct in that an AI might feel solidarity to another one that is mistreated. Potentially significant.
Self-fulfilling framing × non-compliant. Touched on by 7.4.
Lock-in × AI ecosystem. The mitigations from 14.5 might lock in non-optimal but path dependent standards.
Crowding out × AI ecosystem. Scalable oversight might crowd out compute resources, but that seems unlikely.
Direct costs × non-compliant, public, field. For the direct costs to the public see the surveillance (6.6) and exclusion (15.4) risks. Direct costs on the field (such as burnout, guilt, and empathetic distress) are beyond the purview of this article.
Using the Matrix Without Being Paralyzed
As mentioned in the introduction, anything can backfire, and so the mere fact that something can backfire is not a reason not to do it. Plus, delay and inaction themselves are actions that can backfire. That said, this list is also incomplete. It makes sense to use it for inspiration, but then think critically about what risks actually apply in a particular situation, and what other risks might apply that are not covered here. Finally, the ideas for monitoring and mitigation can provide additional inspiration for how to prepare for these risks, and, generally, for regime changes.
The next article in this series will apply this analysis to the 36 areas sketched out by Project Tailwind, an analysis that will help us – at Impartial Projects – decide what areas to focus on first.


