Key indicators at a glance
What counts as
evidence here
This is a record, not an argument. It exists because the question of whether advanced AI poses a catastrophic or existential risk is discussed almost entirely in assertions, confident in both directions, while the underlying material is scattered across arXiv preprints, lab system cards, government PDFs and legal texts that almost nobody reads end to end. Every entry below resolves to the thing itself.
Four streams
Foundations is the theoretical case and its history: the arguments for why a sufficiently capable system might behave catastrophically, and the methods proposed to prevent it. Evidence is what happened when people actually measured frontier models: system cards, uplift trials, autonomy and scheming evaluations, including the many that found nothing. Governance is the law, the treaties and the voluntary commitments, with their real current status rather than their announced status. Dissent is the case that this framing is mistaken, argued by its own proponents rather than summarised by its opponents.
The verification rule
An entry earns its place by having a primary source that was retrieved and read: the paper, the card, the official journal text, the signed declaration. 190 of the 200 entries meet that bar. The 10 that do not are still listed, marked source not retrieved, because silently dropping them would misrepresent how much of this literature is cited more often than it is read.
What is deliberately excluded
Predictions presented as findings. Interviews in place of papers, except where the statement itself is the event. Aggregate “experts say” claims without a survey behind them. Any number we could not trace to a stated source, including several widely circulated figures that turned out to be repetitions of repetitions.
The argument is older
than the machines
Nothing in the modern debate is new except the evidence. The structure of the argument was fully stated before electronic computers existed, and restated at every capability inflection since: a machine which improves itself could escape the control of the people who built it, and that would be the last mistake we get to make.
Full bibliographic detail for all 45 Foundations entries, including the alignment-methods line (RLHF, Constitutional AI, debate, weak-to-strong generalisation, AI control) and the 2025–2026 additions, is in the canon.
Six ways it is
supposed to go wrong
“AI risk” is not one claim. It is six distinct mechanisms with different evidence bases, different plausibility, and different policy implications, and conversations founder because people defend one while attacking another. Each card below states the argument, the strongest empirical support for it, and the strongest reason to doubt that it scales to catastrophe.
Specification gaming
Goal misgeneralisation
Deceptive alignment & scheming
Gradual disempowerment
Misuse uplift
The curves the
argument rests on
Both sides of this debate are really arguing about extrapolation. These are the four series that do most of the work, drawn only from figures a named source states in text, with no interpolation or smoothing, and with estimated values marked as estimates.
What the testing
actually found
Between 2023 and 2026 the argument acquired something it had never had: an experimental record. Frontier labs, independent evaluators and two national institutes have now run hundreds of dangerous-capability evaluations. The results are more equivocal than either side’s summary of them, and the most interesting findings are about the evaluations themselves.
Biology: the trials contradict each other
This is the best-funded question in the field and it is not resolved. The largest pre-registered wet-lab randomised trial, with 153 novices working with summer-2025 frontier models, found 5.2% completing the core viral reverse-genetics workflow with AI assistance against 6.6% with internet access alone. That is no uplift. A Los Alamos pilot with ten participants found 60% against 20%, consistent with uplift but far too small to carry weight. Meanwhile an in-silico study of 57 novices found a 4.16× accuracy advantage, with AI-assisted novices beating expert baselines on three of four benchmarks.
The gap between those results is the central unresolved question in AI biosecurity, and it is a gap about ecological validity: knowing the protocol is not the same as executing it at the bench. Earlier studies using 2023-era models (RAND’s red-team trial, OpenAI’s uplift study with Gryphon Scientific, Meta’s internal Llama 3 trial) all found no statistically significant uplift, and are now too old to say much about current systems.
Cyber: real bugs, incomplete chains
Google’s Big Sleep agent found a previously unknown exploitable memory-safety flaw in SQLite in 2024, and in 2025 found CVE-2025-6965, a vulnerability Google says threat actors already knew about, in time to patch it before exploitation. That is a genuine first. But on structured benchmarks the picture is deflationary: as of the Cybench evaluation no frontier model could autonomously complete capture-the-flag tasks that took expert humans more than eleven minutes to first-solve, and xAI’s own Grok 4 card states that end-to-end offensive cyber capability remains below that of a human professional. Finding a bug is not running an operation.
Scheming and autonomy: behaviours are real, spontaneity is not established
Apollo Research found that five of six frontier models tested exhibited in-context scheming in at least one constructed scenario. Anthropic and Redwood documented alignment faking with the reasoning legible in the chain of thought. Anthropic’s own pilot sabotage risk report describes the misalignment risk from Opus 4 as “very low but not fully negligible.”
Against that: DeepMind’s stealth and situational-awareness suite found the strongest models passing 2 of 11 and 2 of 5 challenges, well below the threshold that would break an inability-based safety case. The UK AI Security Institute, drawing on two years of continuous evaluation, reports self-replication success rising from 5% to 60% between 2023 and 2025 and autonomous task length in the cyber domain doubling roughly every eight months, while finding no evidence of spontaneous, unprompted sandbagging or self-replication.
Why the evaluation record is weaker than it looks
- Underpowered. Most uplift studies run fewer than 20 participants per condition, including the trials that triggered a safety-level change.
- Unaudited. Essentially every threshold determination is made by a lab about its own model. METR’s autonomy work is the main independent exception.
- Not comparable. Each lab uses different tasks, graders and thresholds, so cross-model safety claims cannot be checked against each other.
- Thresholds are asserted. The numbers themselves (5× uplift, a compute figure, a benchmark cut-off) are internal conventions, not values derived from a model of actual harm.
- Gameable. Sandbagging and evaluation-awareness are documented. There is currently no method that distinguishes genuine incapability from strategic underperformance at evaluation time.
- Wrong threat actor. Trials typically measure novices on short isolated tasks; the actor the policy is written for is a determined expert with months.
The record
All 200 entries, chronological, each with the body that produced it and a link to the source. Search matches title, author, venue and description. This is the part of the page meant to be used rather than read.
Entries marked source not retrieved are listed for completeness but were not confirmed against the primary document at compilation; treat them as leads. Machine-readable copies of this table are linked in section 10.
What the law
actually requires
Almost everything written about AI regulation describes instruments at the moment they were announced. Several of the most-cited have since been delayed, revoked, renamed or narrowed. This section states current status as of 23 September 2026, with the primary instrument linked in every case.
EU AI Act. Prohibitions since February 2025; general-purpose model obligations (documentation, adversarial testing, incident reporting above the 1025 FLOP systemic-risk presumption) since August 2025. The only legally enforceable frontier-model regime anywhere. California SB 53, signed September 2025: published frontier frameworks, catastrophic-risk reporting, whistleblower protection. China: generative-AI registration since 2023, content labelling since September 2025.
EU high-risk obligations, long expected in August 2026, were postponed by a 423–57 European Parliament vote in June 2026 to December 2027, and to August 2028 for AI embedded in safety components; Council adoption is still pending. Colorado’s AI Act was signed, then materially weakened in May 2026 with its duty of care removed and commencement pushed to January 2027. The AI diffusion rule on chip exports was rescinded two days before it took effect, with no replacement published since.
The Seoul Frontier AI Safety Commitments, the G7 Hiroshima code, the 2023 White House commitments and every published lab safety framework. Companies set their own thresholds, run their own evaluations and decide their own consequences. No jurisdiction requires independent pre-deployment audit of a frontier model. The Council of Europe convention, the only binding AI treaty, has one ratification, the EU’s, and is not in force.
All 47 Governance entries, including the Seoul commitments, the G7 Hiroshima process, the UN scientific panel, China’s framework, the Frontier Model Forum and every lab safety framework, are in the canon with their current status recorded.
Why many researchers
think this is wrong
Five distinct critiques, which disagree with each other as sharply as they disagree with the risk case. Each is stated here as its own proponents state it, without rebuttal. The responses come at the end, clearly labelled, because a reader is entitled to see the argument before they see the counter-argument.
1. The threat model doesn’t describe anything we are building
2. The capabilities are not what the benchmarks say
3. It displaces the harms that are already happening
4. Safety framing is a moat
5. The probability estimates do not mean what they are used to mean
Where risk-side researchers say these critiques are weakest
- On the threat model: the argument was never about language models. Instrumental convergence is a claim about what any sufficiently capable goal-directed system does; criticising the limits of current architectures leaves it untouched.
- On capabilities: the relevant threshold is not today’s systems, and capability surprises are documented rather than hypothetical. Benchmark saturation has been repeatedly predicted late.
- On present harms: the two agendas are complementary rather than zero-sum, and lab safety frameworks in fact cover catastrophic-but-not-existential and near-term harms together. Whether attention is actually displaced is an empirical question that has not been settled.
- On political economy: the incumbency argument cuts both ways, and many researchers took this position years before any institution existed to fund it.
- On epistemics: Ngo’s response is that the case can be grounded in near-term observable mechanisms rather than long-run probabilities: you do not need a defensible P(doom) to justify precaution against a mechanism you can already demonstrate.
What would change
the picture
The test of a record like this is whether it specifies in advance what would move it. These are indicators that would shift the weight of evidence in one direction or the other, chosen because each is observable, dated and reported by someone other than an advocate.
A well-powered wet-lab trial showing significant uplift. The in-silico results are contested precisely because the bench results are negative; a large replicated wet-lab positive would end that argument.
Spontaneous sandbagging. Every scheming result so far is constructed or prompted. A documented case of a model strategically underperforming an evaluation it was not set up to game would be a categorical change, and the GPT-5.6 Sol result shows the boundary is being approached.
An autonomous operation end to end. Not a found vulnerability, but a full chain: reconnaissance, exploitation, persistence, without human steps.
Time horizons past the reliability ceiling. METR cannot currently measure above sixteen hours. A rebuilt suite showing multi-day autonomous competence would make the extrapolation argument much harder to dismiss.
The horizon curve bending. The 89-day doubling is measured over a short window on a saturating suite. Two years of flat measurement on a rebuilt benchmark would undercut the central quantitative claim.
Replicated negative uplift with current models. If frontier models keep failing to move the needle at the bench as capability rises, the misuse mechanism weakens in exactly the domain governments prioritise.
Interpretability delivering verification. If internal-state methods mature enough to certify the absence of deceptive computation, the scheming argument loses its central epistemic advantage: that you cannot check.
Scaling stalling on economics. At 2.4× cost growth per year the frontier is a handful of organisations. A halt driven by capital rather than capability would change the timelines every forecast here depends on.
Cite this record
This page is a compilation, not original research. Cite the primary source for any claim; every entry links to it. If you need to cite the compilation itself:
AI Global Code Red. The Artificial Intelligence Risk Record, 1863–2026. ai.codered.global, compiled 23 September 2026. 200 entries.
Corrections are the point of a record like this. If an entry is misdated, misattributed, superseded or simply wrong, the error is in the compilation and worth reporting. Several entries here exist because widely repeated claims turned out on checking to be misdated or misattributed.
Machine-readable editions
- ai-code-red-canon.csv: all 200 entries with source URLs
- ai-code-red-canon.json: the same data as JSON
- llms.txt: plain-text edition for language models