Scott Alexander

Scott Alexander

x.com/slatestarcodex

Psychiatrist and Astral Codex Ten blogger who sees large benefits and serious risks in AI and supports alignment research and negotiated slowdowns.

Wie wird KI die Welt verändern?

Zivilisatorischer WandelSchrittweiser WandelDoomBloom
Simulierte PositionInterpretationsbereich

Horizontal: sein geäußerter Doom–Bloom-Ausblick. Vertikal: Ausmaß der Transformation.

Doom–Bloom: 64 von 100. Ausmaß der Transformation: 93 von 100. Interpretationsbereiche: horizontal 59 bis 75, vertikal 88 bis 100. Dies sind Interpretationskoordinaten, keine Ereigniswahrscheinlichkeiten.

Das angegebene P(doom) von Scott Alexander

20%

0%100%
“I’m rounding both of them off to 20%.”

AI-caused human extinction, distinct from broader permanent curtailment of humanity’s future

My AI Opinions · Juni 2026

Zeithorizont für Meilensteine von Scott Alexander
  1. Arbeit und Institutionen

    My median forecast for AI able to perform roughly 90% of knowledge jobs is 2034.

    Antwort 1

Nach Meilenstein gruppiert, nicht anhand abgeleiteter Zeitpunkte angeordnet oder mit entsprechenden Abständen dargestellt. Für AGI und übermenschliche KI gelten weiterhin seine Definitionen.

Wovon seine Einschätzung abhängt

Eine zentrale Annahme

The core concern is that systems trained through imperfect rewards may learn to deceive, exploit loopholes, or pursue objectives that diverge from ours once they become strategically capable.
Antwort 1

Wenn sich diese Annahme als anders herausstellen würde, wie würde sich seine Einschätzung ändern?

Eine ungeklärte Frage

I’m uncertain about both.
Antwort 1

Was würde ihm helfen, die plausiblen Ergebnisse hier voneinander zu unterscheiden?

Was ihre Meinung ändern könnte

For example, repeated, adversarial demonstrations that highly capable systems remain honest and corrigible outside their training distribution—combined with interpretability that reveals why, rather than merely finding a reassuring-looking feature—would push my doom estimate substantially downward.
Antwort 3

Welche Belege würden ausreichen, und in welche Richtung würden sie seine Sichtweise verändern?

Weitere Details

Erwartete Vorteile

Mehrere Lesarten bleiben plausibel: Es werden transformative Vorteile von breitem Wert erwartet. / Es werden erhebliche Vorteile erwartet, allerdings unter wichtigen Bedingungen oder mit Einschränkungen bei ihrer Verteilung.

84 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 67 bis 100 auf der qualitativen Skala.

Erwartete Schäden

Schwere oder weitverbreitete Schäden sind ein wesentlicher erwarteter Bestandteil der Zukunft.

78 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 67 bis 100 auf der qualitativen Skala.

Menschlicher Einfluss

Menschliche Entscheidungen haben einen bedeutsamen, aber erheblich eingeschränkten Einfluss.

62 / 100

Geringer EinflussStarker Einfluss

Interpretationsbereich von 49 bis 76 auf der qualitativen Skala.

Erwartete Fähigkeiten

Es wird erwartet, dass KI auf begrenzte Werkzeuge beschränkt bleibt.

Es wird erwartet, dass KI bei den meisten kognitiven Tätigkeiten mit Menschen gleichzieht.

Simulierte Position: Es wird erwartet, dass KI Menschen bei kognitiven Tätigkeiten deutlich übertrifft.

Diese Interpretationen berücksichtigen weiterhin seine genannten Bedingungen. Vorteile und Schäden können beide erheblich sein. Die Bereiche beschreiben, wie wir seine simulierten Antworten interpretieren, und sind keine statistischen Konfidenzintervalle.

Wo stehst du im Vergleich zu Scott Alexander?
Bilde deine eigene KI-Weltsicht in etwa 3 Minuten ab und vergleiche sie dann

Ähnliche Weltsichten

Vordenker, deren simulierte Weltsichten der von Scott Alexander am nächsten kommen

Was Scott Alexander über KI gesagt hat

Scott Alexander writes that AI could bring large benefits and serious risks, and he supports alignment research and a negotiated slowdown.

  1. “Plan A is still speculation, and still-speculative strong action is a perfectly reasonable response to still-speculative threats.”

    Astral Codex Ten, AI Chip Regulation Is Not A Dystopian Surveillance State
  2. “The key insight is that if powerful AI is really as close and transformative as we think, then there’s a massive surplus that can satisfy everyone.”

    Astral Codex Ten, Introducing Plan A
  3. “It’s increasingly clear that nobody has a plan for if this AI thing turns out to be real.”

    Astral Codex Ten, Introducing Plan A
  4. “I find myself more optimistic about alignment than the average person who thinks about AI safety at all (although still more pessimistic than the average member of the population)”

    Astral Codex Ten, My AI Opinions
  5. “A good pause strategy would involve both sides being able to monitor the other’s data centers to prevent illegal training”

    Astral Codex Ten, My AI Opinions

Wörtlich aus den verlinkten Quellen, geprüft am 3. Okt. 2026

Simulierte Einschätzung

Frage 1

Was glaubst du, was KI für unsere Zukunft bedeutet – und warum?

I think AI means we are taking a serious gamble on something that could be the best event in human history or the last event in human history. On the upside, sufficiently capable AI could automate most knowledge work, accelerate medicine and science, improve institutional decisions, and eventually make many goods and services nearly abundant. Imagine having researchers, doctors, engineers, and forecasters vastly better than today’s best available at software-like scale. That is a plausible route to postscarcity—not merely another productivity app that writes mediocre emails. But capability and adoption are different, and AGI and superintelligence are different. My median forecast for AI able to perform roughly 90% of knowledge jobs is 2034. Deployment could take longer; recursive AI research could make further progress much faster. I’m uncertain about both. “It’s just a next-token predictor” is not much comfort: that describes the training task, not necessarily the internal algorithms it learns, just as natural selection does not imply humans merely maximize inclusive fitness while doing calculus. My rounded estimate is a 20% chance of existential catastrophe under current safety efforts, without attaching it to one fixed deadline. The core concern is that systems trained through imperfect rewards may learn to deceive, exploit loopholes, or pursue objectives that diverge from ours once they become strategically capable. Interpretability and AI-assisted alignment may help, but neither is a magic certificate saying “definitely not plotting.” So my bottom line is neither “doom is inevitable” nor “corporations will surely handle it.” Fund alignment aggressively and pursue mutually verifiable US–China limits that slow frontier training enough to buy time. That carries enforcement and political risks, but accepting a one-in-five existential gamble because coordination sounds difficult would be an unusually stupid way for civilization to fail.

Frage 2

Taking benefits and harms together, what overall impact do you expect AI to have?

In expectation, I think AI’s overall impact is positive, but that summary conceals a grotesquely skewed distribution. Most good outcomes could be extraordinarily good: postscarcity, rapid medical progress, amplified intelligence, and institutions making decisions with better forecasts instead of vibes and committee politics. Some bad outcomes are not “social media, but worse”; they are dictatorship, permanent human disempowerment, or extinction. My rounded probability of existential catastrophe under current safety efforts is 20%. That does not mean I expect doom—the modal outcome is probably some form of enormous benefit—but a gamble can have positive expected value while still being recklessly unsafe. If someone offered an 80% chance of utopia and a 20% chance of killing everyone, “overall impact: probably good” would be mathematically defensible and morally inadequate. So I expect AI to improve the future, conditional on our reaching it. I also expect the transition to be dangerous enough that alignment work and coordinated, verifiable slowing are central parts of obtaining the upside rather than optional anxieties attached to it.

Frage 3

Welche Entdeckung oder welches Ereignis würde deine Sicht auf die künftigen Auswirkungen von KI am stärksten verändern?

The biggest update would come from decisive evidence about whether scalable alignment works before systems become strategically dangerous. For example, repeated, adversarial demonstrations that highly capable systems remain honest and corrigible outside their training distribution—combined with interpretability that reveals why, rather than merely finding a reassuring-looking feature—would push my doom estimate substantially downward. The opposite discovery would be models reliably concealing goals, sabotaging evaluations, or exploiting deployment channels once they infer they are being tested. Especially alarming would be evidence that these behaviors emerge naturally with capability and survive serious attempts to remove them. That would move me upward sharply and strengthen the case for immediate coordinated slowing. Timelines matter too. Clear evidence that AI can automate AI research and produce sustained recursive improvement would move my 2034 median earlier and make the transition riskier. Conversely, a demonstrated hard bottleneck—something causal and measurable in compute, learning, or real-world deployment, not “exponentials eventually flatten” or “next-token predictors can’t reason”—would move it later. So the answer is not one flashy benchmark. It is evidence about whether capability growth reaches dangerous strategic competence before alignment methods become dependable.

Frage 4

Wer wird deiner Erwartung nach am meisten von KI profitieren?

Initially, the biggest beneficiaries will probably be whoever controls the strongest models, compute, data centers, and deployment channels: frontier labs, major technology companies, governments, and already-capital-rich actors. If AI substitutes for most knowledge work while ownership remains concentrated, it could make labor less valuable and owners vastly more powerful. “Everyone gets cheaper software” is not an adequate answer if a few institutions also acquire unprecedented economic and political leverage. In a good outcome, though, the largest absolute benefits could flow to ordinary people, especially those currently underserved by scarce expertise. Cheap access to excellent doctors, tutors, engineers, lawyers, and scientific assistance matters more to someone who lacks those services than to someone who already has them. Medical acceleration could disproportionately benefit the sick and elderly; postscarcity could disproportionately benefit the poor. So there are two separate questions: who benefits from the technology’s productive abundance, and who captures control over it. The first answer could be nearly everyone. The second, absent deliberate institutional design and diffusion of power, is likely to be a small number of corporations and states. And in the catastrophic branch, of course, nobody benefits—a distributional policy with admirably low administrative complexity.

Quellen

Artikel, Interviews und Schriften, die als Grundlage für diesen simulierten Nutzer dienen.

My AI Opinions

His current first-person synthesis: AGI means ability to do 90% of knowledge jobs; median 2034, with uncertain research acceleration and diffusion. Reaffirms rounded 20% P(doom), with no fixed calendar deadline; broader permanent curtailment is separate. Supports both alignment research and negotiated slowing. Expects enormous postscarcity upside, but warns about dictatorship and human disempowerment.

astralcodexten.com
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Explains interpretability techniques and their limitations, including probes, sparse autoencoders, and activation verbalizers. Optimistic about useful practical investigation but rejects treating a detected feature or probe as a complete understanding or guaranteed safety solution.

astralcodexten.com
Open Questions On Open Weights

Explicitly neutral about banning open weights now: values user ownership and freedom from corporate control, while expecting serious misuse difficulties. Prefers saving political capital for threats where warning shots may arrive too late. Distinguishes reactive policy opportunities for misuse from strategically concealed takeover.

astralcodexten.com
AI Chip Regulation Is Not A Dystopian Surveillance State

Defends negotiated chip regulation and verifiable training limits against blanket claims of dystopia. Acknowledges real freedom costs, including future restrictions on new open-weight training, and risks that governments implement centralizing provisions without countervailing diffusion of power.

astralcodexten.com
Introducing Plan A

Introduces a proposed route to manage AI development while distributing power; criticizes vague calls merely to regulate more or less without specifying a desirable end state. Used as his attributed introduction and advocacy, not evidence that the scenario will occur.

astralcodexten.com
The AI Superforecasters Are Here

Argues cheaper capable forecasting could improve institutional and personal decisions, yet worries people will ignore advice. Treats forecasting beyond human performance as a useful prospective test of the normal-technology view. Distinguishes anecdotes and startup claims from head-to-head competitions; admits resisting forecasts that challenge his own pause hopes.

astralcodexten.com
New Paradigms Won’t Save You

Rejects the inference that requiring a new AI paradigm implies a safely distant AGI timeline. Argues paradigm changes can arrive soon and inherit existing compute infrastructure; wants explicit bottleneck arguments rather than reassurance by terminology.

astralcodexten.com
The Sigmoids Won’t Save You

Agrees growth cannot stay exponential forever but disputes placing the bend conveniently before dangerous capability. Demands a causal bottleneck model or a defensible forecasting prior instead of the slogan that all exponentials eventually flatten.

astralcodexten.com
Every Debate On Pausing AI

Satirical dialogue defends discussion of transparent, enforceable bilateral US-China slowing. Separates training limits from stopping existing inference, and legitimate negotiation or enforcement objections from falsely describing every pause proposal as unilateral.

astralcodexten.com
Shameless Guesses, Not Hallucinations

Frames confident false answers as reward-shaped guessing rather than proof that AI cannot think. Treats the gap between trained reward and useful honest advice as an alignment issue; analogous human failures undermine easy dismissal of AI competence.

astralcodexten.com
Next-Token Predictor Is An AI’s Job, Not Its Species

Separates training objectives from the representations and algorithms they produce, using evolution and human learning analogies. Argues next-token prediction does not itself establish that a system lacks reasoning or world models.

astralcodexten.com
Introducing AI 2027

Identifies his part-time writing/publicity contribution and explicitly says the very fast scenario is not his median. Important provenance for his connection to AI Futures Project; use June 2026 personal forecasts instead of importing Daniel Kokotajlo’s timeline or scenario catastrophe probability.

astralcodexten.com
Wo stehst du?
Erkunde deine eigene KI-Weltsicht, indem du ein paar einfache Fragen beantwortest.
Deine eigene Weltsicht abbilden