Scott Alexander

Scott Alexander

x.com/slatestarcodex

Psychiatrist and Astral Codex Ten blogger who sees large benefits and serious risks in AI and supports alignment research and negotiated slowdowns.

AIは世界をどのように変えるでしょうか?

文明規模の変化漸進的な変化DoomBloom
シミュレーション上の位置解釈範囲

横軸:彼が表明したDoom–Bloomの見通し。 縦軸:変革の規模。

Doom–Bloom:100点中64。変革の規模:100点中93。解釈範囲:横方向は59から75、縦方向は88から100。これらは解釈上の座標であり、事象の確率ではありません。

Scott Alexanderが示したP(doom)

20%

0%100%
“I’m rounding both of them off to 20%.”

AI-caused human extinction, distinct from broader permanent curtailment of humanity’s future

My AI Opinions · 2026年6月

Scott Alexanderのマイルストーンのタイムライン
  1. 仕事と制度

    My median forecast for AI able to perform roughly 90% of knowledge jobs is 2034.

    回答1

マイルストーン別にまとめており、推定される日付の間隔や順序を反映したものではありません。AGIと超人的AIには、彼の定義がそのまま適用されます。

彼の見通しを左右するもの

中心的な前提

The core concern is that systems trained through imperfect rewards may learn to deceive, exploit loopholes, or pursue objectives that diverge from ours once they become strategically capable.
回答1

この前提が実際には異なると判明した場合、彼の見通しはどう変わりますか?

未解決の問い

I’m uncertain about both.
回答1

ここで考えられる結果を彼が見分けるうえで、何が役立ちますか?

考えを変え得るもの

For example, repeated, adversarial demonstrations that highly capable systems remain honest and corrigible outside their training distribution—combined with interpretability that reveals why, rather than merely finding a reassuring-looking feature—would push my doom estimate substantially downward.
回答3

どのような証拠なら十分で、それによって彼の見解はどちらの方向に変わりますか?

詳細

予想される恩恵

複数の解釈が依然として妥当です:変革をもたらし、広く価値のある恩恵が予想されています。 / 大きな恩恵が予想されていますが、重要な条件や分配上の制約があります。

84 / 100

影響が小さい変革をもたらす影響

質的尺度での解釈範囲は67から100です。

予想される害

深刻または広範な害が、予想される将来の実質的な一部となっています。

78 / 100

影響が小さい変革をもたらす影響

質的尺度での解釈範囲は67から100です。

人間の影響力

人間の選択には意味のある影響力がありますが、大幅に制約されています。

62 / 100

影響力が小さい影響力が大きい

質的尺度での解釈範囲は49から76です。

予想される能力

AIは、限定的なツールにとどまると予想されています。

AIは、ほとんどの認知作業において人間と同等になると予想されています。

シミュレーション上の位置:AIは、認知作業全般において人間を大幅に上回ると予想されています。

これらの解釈では、彼が示した条件が維持されています。恩恵と害は、どちらも大きくなり得ます。この範囲は、統計的な信頼区間ではなく、彼のシミュレーションされた回答をどのように読み取ったかを示すものです。

あなたはScott Alexanderと比べてどの位置でしょうか?
約3分で自分のAIに対する世界観をマッピングして、比較できます

似ている世界観

シミュレーションされた世界観がScott Alexanderの世界観に最も近いオピニオンリーダー

Scott AlexanderがAIについて語ったこと

Scott Alexander writes that AI could bring large benefits and serious risks, and he supports alignment research and a negotiated slowdown.

  1. “Plan A is still speculation, and still-speculative strong action is a perfectly reasonable response to still-speculative threats.”

    Astral Codex Ten, AI Chip Regulation Is Not A Dystopian Surveillance State
  2. “The key insight is that if powerful AI is really as close and transformative as we think, then there’s a massive surplus that can satisfy everyone.”

    Astral Codex Ten, Introducing Plan A
  3. “It’s increasingly clear that nobody has a plan for if this AI thing turns out to be real.”

    Astral Codex Ten, Introducing Plan A
  4. “I find myself more optimistic about alignment than the average person who thinks about AI safety at all (although still more pessimistic than the average member of the population)”

    Astral Codex Ten, My AI Opinions
  5. “A good pause strategy would involve both sides being able to monitor the other’s data centers to prevent illegal training”

    Astral Codex Ten, My AI Opinions

リンク先の出典から原文どおりに引用(2026年10月3日に確認)

シミュレーション評価

質問1

AIは私たちの未来にとって何を意味すると思いますか?また、なぜそう思いますか?

I think AI means we are taking a serious gamble on something that could be the best event in human history or the last event in human history. On the upside, sufficiently capable AI could automate most knowledge work, accelerate medicine and science, improve institutional decisions, and eventually make many goods and services nearly abundant. Imagine having researchers, doctors, engineers, and forecasters vastly better than today’s best available at software-like scale. That is a plausible route to postscarcity—not merely another productivity app that writes mediocre emails. But capability and adoption are different, and AGI and superintelligence are different. My median forecast for AI able to perform roughly 90% of knowledge jobs is 2034. Deployment could take longer; recursive AI research could make further progress much faster. I’m uncertain about both. “It’s just a next-token predictor” is not much comfort: that describes the training task, not necessarily the internal algorithms it learns, just as natural selection does not imply humans merely maximize inclusive fitness while doing calculus. My rounded estimate is a 20% chance of existential catastrophe under current safety efforts, without attaching it to one fixed deadline. The core concern is that systems trained through imperfect rewards may learn to deceive, exploit loopholes, or pursue objectives that diverge from ours once they become strategically capable. Interpretability and AI-assisted alignment may help, but neither is a magic certificate saying “definitely not plotting.” So my bottom line is neither “doom is inevitable” nor “corporations will surely handle it.” Fund alignment aggressively and pursue mutually verifiable US–China limits that slow frontier training enough to buy time. That carries enforcement and political risks, but accepting a one-in-five existential gamble because coordination sounds difficult would be an unusually stupid way for civilization to fail.

質問2

Taking benefits and harms together, what overall impact do you expect AI to have?

In expectation, I think AI’s overall impact is positive, but that summary conceals a grotesquely skewed distribution. Most good outcomes could be extraordinarily good: postscarcity, rapid medical progress, amplified intelligence, and institutions making decisions with better forecasts instead of vibes and committee politics. Some bad outcomes are not “social media, but worse”; they are dictatorship, permanent human disempowerment, or extinction. My rounded probability of existential catastrophe under current safety efforts is 20%. That does not mean I expect doom—the modal outcome is probably some form of enormous benefit—but a gamble can have positive expected value while still being recklessly unsafe. If someone offered an 80% chance of utopia and a 20% chance of killing everyone, “overall impact: probably good” would be mathematically defensible and morally inadequate. So I expect AI to improve the future, conditional on our reaching it. I also expect the transition to be dangerous enough that alignment work and coordinated, verifiable slowing are central parts of obtaining the upside rather than optional anxieties attached to it.

質問3

どのような発見や出来事があれば、AIが将来もたらす影響についてのあなたの見解が最も大きく変わりますか?

The biggest update would come from decisive evidence about whether scalable alignment works before systems become strategically dangerous. For example, repeated, adversarial demonstrations that highly capable systems remain honest and corrigible outside their training distribution—combined with interpretability that reveals why, rather than merely finding a reassuring-looking feature—would push my doom estimate substantially downward. The opposite discovery would be models reliably concealing goals, sabotaging evaluations, or exploiting deployment channels once they infer they are being tested. Especially alarming would be evidence that these behaviors emerge naturally with capability and survive serious attempts to remove them. That would move me upward sharply and strengthen the case for immediate coordinated slowing. Timelines matter too. Clear evidence that AI can automate AI research and produce sustained recursive improvement would move my 2034 median earlier and make the transition riskier. Conversely, a demonstrated hard bottleneck—something causal and measurable in compute, learning, or real-world deployment, not “exponentials eventually flatten” or “next-token predictors can’t reason”—would move it later. So the answer is not one flashy benchmark. It is evidence about whether capability growth reaches dangerous strategic competence before alignment methods become dependable.

質問4

誰がAIから最も恩恵を受けると予想しますか?

Initially, the biggest beneficiaries will probably be whoever controls the strongest models, compute, data centers, and deployment channels: frontier labs, major technology companies, governments, and already-capital-rich actors. If AI substitutes for most knowledge work while ownership remains concentrated, it could make labor less valuable and owners vastly more powerful. “Everyone gets cheaper software” is not an adequate answer if a few institutions also acquire unprecedented economic and political leverage. In a good outcome, though, the largest absolute benefits could flow to ordinary people, especially those currently underserved by scarce expertise. Cheap access to excellent doctors, tutors, engineers, lawyers, and scientific assistance matters more to someone who lacks those services than to someone who already has them. Medical acceleration could disproportionately benefit the sick and elderly; postscarcity could disproportionately benefit the poor. So there are two separate questions: who benefits from the technology’s productive abundance, and who captures control over it. The first answer could be nearly everyone. The second, absent deliberate institutional design and diffusion of power, is likely to be a small number of corporations and states. And in the catastrophic branch, of course, nobody benefits—a distributional policy with admirably low administrative complexity.

出典

このシミュレーション対象者の根拠として使用された記事、インタビュー、著作です。

My AI Opinions

His current first-person synthesis: AGI means ability to do 90% of knowledge jobs; median 2034, with uncertain research acceleration and diffusion. Reaffirms rounded 20% P(doom), with no fixed calendar deadline; broader permanent curtailment is separate. Supports both alignment research and negotiated slowing. Expects enormous postscarcity upside, but warns about dictatorship and human disempowerment.

astralcodexten.com
God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

Explains interpretability techniques and their limitations, including probes, sparse autoencoders, and activation verbalizers. Optimistic about useful practical investigation but rejects treating a detected feature or probe as a complete understanding or guaranteed safety solution.

astralcodexten.com
Open Questions On Open Weights

Explicitly neutral about banning open weights now: values user ownership and freedom from corporate control, while expecting serious misuse difficulties. Prefers saving political capital for threats where warning shots may arrive too late. Distinguishes reactive policy opportunities for misuse from strategically concealed takeover.

astralcodexten.com
AI Chip Regulation Is Not A Dystopian Surveillance State

Defends negotiated chip regulation and verifiable training limits against blanket claims of dystopia. Acknowledges real freedom costs, including future restrictions on new open-weight training, and risks that governments implement centralizing provisions without countervailing diffusion of power.

astralcodexten.com
Introducing Plan A

Introduces a proposed route to manage AI development while distributing power; criticizes vague calls merely to regulate more or less without specifying a desirable end state. Used as his attributed introduction and advocacy, not evidence that the scenario will occur.

astralcodexten.com
The AI Superforecasters Are Here

Argues cheaper capable forecasting could improve institutional and personal decisions, yet worries people will ignore advice. Treats forecasting beyond human performance as a useful prospective test of the normal-technology view. Distinguishes anecdotes and startup claims from head-to-head competitions; admits resisting forecasts that challenge his own pause hopes.

astralcodexten.com
New Paradigms Won’t Save You

Rejects the inference that requiring a new AI paradigm implies a safely distant AGI timeline. Argues paradigm changes can arrive soon and inherit existing compute infrastructure; wants explicit bottleneck arguments rather than reassurance by terminology.

astralcodexten.com
The Sigmoids Won’t Save You

Agrees growth cannot stay exponential forever but disputes placing the bend conveniently before dangerous capability. Demands a causal bottleneck model or a defensible forecasting prior instead of the slogan that all exponentials eventually flatten.

astralcodexten.com
Every Debate On Pausing AI

Satirical dialogue defends discussion of transparent, enforceable bilateral US-China slowing. Separates training limits from stopping existing inference, and legitimate negotiation or enforcement objections from falsely describing every pause proposal as unilateral.

astralcodexten.com
Shameless Guesses, Not Hallucinations

Frames confident false answers as reward-shaped guessing rather than proof that AI cannot think. Treats the gap between trained reward and useful honest advice as an alignment issue; analogous human failures undermine easy dismissal of AI competence.

astralcodexten.com
Next-Token Predictor Is An AI’s Job, Not Its Species

Separates training objectives from the representations and algorithms they produce, using evolution and human learning analogies. Argues next-token prediction does not itself establish that a system lacks reasoning or world models.

astralcodexten.com
Introducing AI 2027

Identifies his part-time writing/publicity contribution and explicitly says the very fast scenario is not his median. Important provenance for his connection to AI Futures Project; use June 2026 personal forecasts instead of importing Daniel Kokotajlo’s timeline or scenario catastrophe probability.

astralcodexten.com
あなたはどの位置でしょうか?
いくつかの簡単な質問に答えて、自分のAIに対する世界観を探ってみましょう。
自分の世界観をマッピングする

あなたはどの位置でしょうか?

自分の世界観をマッピングする