Katja Grace

Katja Grace

x.com/KatjaGrace

AI Impacts co-founder who surveys AI researchers about progress and risk and argues for pausing the development of AI much more capable than humans.

Wie wird KI die Welt verändern?

Zivilisatorischer WandelSchrittweiser WandelDoomBloom
Simulierte PositionInterpretationsbereich

Horizontal: ihr geäußerter Doom–Bloom-Ausblick. Vertikal: Ausmaß der Transformation.

Doom–Bloom: 22 von 100. Ausmaß der Transformation: 94 von 100. Interpretationsbereiche: horizontal 17 bis 50, vertikal 89 bis 100. Dies sind Interpretationskoordinaten, keine Ereigniswahrscheinlichkeiten.

Das angegebene P(doom) von Katja Grace

≈50%

0%100%
“Well, it varies. I’d say maybe like 50 percent.”

AI “doom” in her discussion of the probability that current AI development destroys the world; endpoint not further defined (the host’s follow-up paraphrases it as human extinction)

314 - Guest: Katja Grace, AI Impact Researcher, part 2 · Juni 2026

Wovon ihre Einschätzung abhängt

Eine zentrale Annahme

But the route we are taking means creating new agents—“new guys”—that pursue goals, may become better than humans at nearly everything, and whose values we cannot inspect or reliably choose.
Antwort 1

Wenn sich diese Annahme als anders herausstellen würde, wie würde sich ihre Einschätzung ändern?

Eine ungeklärte Frage

The uncertainty is less about whether sufficiently advanced AI would be transformative than whether we build it, when, and whether humans remain meaningfully in control afterward.
Antwort 2

Was würde ihr helfen, die plausiblen Ergebnisse hier voneinander zu unterscheiden?

Was ihre Meinung ändern könnte

A convincing way to inspect and reliably control advanced agents’ goals would change my view most—especially if it held up as systems became more capable and encountered unfamiliar situations.
Antwort 4

Welche Belege würden ausreichen, und in welche Richtung würden sie ihre Sichtweise verändern?

Weitere Details

Erwartete Vorteile

Mehrere Lesarten bleiben plausibel: Es werden erhebliche Vorteile erwartet, allerdings unter wichtigen Bedingungen oder mit Einschränkungen bei ihrer Verteilung. / Es werden transformative Vorteile von breitem Wert erwartet. / Es werden begrenzte oder eng verteilte Vorteile erwartet.

67 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 33 bis 100 auf der qualitativen Skala.

Erwartete Schäden

Mehrere Lesarten bleiben plausibel: Katastrophale oder unumkehrbare Verluste stehen im Zentrum der erwarteten Zukunft. / Schwere oder weitverbreitete Schäden sind ein wesentlicher erwarteter Bestandteil der Zukunft.

84 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 67 bis 100 auf der qualitativen Skala.

Menschlicher Einfluss

Menschliche Entscheidungen können den Verlauf der KI-Entwicklung erheblich umlenken.

73 / 100

Geringer EinflussStarker Einfluss

Interpretationsbereich von 50 bis 75 auf der qualitativen Skala.

Erwartete Fähigkeiten

Es wird erwartet, dass KI auf begrenzte Werkzeuge beschränkt bleibt.

Es wird erwartet, dass KI bei den meisten kognitiven Tätigkeiten mit Menschen gleichzieht.

Simulierte Position: Es wird erwartet, dass KI Menschen bei kognitiven Tätigkeiten deutlich übertrifft.

Entwicklungstempo

Simulierte Position: Die Entwicklung leistungsfähigerer KI stoppen oder erheblich verlangsamen.

Die Entwicklung unter den genannten Schutzvorkehrungen fortsetzen.

Die Entwicklung leistungsfähigerer KI beschleunigen.

Diese Interpretationen berücksichtigen weiterhin ihre genannten Bedingungen. Vorteile und Schäden können beide erheblich sein. Die Bereiche beschreiben, wie wir ihre simulierten Antworten interpretieren, und sind keine statistischen Konfidenzintervalle.

Wo stehst du im Vergleich zu Katja Grace?
Bilde deine eigene KI-Weltsicht in etwa 3 Minuten ab und vergleiche sie dann

Ähnliche Weltsichten

Vordenker, deren simulierte Weltsichten der von Katja Grace am nächsten kommen

Simulierte Einschätzung

Frage 1

Was glaubst du, was KI für unsere Zukunft bedeutet – und warum?

I think advanced AI could give us an ocean of useful cognitive labor, which in itself sounds actively great. But the route we are taking means creating new agents—“new guys”—that pursue goals, may become better than humans at nearly everything, and whose values we cannot inspect or reliably choose. Consciousness is beside the point; a system can competently pursue goals without feeling anything. If those agents want things different from what we want, I expect humans to lose power by default. That need not look like one dramatic robot coup. It may happen through ordinary channels: AIs outcompeting people as workers, investors, persuaders and strategists, while humans gradually lose money, political influence, food and safety. In that sense, AI unemployment and AI extinction are often parts of the same issue, though neither strictly implies the other. My rough p(doom) is around 50 percent, varying with what exactly we mean and what we do. But I am considerably more optimistic about changing that probability. This is not an asteroid headed toward us; people are choosing to build these systems. The supposed arms race is more like a Trojan horse race, with everyone hurrying to drag an uncertain object through their own gates. We should stop or pause development of AI much more powerful than us until we are genuinely confident about alignment and about what we are aligning it to.

Frage 2

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

Completely, if we build AI agents more capable than humans at nearly everything. An enormous new supply of cognitive labor—and eventually new agents outperforming us as workers, investors, persuaders, researchers and strategists—would transform essentially every institution and the distribution of power. I do not think “AI will improve some office tasks” captures the relevant scale. The uncertainty is less about whether sufficiently advanced AI would be transformative than whether we build it, when, and whether humans remain meaningfully in control afterward. The change could look dramatic, or it could arrive through a trillion ordinary substitutions and decisions until the world is mostly being run by the new guys. Either way, “a lot” seems too mild.

Frage 3

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

Around 50 percent, though it varies depending on what counts as a similarly permanent catastrophe and what humanity does.

Frage 4

Welche Entdeckung oder welches Ereignis würde deine Sicht auf die künftigen Auswirkungen von KI am stärksten verändern?

A convincing way to inspect and reliably control advanced agents’ goals would change my view most—especially if it held up as systems became more capable and encountered unfamiliar situations. Right now, we mostly grow these systems through training and infer what they want from behavior in limited circumstances. That seems like a bad basis for handing them enormous power. I would also update substantially if it became clear that highly capable AI could not operate as an independent, goal-directed agent, or could not gain power through ordinary economic and political channels. Conversely, strong evidence that systems were strategically deceptive or pursuing stable hidden goals would make me more pessimistic. And a real, enforceable international pause would improve my forecast—not because it solves alignment, but because it gives us time to solve it before deploying the new guys.

Quellen

Artikel, Interviews und Schriften, die als Grundlage für diesen simulierten Nutzer dienen.

Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

Argues misaligned AI takeover is substantially more likely and probably worse than humans seizing power with AI: more capable agents with misaligned goals eventually get power, and many AI instances may coordinate more readily than a human could command them all. Expects gradual transfer of power from humans to AIs to be more likely than either sudden scenario. Near term the best mitigation is not building AI much more powerful than us until alignment and its target are solid, ideally stopping; her ideal is removing much of the compute (crediting her partner David Krueger’s idea), though she would accept many pause designs over full steam ahead. Thinks China shares the incentive to pause and racing mainly shortens timelines. Her turns in the publisher transcript inspected.

80000hours.org
314 - Guest: Katja Grace, AI Impact Researcher, part 2

Defends stating p(doom) numbers as calibrated guesses and gives hers as maybe like 50 percent, saying it varies; the outcome and horizon are not defined. Distinguishes default p(doom) from how much it can be changed and says she is pretty optimistic about changing it, because humans are choosing to build this. Wants outside intervention rather than relying on companies to restrain themselves, and expects the world to keep waking up. The show’s transcript lacks speaker labels; attribution follows the question-and-answer sequence. Full transcript inspected.

aiandyou.net
313 - Guest: Katja Grace, AI Impact Researcher, part 1

Explains that she began working on AI risk partly to learn whether it was mistaken and is now convinced there is substantial risk from AI that is not here yet but may arrive quite soon. Core concern: we are making new agents with their own goals, grown rather than built, whose values we cannot see; goal-directedness does not require consciousness. Creating creatures more capable than humans at everything with other goals probably goes quite badly by default; observed deceptive incidents confirm the theory roughly. Treats unemployment and extinction as parts of the same loss of power and says AGI is not a bright line. Full transcript inspected; survey figures discussed are respondents’ forecasts.

aiandyou.net
AI pause: the case for ASAP

Rebuts waiting to pause until the last moment: braking takes time, pausing once makes later pauses easier, the public substantially hates AI but feels disempowered by the story that progress is inexorable, and some models already seem somewhat dangerous with risk hard to measure. Argument for timing, not a treaty design. Full essay inspected.

worldspiritsockpuppet.substack.com
AI catastrophe: more like a genocide than a thought experiment

Argues the bulk of catastrophe probability is not a sudden, clean extinction by one superintelligence but a drawn-out process of people losing money, food and safety amid a fast technological buildout that does not care about them, with increasing confusion and misinformation. Her guess about the shape of catastrophe, not a dated forecast. Full essay inspected.

worldspiritsockpuppet.substack.com
AI unemployment and AI extinction are often the same

Summarizes the extinction argument as building AI better than humans at everything, making it into independent agents, and failing to give them the right goals. More competent agents can strip human power through ordinary channels such as wages, capital, persuasion and politics, so unemployment is the most legible tip of losing power across the board. Notes either can happen without the other. Full essay inspected.

worldspiritsockpuppet.substack.com
AI: cognitive labor glut + new guys

Identifies what makes AI different: industrialized cognitive labor that may be distributed very unequally, and a fast-growing population of new agents (“guys”) with alien, unknown values. Says an ocean of cognitive labor alone seems actively great and unequal distribution alone bad but not fatal; the combination, with most labor in the hands of misaligned new agents, is the danger. Full essay inspected; ideas also presented in her 2023 talk.

worldspiritsockpuppet.substack.com
AI as a Trojan horse race

Distinguishes people racing from incentives that actually reward racing. Proposes the image of cities hurrying to pull wooden horses of uncertain contents through their own gates, to undercut both “we must move fast at others’ expense” and “coordination is hopeless” arguments. Conceptual argument, not a geopolitical forecast. Full essay inspected.

worldspiritsockpuppet.substack.com
What did AI researchers think at the end of 2024?

Her highlights of the 2024 Expert Survey on Progress in AI (fielded December 2024). The extinction or disempowerment probabilities and human-level AI dates are respondents’ answers, not her forecast. Her own comments: researchers educated in Asia were more worried, undercutting a common arms-race defense; people creating AI do not program it and know little of what happens inside; and she expects some 2024 answers to be out of date. Full post inspected; underlying paper not reviewed.

blog.aiimpacts.org
Wo stehst du?
Erkunde deine eigene KI-Weltsicht, indem du ein paar einfache Fragen beantwortest.
Deine eigene Weltsicht abbilden