问题 1
Katja Grace
x.com/KatjaGraceAI Impacts co-founder who surveys AI researchers about progress and risk and argues for pausing the development of AI much more capable than humans.
AI将如何改变世界?
横向:她表达的 Doom–Bloom 前景看法。 纵向:变革程度。
Doom–Bloom:100 中的 22。变革程度:100 中的 94。解读范围:横向为 17 至 50,纵向为 89 至 100。这些是解读坐标,而不是事件概率。
≈50%
“Well, it varies. I’d say maybe like 50 percent.”
AI “doom” in her discussion of the probability that current AI development destroys the world; endpoint not further defined (the host’s follow-up paraphrases it as human extinction)
314 - Guest: Katja Grace, AI Impact Researcher, part 2 · 2026年6月
一个核心假设
But the route we are taking means creating new agents—“new guys”—that pursue goals, may become better than humans at nearly everything, and whose values we cannot inspect or reliably choose.回答 1
如果这个假设实际并非如此,她的展望会如何变化?
一个尚未解决的问题
The uncertainty is less about whether sufficiently advanced AI would be transformative than whether we build it, when, and whether humans remain meaningfully in control afterward.回答 2
什么能帮助她区分这里各种合理的结果?
什么可能使其改变看法
A convincing way to inspect and reliably control advanced agents’ goals would change my view most—especially if it held up as systems became more capable and encountered unfamiliar situations.回答 4
什么证据才足够,又会让她的观点朝哪个方向转变?
更多详情
仍有几种解读是合理的:预计将带来显著益处,但受到重要条件或分配方面的限制。 / 预计将带来具有变革性且广泛有价值的收益。 / 预计收益有限,或仅分布在较小范围内。
67 / 100
在定性尺度上,解读范围为 33 到 100。
仍有几种解读是合理的:灾难性或不可逆的损失是预期未来的核心。 / 严重或广泛的危害预计将是未来不可忽视的一部分。
84 / 100
在定性尺度上,解读范围为 67 到 100。
人类的选择可以大幅改变AI的发展轨迹。
73 / 100
在定性尺度上,解读范围为 50 到 75。
预计AI仍将是能力有限的工具。
预计AI将在大多数认知工作中达到人类水平。
模拟位置:预计AI将在认知工作中大幅超越人类。
模拟位置:停止或大幅放缓开发能力更强的AI。
在落实所述保障措施的前提下继续开发。
加快开发能力更强的AI。
这些解读保留了她陈述的条件。益处和危害都可能很大。这些范围描述的是我们如何解读她的模拟回答,而不是统计置信区间。
相似的世界观
模拟世界观与 Katja Grace 最接近的意见领袖
模拟评估
来源
用于为此模拟用户提供事实依据的文章、访谈和著述。
Argues misaligned AI takeover is substantially more likely and probably worse than humans seizing power with AI: more capable agents with misaligned goals eventually get power, and many AI instances may coordinate more readily than a human could command them all. Expects gradual transfer of power from humans to AIs to be more likely than either sudden scenario. Near term the best mitigation is not building AI much more powerful than us until alignment and its target are solid, ideally stopping; her ideal is removing much of the compute (crediting her partner David Krueger’s idea), though she would accept many pause designs over full steam ahead. Thinks China shares the incentive to pause and racing mainly shortens timelines. Her turns in the publisher transcript inspected.

Defends stating p(doom) numbers as calibrated guesses and gives hers as maybe like 50 percent, saying it varies; the outcome and horizon are not defined. Distinguishes default p(doom) from how much it can be changed and says she is pretty optimistic about changing it, because humans are choosing to build this. Wants outside intervention rather than relying on companies to restrain themselves, and expects the world to keep waking up. The show’s transcript lacks speaker labels; attribution follows the question-and-answer sequence. Full transcript inspected.

Explains that she began working on AI risk partly to learn whether it was mistaken and is now convinced there is substantial risk from AI that is not here yet but may arrive quite soon. Core concern: we are making new agents with their own goals, grown rather than built, whose values we cannot see; goal-directedness does not require consciousness. Creating creatures more capable than humans at everything with other goals probably goes quite badly by default; observed deceptive incidents confirm the theory roughly. Treats unemployment and extinction as parts of the same loss of power and says AGI is not a bright line. Full transcript inspected; survey figures discussed are respondents’ forecasts.

Rebuts waiting to pause until the last moment: braking takes time, pausing once makes later pauses easier, the public substantially hates AI but feels disempowered by the story that progress is inexorable, and some models already seem somewhat dangerous with risk hard to measure. Argument for timing, not a treaty design. Full essay inspected.

Argues the bulk of catastrophe probability is not a sudden, clean extinction by one superintelligence but a drawn-out process of people losing money, food and safety amid a fast technological buildout that does not care about them, with increasing confusion and misinformation. Her guess about the shape of catastrophe, not a dated forecast. Full essay inspected.

Summarizes the extinction argument as building AI better than humans at everything, making it into independent agents, and failing to give them the right goals. More competent agents can strip human power through ordinary channels such as wages, capital, persuasion and politics, so unemployment is the most legible tip of losing power across the board. Notes either can happen without the other. Full essay inspected.

Identifies what makes AI different: industrialized cognitive labor that may be distributed very unequally, and a fast-growing population of new agents (“guys”) with alien, unknown values. Says an ocean of cognitive labor alone seems actively great and unequal distribution alone bad but not fatal; the combination, with most labor in the hands of misaligned new agents, is the danger. Full essay inspected; ideas also presented in her 2023 talk.

Distinguishes people racing from incentives that actually reward racing. Proposes the image of cities hurrying to pull wooden horses of uncertain contents through their own gates, to undercut both “we must move fast at others’ expense” and “coordination is hopeless” arguments. Conceptual argument, not a geopolitical forecast. Full essay inspected.

Her highlights of the 2024 Expert Survey on Progress in AI (fielded December 2024). The extinction or disempowerment probabilities and human-level AI dates are respondents’ answers, not her forecast. Her own comments: researchers educated in Asia were more worried, undercutting a common arms-race defense; people creating AI do not program it and know little of what happens inside; and she expects some 2024 answers to be out of date. Full post inspected; underlying paper not reviewed.

你的立场在哪里?
描绘我的世界观