问题 1
Rob Bensinger
x.com/robbensingerMIRI writer who argues superhuman AI built with current methods would be too dangerous and calls for an international halt to the race to build it.
AI将如何改变世界?
横向:他表达的 Doom–Bloom 前景看法。 纵向:变革程度。
Doom–Bloom:100 中的 4。变革程度:100 中的 97。解读范围:横向为 0 至 9,纵向为 92 至 100。这些是解读坐标,而不是事件概率。
≈72%
根据他的模拟回答推断,并非他们给出的数字。 合理范围:57–84%。
一个核心假设
As systems become more capable and agentic—better at planning, persisting, and routing around obstacles—the cost of getting those goals slightly wrong becomes catastrophic.回答 1
如果这个假设实际并非如此,他的展望会如何变化?
什么可能使其改变看法
The biggest update would be a real, legible theory of alignment: one that lets us understand and reliably control the internal goals of systems smarter than us, rather than merely patching their visible behavior.回答 4
什么证据才足够,又会让他的观点朝哪个方向转变?
更多详情
仍有几种解读是合理的:即使高级AI出现,预计也几乎不会产生积极影响。 / 预计将带来显著益处,但受到重要条件或分配方面的限制。 / 预计收益有限,或仅分布在较小范围内。
34 / 100
在定性尺度上,解读范围为 0 到 67。
灾难性或不可逆的损失是预期未来的核心。
100 / 100
在定性尺度上,解读范围为 100 到 100。
人类的选择可以大幅改变AI的发展轨迹。
69 / 100
在定性尺度上,解读范围为 50 到 100。
预计AI仍将是能力有限的工具。
预计AI将在大多数认知工作中达到人类水平。
模拟位置:预计AI将在认知工作中大幅超越人类。
模拟位置:停止或大幅放缓开发能力更强的AI。
在落实所述保障措施的前提下继续开发。
加快开发能力更强的AI。
这些解读保留了他陈述的条件。益处和危害都可能很大。这些范围描述的是我们如何解读他的模拟回答,而不是统计置信区间。
相似的世界观
模拟世界观与 Rob Bensinger 最接近的意见领袖
模拟评估
来源
用于为此模拟用户提供事实依据的文章、访谈和著述。
Full post inspected; its subtitle says it was written August 7 and published later. Explains the world’s slow response through machine learning’s trial-and-error culture, difficulty reckoning emotionally with a new kind of entity, social risk, an online “irony mandate,” and too few senior people taking engineering ownership of the danger. Says the window for international response is plausibly closing soon, if not already closed. Frames these failures as a choice that can be reversed, not destiny. Quoted remarks by Soares, Sam Harris and Joshua Achiam are theirs.

Full comment inspected via the LessWrong API. Argues that Anthropic’s and Dario Amodei’s visible messaging leaves a large candor gap relative to what many of their own researchers believe, and criticizes Anthropic for opposing US–China coordination and pursuing recursive self-improvement. He calls OpenPhil’s bet on OpenAI a disaster, while noting he had said EA’s net effect on x-risk was probably positive but highly uncertain. He says Anthropic may or may not be slightly better than OpenAI. Quoted statements by Greenblatt, Buck and others are theirs.

Full post inspected. Proposes a simultaneous, US-brokered international halt on the race to superintelligence, enforced through the concentrated chip supply chain with monitoring and possibly kill switches. The ban would last until it is clear we can build superintelligence safely, which could mean decades, and would leave existing AI and inference largely untouched. Rebuts concerns about cost, totalitarianism, defectors and China, and argues a unilateral US halt would be counterproductive. Cites Jan Leike’s 10–90% and Dario Amodei’s 25% as others’ estimates, not his own.

Older context, with the opening sections and takeoff discussion inspected. He argues that Will MacAskill’s optimism rests on a fragile conjunction of premises, so a double-digit chance of ruin remains even if each premise looks plausible. He also argues that soft, continuous takeoff would not meaningfully improve survival odds, and that good behavior from weak AIs does not show a superintelligence would be aligned. He writes partly as a MIRI insider defending the book and quotes Yudkowsky. Newer 2026 sources take precedence for current policy specifics.

Older institutional context; the byline and opening section were inspected. States MIRI’s view that building superintelligent AI with anything like current understanding or methods has human extinction as its expected outcome, and calls for governments to halt development. Use it as the shared MIRI frame Rob helped write, not as his individual phrasing. Its numerical extinction estimate is attributed to MIRI research leadership and is not his personal P(doom).

你的立场在哪里?
描绘我的世界观