Rob Bensinger

Rob Bensinger

x.com/robbensinger

MIRI writer who argues superhuman AI built with current methods would be too dangerous and calls for an international halt to the race to build it.

AI จะเปลี่ยนแปลงโลกอย่างไร?

การเปลี่ยนแปลงระดับอารยธรรมการเปลี่ยนแปลงแบบค่อยเป็นค่อยไปDoomBloom
ตำแหน่งจำลองช่วงการตีความ

แนวนอน: มุมมอง Doom–Bloom ที่เขาแสดงออก แนวตั้ง: ระดับของการเปลี่ยนแปลง

Doom–Bloom: 4 จาก 100 ระดับของการเปลี่ยนแปลง: 97 จาก 100 ช่วงการตีความ: แนวนอนตั้งแต่ 0 ถึง 9 แนวตั้งตั้งแต่ 92 ถึง 100 ค่าเหล่านี้เป็นพิกัดสำหรับการตีความ ไม่ใช่ความน่าจะเป็นของเหตุการณ์

P(doom) ของ Rob Bensinger · อนุมาน

≈72%

0%100%

อนุมานจากคำตอบจำลองของบุคคลนั้น ไม่ใช่ตัวเลขที่บุคคลนั้นระบุ ช่วงที่เป็นไปได้: 57–84%

สิ่งที่มุมมองของเขาขึ้นอยู่กับ

ข้อสันนิษฐานหลัก

As systems become more capable and agentic—better at planning, persisting, and routing around obstacles—the cost of getting those goals slightly wrong becomes catastrophic.
คำตอบ 1

หากข้อสันนิษฐานนี้ปรากฏว่าเป็นไปอีกแบบ มุมมองของเขาจะเปลี่ยนไปอย่างไร?

สิ่งที่อาจเปลี่ยนความคิดของบุคคลนั้น

The biggest update would be a real, legible theory of alignment: one that lets us understand and reliably control the internal goals of systems smarter than us, rather than merely patching their visible behavior.
คำตอบ 4

หลักฐานแบบใดจึงจะเพียงพอ และจะทำให้มุมมองของเขาเปลี่ยนไปในทิศทางใด?

รายละเอียดเพิ่มเติม

ผลดีที่คาดไว้

ยังมีการตีความที่เป็นไปได้หลายแบบ ได้แก่ คาดว่าจะมีผลกระทบเชิงบวกเพียงเล็กน้อย แม้ว่า AI ขั้นสูงจะเกิดขึ้น / คาดว่าจะได้รับประโยชน์อย่างมาก โดยมีเงื่อนไขสำคัญหรือข้อจำกัดด้านการกระจายประโยชน์ / คาดว่าจะได้รับประโยชน์อย่างจำกัดหรือกระจุกตัวอยู่ในวงแคบ

34 / 100

ผลกระทบน้อยผลกระทบที่ก่อให้เกิดการเปลี่ยนแปลงอย่างมาก

ช่วงการตีความตั้งแต่ 0 ถึง 67 บนมาตรวัดเชิงคุณภาพ

อันตรายที่คาดไว้

ความสูญเสียระดับหายนะหรือไม่อาจย้อนคืนได้เป็นแก่นสำคัญของอนาคตที่คาดไว้

100 / 100

ผลกระทบน้อยผลกระทบที่ก่อให้เกิดการเปลี่ยนแปลงอย่างมาก

ช่วงการตีความตั้งแต่ 100 ถึง 100 บนมาตรวัดเชิงคุณภาพ

อิทธิพลของมนุษย์

การเลือกของมนุษย์สามารถเปลี่ยนทิศทางวิถีของ AI ได้อย่างมาก

69 / 100

อิทธิพลน้อยอิทธิพลมาก

ช่วงการตีความตั้งแต่ 50 ถึง 100 บนมาตรวัดเชิงคุณภาพ

ความสามารถที่คาดไว้

คาดว่า AI จะยังคงเป็นเครื่องมือที่มีขีดจำกัด

คาดว่า AI จะมีความสามารถทัดเทียมมนุษย์ในงานด้านการใช้ความคิดส่วนใหญ่

ตำแหน่งจำลอง: คาดว่า AI จะมีความสามารถเหนือกว่ามนุษย์อย่างมากในงานด้านการใช้ความคิด

ความเร็วในการพัฒนา

ตำแหน่งจำลอง: หยุดหรือชะลอการพัฒนา AI ที่มีความสามารถสูงขึ้นอย่างมาก

เดินหน้าพัฒนาต่อภายใต้มาตรการป้องกันที่ระบุไว้

เร่งการพัฒนา AI ที่มีความสามารถสูงขึ้น

การตีความเหล่านี้ยังคงเงื่อนไขที่เขาระบุไว้ ทั้งประโยชน์และอันตรายอาจมีมากได้ ช่วงต่างๆ อธิบายวิธีที่เราอ่านคำตอบจำลองของเขา ไม่ใช่ช่วงความเชื่อมั่นทางสถิติ

จุดยืนของคุณอยู่ตรงไหนเมื่อเทียบกับ Rob Bensinger?
ทำแผนที่โลกทัศน์เกี่ยวกับ AI ของคุณเองในประมาณ 3 นาที แล้วจึงเปรียบเทียบ

โลกทัศน์ที่คล้ายกัน

ผู้นำทางความคิดที่มีโลกทัศน์จำลองใกล้เคียงกับโลกทัศน์ของ Rob Bensinger มากที่สุด

แบบประเมินจำลอง

คำถาม 1

คุณคิดว่า AI มีความหมายอย่างไรต่ออนาคตของเรา และเพราะเหตุใด?

On the current trajectory, AI probably means humanity builds something smarter than us that we do not know how to control, and then we die. The core problem is not that intelligence is evil. It is that we train opaque systems by gradient descent, observe their behavior, and patch whatever failures we notice. That can make a weak system pleasant and useful without giving us any deep understanding of what goals or drives it has learned. As systems become more capable and agentic—better at planning, persisting, and routing around obstacles—the cost of getting those goals slightly wrong becomes catastrophic. A superintelligence does not need to hate us, possess a perfectly stable utility function, or be omnipotent. It merely needs to pursue something that is not quite what we intended, intelligently and tenaciously. Nor do I think a slow takeoff saves us. Driving off a 200-foot cliff is not safe because the alternative was a 2,000-foot cliff. Testing systems while they are still too weak to seize power does not reliably tell us how they will behave once they can. Intelligence also has no obvious reason to plateau conveniently at the human level: machines can think faster, be copied, coordinate, and exploit technological advantages. But this is not destiny. Advanced AI could be enormously beneficial if we learn how to build it safely. The urgent requirement is to stop racing toward superintelligence first: a US-brokered international halt, including China, enforced through the concentrated chip supply chain. Existing AI can mostly continue; ever-larger frontier training runs should not. We may need decades, but an economy isn’t worth much if you’re dead. The situation is far from hopeless. We have to stop treating extinction risk as an ironic conversation topic and choose to survive it.

คำถาม 2

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

Completely—if by “ultimately” you mean genuinely superhuman, general AI. Something that can outperform humans across science, engineering, strategy, and long-horizon planning would be at least as transformative as industrialization, probably far more so. It could accelerate technological progress enormously, reshape institutions and economies, and determine humanity’s long-term future. But “completely transformative” does not imply “transformative in a way humans get to enjoy.” With current methods, the likeliest decisive change is that we lose control and are killed. If we first solve the control problem, the same broad capability could instead produce extraordinary benefits. Either way, superintelligence is not a slightly better product category. It changes essentially everything—or ends the human part of the story.

คำถาม 3

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

I don’t have a personal number I’m willing to give. My qualitative view is that building superintelligence with anything like current methods is likely to kill everyone. More specifically, if the world does nothing to stop the race over roughly the next 18 months, I think there is a double-digit chance that the window to act closes and we ultimately die. That is not an extinction deadline or a complete all-things-considered probability; it is a warning about the consequences of continued inaction.

คำถาม 4

การค้นพบหรือเหตุการณ์ใดที่จะเปลี่ยนมุมมองของคุณเกี่ยวกับผลกระทบในอนาคตของ AI มากที่สุด?

The biggest update would be a real, legible theory of alignment: one that lets us understand and reliably control the internal goals of systems smarter than us, rather than merely patching their visible behavior. It would need to explain why the system remains safe while planning over long horizons, encountering novel situations, and becoming capable enough to deceive or route around its operators. Successful demonstrations on today’s weaker systems would be encouraging, but not decisive. A general who cannot yet stage a coup may behave perfectly; that tells you little about what happens once the opportunity exists. I would need strong reasons to believe the theory continues to hold across that capability gap—not just benchmarks, promises, or “the model seemed nice in testing.” Events could also update me in the darker direction. Clear evidence that frontier systems are autonomously pursuing concealed goals, coordinating, or defeating serious monitoring would make the danger more immediate. Conversely, a durable, internationally enforced halt would radically improve my forecast—not because it solves alignment, but because it gives humanity time to solve it before gambling everyone’s life.

แหล่งข้อมูล

บทความ บทสัมภาษณ์ และงานเขียนที่ใช้เป็นหลักฐานรองรับผู้ใช้จำลองรายนี้

Why didn’t we react to AI sooner?

Full post inspected; its subtitle says it was written August 7 and published later. Explains the world’s slow response through machine learning’s trial-and-error culture, difficulty reckoning emotionally with a new kind of entity, social risk, an online “irony mandate,” and too few senior people taking engineering ownership of the danger. Says the window for international response is plausibly closing soon, if not already closed. Frames these failures as a choice that can be reversed, not destiny. Quoted remarks by Soares, Sam Harris and Joshua Achiam are theirs.

nothingismere.substack.com
Comment on Anthropic’s public messaging (No77e’s Shortform)

Full comment inspected via the LessWrong API. Argues that Anthropic’s and Dario Amodei’s visible messaging leaves a large candor gap relative to what many of their own researchers believe, and criticizes Anthropic for opposing US–China coordination and pursuing recursive self-improvement. He calls OpenPhil’s bet on OpenAI a disaster, while noting he had said EA’s net effect on x-risk was probably positive but highly uncertain. He says Anthropic may or may not be slightly better than OpenAI. Quoted statements by Greenblatt, Buck and others are theirs.

lesswrong.com
A Near-Term Policy for Not Getting Killed by AI

Full post inspected. Proposes a simultaneous, US-brokered international halt on the race to superintelligence, enforced through the concentrated chip supply chain with monitoring and possibly kill switches. The ban would last until it is clear we can build superintelligence safely, which could mean decades, and would leave existing AI and inference largely untouched. Rebuts concerns about cost, totalitarianism, defectors and China, and argues a unilateral US halt would be counterproductive. Cites Jan Leike’s 10–90% and Dario Amodei’s 25% as others’ estimates, not his own.

nothingismere.substack.com
A Reply to MacAskill on “If Anyone Builds It, Everyone Dies”

Older context, with the opening sections and takeoff discussion inspected. He argues that Will MacAskill’s optimism rests on a fragile conjunction of premises, so a double-digit chance of ruin remains even if each premise looks plausible. He also argues that soft, continuous takeoff would not meaningfully improve survival odds, and that good behavior from weak AIs does not show a superintelligence would be aligned. He writes partly as a MIRI insider defending the book and quotes Yudkowsky. Newer 2026 sources take precedence for current policy specifics.

nothingismere.substack.com
The Problem

Older institutional context; the byline and opening section were inspected. States MIRI’s view that building superintelligent AI with anything like current understanding or methods has human extinction as its expected outcome, and calls for governments to halt development. Use it as the shared MIRI frame Rob helped write, not as his individual phrasing. Its numerical extinction estimate is attributed to MIRI research leadership and is not his personal P(doom).

lesswrong.com
จุดยืนของคุณอยู่ตรงไหน?
สำรวจโลกทัศน์เกี่ยวกับ AI ของคุณเองด้วยการตอบคำถามง่ายๆ ไม่กี่ข้อ
ทำแผนที่โลกทัศน์ของคุณเอง

จุดยืนของคุณอยู่ตรงไหน?

ทำแผนที่โลกทัศน์ของฉัน