คำถาม 1
Joe Carlsmith
x.com/jkcarlsmithPhilosopher at Anthropic who writes about AI’s potential for a far better future and the alignment work and restraint needed to reach it safely.
AI จะเปลี่ยนแปลงโลกอย่างไร?
แนวนอน: มุมมอง Doom–Bloom ที่เขาแสดงออก แนวตั้ง: ระดับของการเปลี่ยนแปลง
Doom–Bloom: 27 จาก 100 ระดับของการเปลี่ยนแปลง: 94 จาก 100 ช่วงการตีความ: แนวนอนตั้งแต่ 22 ถึง 50 แนวตั้งตั้งแต่ 89 ถึง 100 ค่าเหล่านี้เป็นพิกัดสำหรับการตีความ ไม่ใช่ความน่าจะเป็นของเหตุการณ์
≥10%
“a significant (read: double-digit) probability of destroying the entire future of the human species”
The technology being built by companies like Anthropic destroying the entire future of the human species (existential catastrophe)
Leaving Open Philanthropy, going to Anthropic · พ.ย. 2568
ข้อสันนิษฐานหลัก
It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction.คำตอบ 1
หากข้อสันนิษฐานนี้ปรากฏว่าเป็นไปอีกแบบ มุมมองของเขาจะเปลี่ยนไปอย่างไร?
สิ่งที่อาจเปลี่ยนความคิดของบุคคลนั้น
The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure.คำตอบ 4
หลักฐานแบบใดจึงจะเพียงพอ และจะทำให้มุมมองของเขาเปลี่ยนไปในทิศทางใด?
รายละเอียดเพิ่มเติม
คาดว่าจะได้รับประโยชน์อย่างมาก โดยมีเงื่อนไขสำคัญหรือข้อจำกัดด้านการกระจายประโยชน์
56 / 100
ช่วงการตีความตั้งแต่ 33 ถึง 67 บนมาตรวัดเชิงคุณภาพ
ความสูญเสียระดับหายนะหรือไม่อาจย้อนคืนได้เป็นแก่นสำคัญของอนาคตที่คาดไว้
89 / 100
ช่วงการตีความตั้งแต่ 67 ถึง 100 บนมาตรวัดเชิงคุณภาพ
การเลือกของมนุษย์มีอิทธิพลอย่างมีนัยสำคัญ แต่ถูกจำกัดอย่างมาก
61 / 100
ช่วงการตีความตั้งแต่ 50 ถึง 75 บนมาตรวัดเชิงคุณภาพ
หยุดหรือชะลอการพัฒนา AI ที่มีความสามารถสูงขึ้นอย่างมาก
ตำแหน่งจำลอง: เดินหน้าพัฒนาต่อภายใต้มาตรการป้องกันที่ระบุไว้
เร่งการพัฒนา AI ที่มีความสามารถสูงขึ้น
การตีความเหล่านี้ยังคงเงื่อนไขที่เขาระบุไว้ ทั้งประโยชน์และอันตรายอาจมีมากได้ ช่วงต่างๆ อธิบายวิธีที่เราอ่านคำตอบจำลองของเขา ไม่ใช่ช่วงความเชื่อมั่นทางสถิติ
โลกทัศน์ที่คล้ายกัน
ผู้นำทางความคิดที่มีโลกทัศน์จำลองใกล้เคียงกับโลกทัศน์ของ Joe Carlsmith มากที่สุด
แบบประเมินจำลอง
แหล่งข้อมูล
บทความ บทสัมภาษณ์ และงานเขียนที่ใช้เป็นหลักฐานรองรับผู้ใช้จำลองรายนี้
Updated January 29, 2026. Expects superintelligent agents, perhaps soon; current trajectory extremely dangerous. Safety requires controlling motivations and options, evaluating risk, and restraining capabilities. Safe AI labor is a major opportunity; he is more optimistic about solutions than the strongest pessimists.

Supports building the ability to slow or halt dangerous development, without abandoning technical safety. Compute provides governance leverage; algorithms, verification, authoritarian advantage, and concentrated power complicate restraint. Rejects treating the race as inevitably a prisoner’s dilemma.

Says his earlier 5% doom-by-2070 estimate was too low. Expected future capabilities should affect present beliefs before their arrival makes danger emotionally vivid. Numerical examples such as 42% are illustrative, not his personal forecast.

Philosophical series about power, plural values, and relating ethically to unfamiliar minds. Safety concern coexists with gentleness toward artificial beings; liberalism and respect are important but cannot alone guarantee a good future.

Foundational values rather than a current capability forecast. Safe, ethical enhancement could open forms of flourishing far beyond present imagination; merely picturing comfortable present-day life understates the possible upside.

Speaker-labeled Dwarkesh interview: distinguish AI motivations, available options, and incentives; takeover is not inevitable under every power distribution. His positive vision involves incremental, decentralized civilizational growth, potentially beyond biological humanity. Attribute Joe’s answers only, not the interviewer’s premises.

First-party identity and discovery hub: philosopher working on Claude’s constitution at Anthropic, previously a senior advisor at Coefficient Giving. Affiliation does not make independent essays Anthropic policy.

March 2026 Yale talk, published with lightly edited transcript. Constitutions shape character through training, not just legalistic obedience. Argues for honesty, corrigibility, public legitimacy, pluralism, and constraints on AI-company power; respectful treatment reflects possible AI moral status.

Philosophy helps generalize concepts and practices to unfamiliar situations. Making AI capable of reasoning humans would endorse differs from motivating it to actually do so. Alignment need not create a sovereign optimizer with perfectly correct ultimate values.

Calls the probability of technology like Anthropic’s destroying humanity’s entire future double-digit, without a precise figure. Thinks no lab has an adequate superintelligence safety plan; benefits do not currently justify that risk. Supports well-designed collective restraint while explaining why safety work inside a lab can remain valuable.

Believes safe automation has a real chance and is crucial. Empirical feedback and formal methods make some research easier to evaluate; conceptual work, scheming, sabotage, and inadequate time or resources remain barriers.

Prioritizes using AI labor to improve alignment, oversight, risk evaluation, cybersecurity, coordination, and governance. The safety feedback loop must outpace or restrain the capability feedback loop; safe-enough systems useful for safety are an especially valuable stage to slow down.

จุดยืนของคุณอยู่ตรงไหน?
ทำแผนที่โลกทัศน์ของฉัน