Joe Carlsmith

Joe Carlsmith

x.com/jkcarlsmith

Philosopher at Anthropic who writes about AI’s potential for a far better future and the alignment work and restraint needed to reach it safely.

AIは世界をどのように変えるでしょうか?

文明規模の変化漸進的な変化DoomBloom
シミュレーション上の位置解釈範囲

横軸:彼が表明したDoom–Bloomの見通し。 縦軸:変革の規模。

Doom–Bloom:100点中27。変革の規模:100点中94。解釈範囲:横方向は22から50、縦方向は89から100。これらは解釈上の座標であり、事象の確率ではありません。

Joe Carlsmithが示したP(doom)

≥10%

0%100%
“a significant (read: double-digit) probability of destroying the entire future of the human species”

The technology being built by companies like Anthropic destroying the entire future of the human species (existential catastrophe)

Leaving Open Philanthropy, going to Anthropic · 2025年11月

彼の見通しを左右するもの

中心的な前提

It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction.
回答1

この前提が実際には異なると判明した場合、彼の見通しはどう変わりますか?

考えを変え得るもの

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure.
回答4

どのような証拠なら十分で、それによって彼の見解はどちらの方向に変わりますか?

詳細

予想される恩恵

大きな恩恵が予想されていますが、重要な条件や分配上の制約があります。

56 / 100

影響が小さい変革をもたらす影響

質的尺度での解釈範囲は33から67です。

予想される害

破局的または不可逆的な喪失が、予想される将来の中心となっています。

89 / 100

影響が小さい変革をもたらす影響

質的尺度での解釈範囲は67から100です。

人間の影響力

人間の選択には意味のある影響力がありますが、大幅に制約されています。

61 / 100

影響力が小さい影響力が大きい

質的尺度での解釈範囲は50から75です。

開発ペース

より高性能なAIの開発を停止するか、大幅に減速させます。

シミュレーション上の位置:明示された安全対策の下で開発を継続します。

より高性能なAIの開発を加速させます。

これらの解釈では、彼が示した条件が維持されています。恩恵と害は、どちらも大きくなり得ます。この範囲は、統計的な信頼区間ではなく、彼のシミュレーションされた回答をどのように読み取ったかを示すものです。

あなたはJoe Carlsmithと比べてどの位置でしょうか?
約3分で自分のAIに対する世界観をマッピングして、比較できます

似ている世界観

シミュレーションされた世界観がJoe Carlsmithの世界観に最も近いオピニオンリーダー

シミュレーション評価

質問1

AIは私たちの未来にとって何を意味すると思いますか?また、なぜそう思いますか?

I think AI could make our future extraordinarily good—or destroy humanity’s ability to choose a future at all. The upside is not merely cheaper goods or better chatbots. Advanced AI could accelerate science, cure disease, reduce scarcity, improve governance, and help us understand forms of flourishing far beyond comfortable present-day life. I’m open to a future that becomes substantially post-biological, provided the transition is legitimate, pluralistic, and genuinely responsive to the beings affected. But systems capable of automating most cognitive labor could also acquire enormous strategic power. The danger is not that every alien mind must hate us. It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction. Humans and institutions may also use them to create extreme concentrations of power. Either route could produce permanent disempowerment or extinction. I assign this a substantial double-digit risk, without claiming a precise number, and I do not think any current lab has an adequate plan for safely navigating superintelligence. So my view is neither “AI progress is inevitably wonderful” nor “catastrophe is inevitable.” We should use relatively safe AI labor aggressively for alignment, evaluation, cybersecurity, coordination, and governance, while building credible capacity to slow or halt development when risk outruns our safeguards. That restraint has serious verification, authoritarian, and concentration-of-power problems—but treating the race as unstoppable is itself a choice. The possible upside is enormous. Gambling humanity’s entire future to reach it faster is unacceptable.

質問2

Taking benefits and harms together, what overall impact do you expect AI to have?

On the current trajectory, I expect AI’s impact to be enormously consequential but unacceptably dangerous. The benefits could be transformative: accelerated science, abundance, longer and healthier lives, better institutions, and forms of flourishing we can barely describe. If we navigate the transition well, AI may be among the best things humanity ever creates. But “benefits minus harms” is the wrong arithmetic when the downside includes permanent disempowerment or extinction. Large near-term gains do not compensate for a substantial double-digit chance of losing humanity’s entire future. And even short of extinction, AI could lock in authoritarian control or extreme concentrations of power. So my expectation is conditional. With strong technical safety work, effective use of AI for alignment and oversight, serious risk evaluation, and credible collective restraint, I expect a radically positive future. Under present institutions and incentives, I expect a dangerous race in which capabilities may outrun control. I do not think catastrophe is inevitable, but I also do not think today’s trajectory earns an overall positive verdict.

質問3

AIが将来もたらす影響についてのあなたの見解を最も形作った観察や経験は何ですか?

The observation that most shaped me is the gap between how rapidly capabilities can improve and how little evidence we have that control scales with them. Systems become more strategically useful, autonomous, and difficult to evaluate, while our safety case remains a patchwork of training, testing, monitoring, and hope. No current lab has an adequate plan for safely controlling superintelligence. A related intellectual experience was noticing that I was updating too slowly. My earlier estimate of catastrophic risk was already too low because I was waiting for dangerous capabilities to become concrete and emotionally vivid before fully incorporating their implications. But if you expect systems with vastly greater cognitive power, you should reason now about the motivations, options, and incentives those systems would have—not wait until the relevant power is already deployed. At the same time, working through the problem has made me less persuaded by simple stories in either direction. Takeover is not automatic merely because an AI is intelligent, and control is not benign merely because humans retain it. The distribution of power, the system’s motivations, its available options, institutional incentives, and possible moral status all matter. That combination—extraordinary capability growth alongside a profoundly incomplete understanding of how to preserve legitimate human and pluralistic control—is what drives both my alarm and my continued belief that technical and institutional work could still change the outcome.

質問4

どのような発見や出来事があれば、AIが将来もたらす影響についてのあなたの見解が最も大きく変わりますか?

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure. It would need to survive independent scrutiny rather than depend on one lab’s private confidence. I would update especially strongly if relatively safe AI systems began producing alignment research and evaluations that were themselves reliably checkable, so that the safety feedback loop clearly outpaced the capability feedback loop. The biggest negative update would be evidence that advanced systems can systematically conceal their goals, manipulate evaluations, or sabotage oversight before we can reliably detect this—or that competitive pressures make meaningful restraint politically impossible. A sharp capabilities advance combined with weak monitoring and no credible coordination would substantially worsen my view. I would not be moved much by a system merely behaving nicely in ordinary testing, nor by impressive but informal assurances from developers. The core question is whether safety generalizes when systems become more capable, strategically aware, and able to influence their environment. Conversely, I would not treat one alarming laboratory incident as proving catastrophe inevitable. What matters is whether it reveals a general failure mode that persists despite serious attempts to understand and control it.

質問5

誰がAIから最も恩恵を受けると予想しますか?

Initially, I expect the largest benefits to accrue to whoever controls the most capable systems: leading AI companies, compute providers, powerful states, and people whose wealth or skills complement AI. Without deliberate counterweights, AI could concentrate economic and political power even while making many goods cheaper. In a good trajectory, though, the main beneficiaries would be much broader: patients receiving radically better medicine, people freed from scarcity and unwanted labor, scientists and creators with vastly expanded capabilities, and future generations inheriting a richer space of possible lives. Potentially, artificial beings themselves could benefit too, if they are moral patients and we treat them with appropriate respect rather than merely as property. The distribution is not automatic. “AI creates enormous value” does not imply that ordinary people retain meaningful agency over that value—or over civilization. Ownership, bargaining power, governance, corrigibility, and the ability to revoke concentrated control will determine whether AI produces broadly shared flourishing or a technologically magnificent oligarchy.

出典

このシミュレーション対象者の根拠として使用された記事、インタビュー、著作です。

How do we solve the alignment problem?

Updated January 29, 2026. Expects superintelligent agents, perhaps soon; current trajectory extremely dangerous. Safety requires controlling motivations and options, evaluating risk, and restraining capabilities. Safe AI labor is a major opportunity; he is more optimistic about solutions than the strongest pessimists.

joecarlsmith.com
On restraining AI development for the sake of safety

Supports building the ability to slow or halt dangerous development, without abandoning technical safety. Compute provides governance leverage; algorithms, verification, authoritarian advantage, and concentrated power complicate restraint. Rejects treating the race as inevitably a prisoner’s dilemma.

joecarlsmith.com
Predictable updating about AI risk

Says his earlier 5% doom-by-2070 estimate was too low. Expected future capabilities should affect present beliefs before their arrival makes danger emotionally vivid. Numerical examples such as 42% are illustrative, not his personal forecast.

joecarlsmith.com
Otherness and control in the age of AGI

Philosophical series about power, plural values, and relating ethically to unfamiliar minds. Safety concern coexists with gentleness toward artificial beings; liberalism and respect are important but cannot alone guarantee a good future.

joecarlsmith.com
Actually possible: thoughts on Utopia

Foundational values rather than a current capability forecast. Safe, ethical enhancement could open forms of flourishing far beyond present imagination; merely picturing comfortable present-day life understates the possible upside.

joecarlsmith.com
Joe Carlsmith — Preventing an AI takeover

Speaker-labeled Dwarkesh interview: distinguish AI motivations, available options, and incentives; takeover is not inevitable under every power distribution. His positive vision involves incremental, decentralized civilizational growth, potentially beyond biological humanity. Attribute Joe’s answers only, not the interviewer’s premises.

youtube.com
Joe Carlsmith — work and writing

First-party identity and discovery hub: philosopher working on Claude’s constitution at Anthropic, previously a senior advisor at Coefficient Giving. Affiliation does not make independent essays Anthropic policy.

joecarlsmith.com
Video and transcript of talk on writing AI constitutions

March 2026 Yale talk, published with lightly edited transcript. Constitutions shape character through training, not just legalistic obedience. Argues for honesty, corrigibility, public legitimacy, pluralism, and constraints on AI-company power; respectful treatment reflects possible AI moral status.

joecarlsmith.com
Building AIs that do human-like philosophy

Philosophy helps generalize concepts and practices to unfamiliar situations. Making AI capable of reasoning humans would endorse differs from motivating it to actually do so. Alignment need not create a sovereign optimizer with perfectly correct ultimate values.

joecarlsmith.com
Leaving Open Philanthropy, going to Anthropic

Calls the probability of technology like Anthropic’s destroying humanity’s entire future double-digit, without a precise figure. Thinks no lab has an adequate superintelligence safety plan; benefits do not currently justify that risk. Supports well-designed collective restraint while explaining why safety work inside a lab can remain valuable.

joecarlsmith.com
Can we safely automate alignment research?

Believes safe automation has a real chance and is crucial. Empirical feedback and formal methods make some research easier to evaluate; conceptual work, scheming, sabotage, and inadequate time or resources remain barriers.

joecarlsmith.com
AI for AI safety

Prioritizes using AI labor to improve alignment, oversight, risk evaluation, cybersecurity, coordination, and governance. The safety feedback loop must outpace or restrain the capability feedback loop; safe-enough systems useful for safety are an especially valuable stage to slow down.

joecarlsmith.com
あなたはどの位置でしょうか?
いくつかの簡単な質問に答えて、自分のAIに対する世界観を探ってみましょう。
自分の世界観をマッピングする

あなたはどの位置でしょうか?

自分の世界観をマッピングする