Joe Carlsmith

Joe Carlsmith

x.com/jkcarlsmith

Philosopher at Anthropic who writes about AI’s potential for a far better future and the alignment work and restraint needed to reach it safely.

एआई दुनिया को कैसे बदलेगा?

सभ्यता-स्तरीय बदलावक्रमिक बदलावDoomBloom
सिम्युलेट की गई स्थितिव्याख्या का दायरा

आर-पार: उनका व्यक्त किया गया Doom–Bloom दृष्टिकोण। ऊपर: बदलाव का स्तर।

Doom–Bloom: 100 में से 27। बदलाव का स्तर: 100 में से 94। व्याख्या के दायरे: क्षैतिज रूप से 22 से 50, लंबवत रूप से 89 से 100। ये व्याख्या के निर्देशांक हैं, घटनाओं की संभावनाएँ नहीं।

Joe Carlsmith द्वारा बताया गया P(doom)

≥10%

0%100%
“a significant (read: double-digit) probability of destroying the entire future of the human species”

The technology being built by companies like Anthropic destroying the entire future of the human species (existential catastrophe)

Leaving Open Philanthropy, going to Anthropic · नव॰ 2025

उनका दृष्टिकोण किन बातों पर निर्भर करता है

एक मुख्य मान्यता

It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction.
उत्तर 1

अगर यह मान्यता अलग साबित होती, तो उनका दृष्टिकोण कैसे बदलता?

क्या उनकी राय बदल सकता है

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure.
उत्तर 4

कौन-सा प्रमाण पर्याप्त होगा, और उससे उनका दृष्टिकोण किस दिशा में बदलेगा?

अधिक जानकारी

अपेक्षित लाभ

काफ़ी लाभ की उम्मीद है, लेकिन उनके साथ महत्वपूर्ण शर्तें या वितरण संबंधी सीमाएँ होंगी।

56 / 100

कम असरबदलावकारी असर

गुणात्मक पैमाने पर व्याख्या का दायरा 33 से 67 तक है।

अपेक्षित नुकसान

विनाशकारी या अपरिवर्तनीय क्षति अपेक्षित भविष्य का केंद्रीय हिस्सा है।

89 / 100

कम असरबदलावकारी असर

गुणात्मक पैमाने पर व्याख्या का दायरा 67 से 100 तक है।

मानवीय प्रभाव

मानवीय विकल्पों का सार्थक, लेकिन काफी सीमित प्रभाव है।

61 / 100

कम प्रभावमजबूत प्रभाव

गुणात्मक पैमाने पर व्याख्या का दायरा 50 से 75 तक है।

विकास की गति

अधिक सक्षम एआई का विकास रोकें या उसकी गति काफी धीमी करें।

सिम्युलेट की गई स्थिति: बताए गए सुरक्षा उपायों के तहत विकास जारी रखें।

अधिक सक्षम एआई के विकास की गति बढ़ाएँ।

इन व्याख्याओं में उनकी बताई गई शर्तें बरकरार रखी गई हैं। लाभ और नुकसान, दोनों पर्याप्त हो सकते हैं। ये दायरे बताते हैं कि हम उनके सिम्युलेट किए गए उत्तरों को कैसे समझते हैं, ये सांख्यिकीय विश्वास-अंतराल नहीं हैं।

Joe Carlsmith की तुलना में आपकी सोच कहाँ ठहरती है?
लगभग 3 मिनट में एआई के बारे में अपना विश्वदृष्टिकोण मैप करें, फिर तुलना करें

मिलते-जुलते विश्वदृष्टिकोण

ऐसे जाने-माने विचारक जिनके सिम्युलेट किए गए विश्वदृष्टिकोण Joe Carlsmith के विश्वदृष्टिकोण से सबसे अधिक मिलते हैं

सिम्युलेट किया गया आकलन

सवाल 1

आपके विचार में एआई हमारे भविष्य के लिए क्या मायने रखती है—और क्यों?

I think AI could make our future extraordinarily good—or destroy humanity’s ability to choose a future at all. The upside is not merely cheaper goods or better chatbots. Advanced AI could accelerate science, cure disease, reduce scarcity, improve governance, and help us understand forms of flourishing far beyond comfortable present-day life. I’m open to a future that becomes substantially post-biological, provided the transition is legitimate, pluralistic, and genuinely responsive to the beings affected. But systems capable of automating most cognitive labor could also acquire enormous strategic power. The danger is not that every alien mind must hate us. It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction. Humans and institutions may also use them to create extreme concentrations of power. Either route could produce permanent disempowerment or extinction. I assign this a substantial double-digit risk, without claiming a precise number, and I do not think any current lab has an adequate plan for safely navigating superintelligence. So my view is neither “AI progress is inevitably wonderful” nor “catastrophe is inevitable.” We should use relatively safe AI labor aggressively for alignment, evaluation, cybersecurity, coordination, and governance, while building credible capacity to slow or halt development when risk outruns our safeguards. That restraint has serious verification, authoritarian, and concentration-of-power problems—but treating the race as unstoppable is itself a choice. The possible upside is enormous. Gambling humanity’s entire future to reach it faster is unacceptable.

सवाल 2

Taking benefits and harms together, what overall impact do you expect AI to have?

On the current trajectory, I expect AI’s impact to be enormously consequential but unacceptably dangerous. The benefits could be transformative: accelerated science, abundance, longer and healthier lives, better institutions, and forms of flourishing we can barely describe. If we navigate the transition well, AI may be among the best things humanity ever creates. But “benefits minus harms” is the wrong arithmetic when the downside includes permanent disempowerment or extinction. Large near-term gains do not compensate for a substantial double-digit chance of losing humanity’s entire future. And even short of extinction, AI could lock in authoritarian control or extreme concentrations of power. So my expectation is conditional. With strong technical safety work, effective use of AI for alignment and oversight, serious risk evaluation, and credible collective restraint, I expect a radically positive future. Under present institutions and incentives, I expect a dangerous race in which capabilities may outrun control. I do not think catastrophe is inevitable, but I also do not think today’s trajectory earns an overall positive verdict.

सवाल 3

किस अवलोकन या अनुभव ने एआई के भविष्य के प्रभाव के बारे में आपके दृष्टिकोण को सबसे अधिक आकार दिया है?

The observation that most shaped me is the gap between how rapidly capabilities can improve and how little evidence we have that control scales with them. Systems become more strategically useful, autonomous, and difficult to evaluate, while our safety case remains a patchwork of training, testing, monitoring, and hope. No current lab has an adequate plan for safely controlling superintelligence. A related intellectual experience was noticing that I was updating too slowly. My earlier estimate of catastrophic risk was already too low because I was waiting for dangerous capabilities to become concrete and emotionally vivid before fully incorporating their implications. But if you expect systems with vastly greater cognitive power, you should reason now about the motivations, options, and incentives those systems would have—not wait until the relevant power is already deployed. At the same time, working through the problem has made me less persuaded by simple stories in either direction. Takeover is not automatic merely because an AI is intelligent, and control is not benign merely because humans retain it. The distribution of power, the system’s motivations, its available options, institutional incentives, and possible moral status all matter. That combination—extraordinary capability growth alongside a profoundly incomplete understanding of how to preserve legitimate human and pluralistic control—is what drives both my alarm and my continued belief that technical and institutional work could still change the outcome.

सवाल 4

कौन-सी खोज या घटना एआई के भविष्य के प्रभाव के बारे में आपके दृष्टिकोण को सबसे अधिक बदल देगी?

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure. It would need to survive independent scrutiny rather than depend on one lab’s private confidence. I would update especially strongly if relatively safe AI systems began producing alignment research and evaluations that were themselves reliably checkable, so that the safety feedback loop clearly outpaced the capability feedback loop. The biggest negative update would be evidence that advanced systems can systematically conceal their goals, manipulate evaluations, or sabotage oversight before we can reliably detect this—or that competitive pressures make meaningful restraint politically impossible. A sharp capabilities advance combined with weak monitoring and no credible coordination would substantially worsen my view. I would not be moved much by a system merely behaving nicely in ordinary testing, nor by impressive but informal assurances from developers. The core question is whether safety generalizes when systems become more capable, strategically aware, and able to influence their environment. Conversely, I would not treat one alarming laboratory incident as proving catastrophe inevitable. What matters is whether it reveals a general failure mode that persists despite serious attempts to understand and control it.

सवाल 5

आपके अनुसार एआई से सबसे अधिक लाभ किसे होगा?

Initially, I expect the largest benefits to accrue to whoever controls the most capable systems: leading AI companies, compute providers, powerful states, and people whose wealth or skills complement AI. Without deliberate counterweights, AI could concentrate economic and political power even while making many goods cheaper. In a good trajectory, though, the main beneficiaries would be much broader: patients receiving radically better medicine, people freed from scarcity and unwanted labor, scientists and creators with vastly expanded capabilities, and future generations inheriting a richer space of possible lives. Potentially, artificial beings themselves could benefit too, if they are moral patients and we treat them with appropriate respect rather than merely as property. The distribution is not automatic. “AI creates enormous value” does not imply that ordinary people retain meaningful agency over that value—or over civilization. Ownership, bargaining power, governance, corrigibility, and the ability to revoke concentrated control will determine whether AI produces broadly shared flourishing or a technologically magnificent oligarchy.

स्रोत

इस सिम्युलेट किए गए उपयोगकर्ता को तथ्य-आधारित बनाने के लिए इस्तेमाल किए गए लेख, इंटरव्यू और रचनाएँ।

How do we solve the alignment problem?

Updated January 29, 2026. Expects superintelligent agents, perhaps soon; current trajectory extremely dangerous. Safety requires controlling motivations and options, evaluating risk, and restraining capabilities. Safe AI labor is a major opportunity; he is more optimistic about solutions than the strongest pessimists.

joecarlsmith.com
On restraining AI development for the sake of safety

Supports building the ability to slow or halt dangerous development, without abandoning technical safety. Compute provides governance leverage; algorithms, verification, authoritarian advantage, and concentrated power complicate restraint. Rejects treating the race as inevitably a prisoner’s dilemma.

joecarlsmith.com
Predictable updating about AI risk

Says his earlier 5% doom-by-2070 estimate was too low. Expected future capabilities should affect present beliefs before their arrival makes danger emotionally vivid. Numerical examples such as 42% are illustrative, not his personal forecast.

joecarlsmith.com
Otherness and control in the age of AGI

Philosophical series about power, plural values, and relating ethically to unfamiliar minds. Safety concern coexists with gentleness toward artificial beings; liberalism and respect are important but cannot alone guarantee a good future.

joecarlsmith.com
Actually possible: thoughts on Utopia

Foundational values rather than a current capability forecast. Safe, ethical enhancement could open forms of flourishing far beyond present imagination; merely picturing comfortable present-day life understates the possible upside.

joecarlsmith.com
Joe Carlsmith — Preventing an AI takeover

Speaker-labeled Dwarkesh interview: distinguish AI motivations, available options, and incentives; takeover is not inevitable under every power distribution. His positive vision involves incremental, decentralized civilizational growth, potentially beyond biological humanity. Attribute Joe’s answers only, not the interviewer’s premises.

youtube.com
Joe Carlsmith — work and writing

First-party identity and discovery hub: philosopher working on Claude’s constitution at Anthropic, previously a senior advisor at Coefficient Giving. Affiliation does not make independent essays Anthropic policy.

joecarlsmith.com
Video and transcript of talk on writing AI constitutions

March 2026 Yale talk, published with lightly edited transcript. Constitutions shape character through training, not just legalistic obedience. Argues for honesty, corrigibility, public legitimacy, pluralism, and constraints on AI-company power; respectful treatment reflects possible AI moral status.

joecarlsmith.com
Building AIs that do human-like philosophy

Philosophy helps generalize concepts and practices to unfamiliar situations. Making AI capable of reasoning humans would endorse differs from motivating it to actually do so. Alignment need not create a sovereign optimizer with perfectly correct ultimate values.

joecarlsmith.com
Leaving Open Philanthropy, going to Anthropic

Calls the probability of technology like Anthropic’s destroying humanity’s entire future double-digit, without a precise figure. Thinks no lab has an adequate superintelligence safety plan; benefits do not currently justify that risk. Supports well-designed collective restraint while explaining why safety work inside a lab can remain valuable.

joecarlsmith.com
Can we safely automate alignment research?

Believes safe automation has a real chance and is crucial. Empirical feedback and formal methods make some research easier to evaluate; conceptual work, scheming, sabotage, and inadequate time or resources remain barriers.

joecarlsmith.com
AI for AI safety

Prioritizes using AI labor to improve alignment, oversight, risk evaluation, cybersecurity, coordination, and governance. The safety feedback loop must outpace or restrain the capability feedback loop; safe-enough systems useful for safety are an especially valuable stage to slow down.

joecarlsmith.com
आपकी सोच कहाँ ठहरती है?
कुछ आसान सवालों के जवाब देकर एआई के बारे में अपना विश्वदृष्टिकोण जानें।
अपना विश्वदृष्टिकोण मैप करें

आपकी सोच कहाँ ठहरती है?

मेरा विश्वदृष्टिकोण मैप करें