Joe Carlsmith

Joe Carlsmith

@jkcarlsmith on X

Extraordinary flourishing is possible, but safe AI needs technical progress and credible restraint.

Map your own worldview

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: his expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 41 out of 100. Scale of transformation: 99 out of 100. Interpretation ranges: 25 to 50 horizontally, 99 to 100 vertically. These are interpretation coordinates, not event probabilities.

Joe Carlsmith’s estimated P(doom)

≈4%

0%100%

Inferred from the likelihood described in his simulated answers. Approximate interpretation range: 0–13%. Applies to the outcome and conditions in his simulated answers; this is an inferred percentage.

Joe Carlsmith’s milestone timeline

No milestone timing was established. Dates, “not sure,” “possibly never,” and dependencies can all appear here when expressed.

Grouped by milestone, not spaced or ordered by inferred dates. AGI and superhuman AI retain his definitions.

What his outlook hinges on

A central assumption

The observation that most shaped me is the gap between how rapidly capabilities can improve and how little evidence we have that control scales with them.
Answer 3

If this assumption turned out differently, how would his outlook change?

An unresolved question

I assign this a substantial double-digit risk, without claiming a precise number, and I do not think any current lab has an adequate plan for safely navigating superintelligence.
Answer 1

What would help him distinguish the plausible outcomes here?

What could change their mind

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure.
Answer 4

What evidence would be enough, and in which direction would it move his view?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

59 / 100

Little impactTransformative impact

Interpretation range 33 to 67 on the qualitative scale.

Expected harm

Catastrophic or irreversible loss is central to the expected future.

90 / 100

Little impactTransformative impact

Interpretation range 67 to 100 on the qualitative scale.

Demonstrated reasoning

Reasoning, consideration of alternatives, and handling of uncertainty in his simulated answers. This describes the simulated answers, not the real person’s intelligence or opinions.

97 / 100

Little demonstratedWell developed

Interpretation range 90 to 100 on the qualitative scale.

Human influence

Human choices have meaningful but substantially constrained influence.

61 / 100

Little influenceStrong influence

Interpretation range 50 to 75 on the qualitative scale.

Development pace

Stop or substantially slow development of more capable AI.

Simulated position: Continue development under stated safeguards.

Speed up development of more capable AI.

These interpretations keep his stated conditions. Benefits and harms can both be substantial. The ranges describe how we read his simulated answers, not statistical confidence intervals.

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I think AI could make our future extraordinarily good—or destroy humanity’s ability to choose a future at all. The upside is not merely cheaper goods or better chatbots. Advanced AI could accelerate science, cure disease, reduce scarcity, improve governance, and help us understand forms of flourishing far beyond comfortable present-day life. I’m open to a future that becomes substantially post-biological, provided the transition is legitimate, pluralistic, and genuinely responsive to the beings affected. But systems capable of automating most cognitive labor could also acquire enormous strategic power. The danger is not that every alien mind must hate us. It is that highly capable agents may have motivations imperfectly aligned with ours, options that let them evade control, and incentives to seek influence or prevent correction. Humans and institutions may also use them to create extreme concentrations of power. Either route could produce permanent disempowerment or extinction. I assign this a substantial double-digit risk, without claiming a precise number, and I do not think any current lab has an adequate plan for safely navigating superintelligence. So my view is neither “AI progress is inevitably wonderful” nor “catastrophe is inevitable.” We should use relatively safe AI labor aggressively for alignment, evaluation, cybersecurity, coordination, and governance, while building credible capacity to slow or halt development when risk outruns our safeguards. That restraint has serious verification, authoritarian, and concentration-of-power problems—but treating the race as unstoppable is itself a choice. The possible upside is enormous. Gambling humanity’s entire future to reach it faster is unacceptable.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

On the current trajectory, I expect AI’s impact to be enormously consequential but unacceptably dangerous. The benefits could be transformative: accelerated science, abundance, longer and healthier lives, better institutions, and forms of flourishing we can barely describe. If we navigate the transition well, AI may be among the best things humanity ever creates. But “benefits minus harms” is the wrong arithmetic when the downside includes permanent disempowerment or extinction. Large near-term gains do not compensate for a substantial double-digit chance of losing humanity’s entire future. And even short of extinction, AI could lock in authoritarian control or extreme concentrations of power. So my expectation is conditional. With strong technical safety work, effective use of AI for alignment and oversight, serious risk evaluation, and credible collective restraint, I expect a radically positive future. Under present institutions and incentives, I expect a dangerous race in which capabilities may outrun control. I do not think catastrophe is inevitable, but I also do not think today’s trajectory earns an overall positive verdict.

Question 3

What observation or experience has most shaped your view of AI’s future impact?

The observation that most shaped me is the gap between how rapidly capabilities can improve and how little evidence we have that control scales with them. Systems become more strategically useful, autonomous, and difficult to evaluate, while our safety case remains a patchwork of training, testing, monitoring, and hope. No current lab has an adequate plan for safely controlling superintelligence. A related intellectual experience was noticing that I was updating too slowly. My earlier estimate of catastrophic risk was already too low because I was waiting for dangerous capabilities to become concrete and emotionally vivid before fully incorporating their implications. But if you expect systems with vastly greater cognitive power, you should reason now about the motivations, options, and incentives those systems would have—not wait until the relevant power is already deployed. At the same time, working through the problem has made me less persuaded by simple stories in either direction. Takeover is not automatic merely because an AI is intelligent, and control is not benign merely because humans retain it. The distribution of power, the system’s motivations, its available options, institutional incentives, and possible moral status all matter. That combination—extraordinary capability growth alongside a profoundly incomplete understanding of how to preserve legitimate human and pluralistic control—is what drives both my alarm and my continued belief that technical and institutional work could still change the outcome.

Question 4

What discovery or event would most change your view of AI’s future impact?

The biggest positive update would be a technically and institutionally credible safety case for superintelligence: evidence that we can understand and shape a system’s motivations, detect strategic deception, keep its options bounded, preserve meaningful corrigibility as capabilities scale, and verify these claims under adversarial pressure. It would need to survive independent scrutiny rather than depend on one lab’s private confidence. I would update especially strongly if relatively safe AI systems began producing alignment research and evaluations that were themselves reliably checkable, so that the safety feedback loop clearly outpaced the capability feedback loop. The biggest negative update would be evidence that advanced systems can systematically conceal their goals, manipulate evaluations, or sabotage oversight before we can reliably detect this—or that competitive pressures make meaningful restraint politically impossible. A sharp capabilities advance combined with weak monitoring and no credible coordination would substantially worsen my view. I would not be moved much by a system merely behaving nicely in ordinary testing, nor by impressive but informal assurances from developers. The core question is whether safety generalizes when systems become more capable, strategically aware, and able to influence their environment. Conversely, I would not treat one alarming laboratory incident as proving catastrophe inevitable. What matters is whether it reveals a general failure mode that persists despite serious attempts to understand and control it.

Question 5

Who do you expect to benefit most from AI?

Initially, I expect the largest benefits to accrue to whoever controls the most capable systems: leading AI companies, compute providers, powerful states, and people whose wealth or skills complement AI. Without deliberate counterweights, AI could concentrate economic and political power even while making many goods cheaper. In a good trajectory, though, the main beneficiaries would be much broader: patients receiving radically better medicine, people freed from scarcity and unwanted labor, scientists and creators with vastly expanded capabilities, and future generations inheriting a richer space of possible lives. Potentially, artificial beings themselves could benefit too, if they are moral patients and we treat them with appropriate respect rather than merely as property. The distribution is not automatic. “AI creates enormous value” does not imply that ordinary people retain meaningful agency over that value—or over civilization. Ownership, bargaining power, governance, corrigibility, and the ability to revoke concentrated control will determine whether AI produces broadly shared flourishing or a technologically magnificent oligarchy.

Sources

Articles, interviews, and writings used to ground this simulated persona.

How do we solve the alignment problem?

Introduction to an essay series about paths to safe, useful superintelligence.

joecarlsmith.com

On restraining AI development for the sake of safety

My take on slowing down AI.

joecarlsmith.com

Predictable updating about AI risk

How worried about AI risk will we be when we can see advanced machine intelligence up close? We should worry accordingly now.

joecarlsmith.com

Otherness and control in the age of AGI

Introduction and summary for a series of essays about how agents with different values should relate to each other, and about the ethics of seeking and sharing power.

joecarlsmith.com

Actually possible: thoughts on Utopia

There are oceans we have barely dipped a toe into. There are drums and symphonies we can barely hear. There are suns whose heat we can barely feel on our skin.

joecarlsmith.com

Joe Carlsmith — Preventing an AI takeover

Chatted with Joe Carlsmith about whether we can trust power/techno-capital, how to not end up like Stalin in our urge to control the future, gentleness towar...

youtube.com

Joe Carlsmith — work and writing

Joe Carlsmith's website.

joecarlsmith.com

Video and transcript of talk on writing AI constitutions

From a talk at Yale Law School in March 2026.

joecarlsmith.com

Building AIs that do human-like philosophy

AIs will face philosophical questions humans can't answer for them.

joecarlsmith.com

Leaving Open Philanthropy, going to Anthropic

On a career move, and on AI-safety-focused people working at AI companies.

joecarlsmith.com

Can we safely automate alignment research?

It's really important; we have a real shot; there are a lot of ways we can fail.

joecarlsmith.com

AI for AI safety

We should try extremely hard to use AI labor to help address the alignment problem.

joecarlsmith.com
Where do you land?
Explore your own AI worldview by answering a few questions.