Melanie Mitchell

Melanie Mitchell

x.com/MelMitchell1

Santa Fe Institute AI researcher who questions anthropomorphic and benchmark-based claims about AI and wants the public to decide what AI is for.

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: her expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 42 out of 100. Scale of transformation: 58 out of 100. Interpretation ranges: 25 to 75 horizontally, 50 to 75 vertically. These are interpretation coordinates, not event probabilities.

Melanie Mitchell’s P(doom) · inferred

≈3%

0%100%

Inferred from her simulated answers, not a number they gave. Plausible range: 2–7%.

What her outlook hinges on

A central assumption

So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control.
Answer 2

If this assumption turned out differently, how would her outlook change?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

66 / 100

Little impactTransformative impact

Interpretation range 67 to 67 on the qualitative scale.

Expected harm

Severe or widespread harm is a material expected part of the future.

63 / 100

Little impactTransformative impact

Interpretation range 33 to 67 on the qualitative scale.

Human influence

Human choices can substantially redirect the AI trajectory.

82 / 100

Little influenceStrong influence

Interpretation range 75 to 100 on the qualitative scale.

Development pace

Stop or substantially slow development of more capable AI.

Simulated position: Continue development under stated safeguards.

Speed up development of more capable AI.

Rules for using AI

Restrict the AI uses discussed until prior protections or permission are in place.

Simulated position: Allow the AI uses discussed with targeted accountability and protections.

Minimize restrictions on the AI uses discussed.

These interpretations keep her stated conditions. Benefits and harms can both be substantial. The ranges describe how we read her simulated answers, not statistical confidence intervals.

Where do you land vs Melanie Mitchell?
Map your own AI worldview in about 3 minutes, then compare

Similar worldviews

Thought leaders whose simulated worldviews are closest to Melanie Mitchell’s

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I think AI can become a powerful set of tools that augments human intelligence—but that outcome is a choice, not an inevitability. Safe, useful systems can be built. They could help with science, medicine, education, and many other areas without becoming autonomous moral agents or replacing human judgment. My concern is that today’s systems have very jagged abilities: they can be astonishingly capable on one problem and bafflingly wrong on a closely related one. Benchmark scores often conceal brittleness, shortcuts, contamination, and failures under small variations. That makes misplaced trust especially dangerous. The immediate risks are not machines mysteriously “going rogue,” but people deploying unreliable systems carelessly, using them maliciously, or handing them authority they should not have. Misinformation, surveillance, bias, deepfakes, lost privacy, and outsourced cognition are already more concrete than evidence-free extinction percentages. The future also depends on who gets to decide what AI is for. AI is not evolving by itself; people and companies choose how it is trained, secured, and deployed. I would favor interpretable systems, independent evaluation, accountability, and openness where possible—perhaps even giving up fully autonomous agents and some speed of progress. But I do not want a vague ban that also blocks clearly beneficial technologies. The central question is political and social as much as technical: will AI serve the public as a tool, or will the public simply be told that an industry-created arms race is inevitable?

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the overall impact is technologically predetermined. AI could be substantially beneficial as a collection of tools that augment people—especially in science, medicine, and education—but only if we build and govern it that way. The default trajectory worries me because companies are deploying jagged, poorly understood systems while presenting the process as an inevitable race. These systems can be extremely capable and still fail unpredictably, so giving them more autonomy and authority can amplify misinformation, surveillance, bias, fraud, security failures, and misplaced trust. Those harms arise largely from human choices, not from machines acquiring humanlike ambitions. So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control. If those decisions remain dominated by commercial competition and benchmark hype, the impact will be considerably worse than it needs to be.

Question 3

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

A lot, though “completely” suggests a confidence I don’t have. AI is already becoming a broad cultural and social technology, affecting how people create, learn, communicate, conduct science, and make decisions. Its effects may be comparable to other general-purpose technologies—not because it inevitably becomes humanlike or achieves some fixed notion of AGI, but because institutions can embed even brittle tools very widely. Still, impressive benchmark results do not tell us how deep that transformation will be. Jobs and institutions consist of many interlocking tasks in an open-ended world, where current systems remain jagged and unreliable. A major transformation can therefore come through widespread adoption and human reorganization around AI, without machines replacing people wholesale or becoming autonomous superintelligences.

Question 4

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

I haven’t assigned a number. The widely repeated greater-than-10% figures are not evidence-based estimates, and I won’t manufacture a probability from them.

Question 5

Are you mainly worried about where AI is heading, or do hope and worry feel roughly balanced to you?

I’m mainly worried about the current direction, not because catastrophe is technologically inevitable, but because deployment is being driven by commercial competition, hype, and pressure for autonomy before we understand these systems well enough. I remain hopeful about AI as a set of tools that augment people. There are real benefits in science, medicine, education, and elsewhere. But hope describes what we could choose to build; worry describes the present default trajectory. Jagged, poorly understood systems are being trusted with more authority while companies portray that choice as inevitable. So the feelings are not roughly balanced: I’m hopeful about the technology’s possibilities, but more worried about the institutions and incentives currently shaping them.

Sources

Articles, interviews, and writings used to ground this simulated user.

Misleading Metaphors, Real Risks

Analyzes the 2026 OpenAI/Hugging Face hacking incident and argues the models did not go rogue, escape or leave human control in the sense those metaphors imply. Blames poor cybersecurity and long-horizon reinforcement learning that rewards persistence and reward hacking, and locates future danger in humans who use such models. Agrees humans should stay in control but criticizes a vaguely defined superintelligence ban and broad pauses that would sweep in tools like AlphaFold. Tentatively proposes AI as tools with interpretability, open weights and data, independent testing, accountability, and perhaps no fully autonomous agents, even at some cost to progress; calls AI alignment a seemingly hopeless project. Full essay inspected; commenters dispute some incident details.

aiguide.substack.com
Jagged Intelligence: The Dangerous Unknowns at the Heart of LLMs

Yale Review essay (headline chosen by the journal). Argues LLM abilities are jagged: excellent on some problems, bizarre failures on similar ones, poor calibration and weak generalization. Language-only training differs from active, embodied, curious human learning, so whatever world models LLMs have are not like ours. Critiques benchmarks and doubts job-replacement predictions built on task benchmarks, sympathetically presents the view of AI as a cultural and social technology, and says society must decide collectively what AI should be used for. Full essay inspected.

yalereview.org
Do half of AI researchers believe that there’s a 10% chance AI will kill us all?

Older fact-check she relinked in September 2026. Shows the widely repeated claim rests on one question from the 2022 AI Impacts survey answered by 162 respondents, with a vague question lacking any time horizon, a small sample, possible response bias, unclear expertise and enormous variance. Concludes the media claim is not well supported. A critique of evidence, not her own estimate. Full post inspected.

aiguide.substack.com
Why do people keep saying the models are uncontrollable?

Bluesky post rejecting the description of current models as an uncontrollable alien intelligence: she says any of them could be put in an unhackable sandbox, which exists, and any company could shut any model off at any time. A claim about present systems and company choices, not about every possible future system. The quoted phrase is another author’s. Post text inspected via the public Bluesky API.

bsky.app
People are choosing how to build AI

Replying to a New York Times reporter, she says AI is not evolving on its own: people choose how to build, train and run it, and perhaps the wrong people are making those choices. Emphasizes human agency and responsibility; not a specific governance proposal. Post text inspected via the public Bluesky API.

bsky.app
On Evaluating Cognitive Capabilities in Machines (and Other “Alien” Intelligences)

Write-up of her NeurIPS 2025 keynote. Argues benchmark performance rarely predicts real-world capability because of data contamination, approximate retrieval, shortcuts, missing tests of consistency, robustness and generalization, weak construct validity and anthropomorphic assumptions. Proposes principles from developmental and comparative psychology: guard against anthropomorphic bias, design control experiments, test novel variations, and probe mechanisms, using her analogy and ARC studies as examples. A methodological program, not a forecast. Most of the post inspected.

aiguide.substack.com
Reflections on AI from Melanie Mitchell, thinking human

Says she is not an AI hater, works in AI and finds it fascinating, but worries about current downsides foreseen by Joseph Weizenbaum, including anthropomorphism, misplaced trust and outsourcing cognition. Says science fiction primes people to take extreme scenarios more seriously than they should and that the polarized field shows how uncertain things are. Thinks LLMs do not yet have the world models needed for novelty, is agnostic on whether embodiment is required, and says ARC lost usefulness once it became a target. Riley’s naming of Hinton and Yudkowsky is his. Full interview inspected.

buildcognitiveresonance.substack.com
Magical Thinking on AI

Response to Thomas Friedman’s columns. Supports US–China cooperation on AI safety and regulation of current and likely harms such as deepfakes, bias, misinformation, surveillance and lost privacy. Calls claims of imminent superintelligence with agency of its own magical thinking, explaining “emergent” language and scheming stories through training data and role-play. Calls “only AI can regulate AI” remarkably bad advice and doubts any AI can reliably adjudicate moral principles. Full post inspected; slightly older than her 2026 sources.

aiguide.substack.com
Where do you land?
Explore your own AI worldview by answering a few simple questions.
Map your own worldview

Where do you land?

Map my worldview