Melanie Mitchell

Melanie Mitchell

x.com/MelMitchell1

Santa Fe Institute AI researcher who questions anthropomorphic and benchmark-based claims about AI and wants the public to decide what AI is for.

Como a IA mudará o mundo?

Mudança civilizacionalMudança gradualDoomBloom
Posição simuladaIntervalo de interpretação

Na horizontal: a perspectiva Doom–Bloom expressa por ela. Para cima: escala da transformação.

Doom–Bloom: 42 de 100. Escala da transformação: 58 de 100. Intervalos de interpretação: 25 a 75 na horizontal, 50 a 75 na vertical. Estas são coordenadas de interpretação, não probabilidades de eventos.

P(doom) de Melanie Mitchell · inferido

≈3%

0%100%

Inferido a partir das respostas simuladas dela, não de um número que essa pessoa forneceu. Intervalo plausível: 2–7%.

Do que a perspectiva dela depende

Uma premissa central

So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control.
Resposta 2

Se essa premissa se revelasse diferente, como a perspectiva dela mudaria?

Mais detalhes

Benefícios esperados

Esperam-se benefícios substanciais, com condições importantes ou limites de distribuição.

66 / 100

Pouco impactoImpacto transformador

Intervalo de interpretação de 67 a 67 na escala qualitativa.

Danos esperados

Danos graves ou generalizados são uma parte relevante do futuro esperado.

63 / 100

Pouco impactoImpacto transformador

Intervalo de interpretação de 33 a 67 na escala qualitativa.

Influência humana

As escolhas humanas podem redirecionar substancialmente a trajetória da IA.

82 / 100

Pouca influênciaForte influência

Intervalo de interpretação de 75 a 100 na escala qualitativa.

Ritmo de desenvolvimento

Interromper ou desacelerar substancialmente o desenvolvimento de uma IA mais capaz.

Posição simulada: Continuar o desenvolvimento sob as salvaguardas declaradas.

Acelerar o desenvolvimento de uma IA mais capaz.

Regras para o uso da IA

Restringir os usos da IA discutidos até que proteções prévias ou uma autorização estejam em vigor.

Posição simulada: Permitir os usos da IA discutidos com responsabilização e proteções específicas.

Minimizar as restrições aos usos da IA discutidos.

Estas interpretações mantêm as condições que ela declarou. Tanto os benefícios quanto os danos podem ser substanciais. Os intervalos descrevem como interpretamos as respostas simuladas dela, não intervalos de confiança estatística.

Onde você se situa em relação a Melanie Mitchell?
Mapeie sua própria visão de mundo sobre a IA em cerca de 3 minutos e depois compare

Visões de mundo semelhantes

Líderes de opinião cujas visões de mundo simuladas são mais próximas da visão de Melanie Mitchell

Avaliação simulada

Pergunta 1

O que você acha que a IA significa para o nosso futuro — e por quê?

I think AI can become a powerful set of tools that augments human intelligence—but that outcome is a choice, not an inevitability. Safe, useful systems can be built. They could help with science, medicine, education, and many other areas without becoming autonomous moral agents or replacing human judgment. My concern is that today’s systems have very jagged abilities: they can be astonishingly capable on one problem and bafflingly wrong on a closely related one. Benchmark scores often conceal brittleness, shortcuts, contamination, and failures under small variations. That makes misplaced trust especially dangerous. The immediate risks are not machines mysteriously “going rogue,” but people deploying unreliable systems carelessly, using them maliciously, or handing them authority they should not have. Misinformation, surveillance, bias, deepfakes, lost privacy, and outsourced cognition are already more concrete than evidence-free extinction percentages. The future also depends on who gets to decide what AI is for. AI is not evolving by itself; people and companies choose how it is trained, secured, and deployed. I would favor interpretable systems, independent evaluation, accountability, and openness where possible—perhaps even giving up fully autonomous agents and some speed of progress. But I do not want a vague ban that also blocks clearly beneficial technologies. The central question is political and social as much as technical: will AI serve the public as a tool, or will the public simply be told that an industry-created arms race is inevitable?

Pergunta 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the overall impact is technologically predetermined. AI could be substantially beneficial as a collection of tools that augment people—especially in science, medicine, and education—but only if we build and govern it that way. The default trajectory worries me because companies are deploying jagged, poorly understood systems while presenting the process as an inevitable race. These systems can be extremely capable and still fail unpredictably, so giving them more autonomy and authority can amplify misinformation, surveillance, bias, fraud, security failures, and misplaced trust. Those harms arise largely from human choices, not from machines acquiring humanlike ambitions. So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control. If those decisions remain dominated by commercial competition and benchmark hype, the impact will be considerably worse than it needs to be.

Pergunta 3

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

A lot, though “completely” suggests a confidence I don’t have. AI is already becoming a broad cultural and social technology, affecting how people create, learn, communicate, conduct science, and make decisions. Its effects may be comparable to other general-purpose technologies—not because it inevitably becomes humanlike or achieves some fixed notion of AGI, but because institutions can embed even brittle tools very widely. Still, impressive benchmark results do not tell us how deep that transformation will be. Jobs and institutions consist of many interlocking tasks in an open-ended world, where current systems remain jagged and unreliable. A major transformation can therefore come through widespread adoption and human reorganization around AI, without machines replacing people wholesale or becoming autonomous superintelligences.

Pergunta 4

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

I haven’t assigned a number. The widely repeated greater-than-10% figures are not evidence-based estimates, and I won’t manufacture a probability from them.

Pergunta 5

Você está principalmente preocupado com o rumo da IA ou esperança e preocupação lhe parecem estar mais ou menos equilibradas?

I’m mainly worried about the current direction, not because catastrophe is technologically inevitable, but because deployment is being driven by commercial competition, hype, and pressure for autonomy before we understand these systems well enough. I remain hopeful about AI as a set of tools that augment people. There are real benefits in science, medicine, education, and elsewhere. But hope describes what we could choose to build; worry describes the present default trajectory. Jagged, poorly understood systems are being trusted with more authority while companies portray that choice as inevitable. So the feelings are not roughly balanced: I’m hopeful about the technology’s possibilities, but more worried about the institutions and incentives currently shaping them.

Fontes

Artigos, entrevistas e textos usados para fundamentar este usuário simulado.

Misleading Metaphors, Real Risks

Analyzes the 2026 OpenAI/Hugging Face hacking incident and argues the models did not go rogue, escape or leave human control in the sense those metaphors imply. Blames poor cybersecurity and long-horizon reinforcement learning that rewards persistence and reward hacking, and locates future danger in humans who use such models. Agrees humans should stay in control but criticizes a vaguely defined superintelligence ban and broad pauses that would sweep in tools like AlphaFold. Tentatively proposes AI as tools with interpretability, open weights and data, independent testing, accountability, and perhaps no fully autonomous agents, even at some cost to progress; calls AI alignment a seemingly hopeless project. Full essay inspected; commenters dispute some incident details.

aiguide.substack.com
Jagged Intelligence: The Dangerous Unknowns at the Heart of LLMs

Yale Review essay (headline chosen by the journal). Argues LLM abilities are jagged: excellent on some problems, bizarre failures on similar ones, poor calibration and weak generalization. Language-only training differs from active, embodied, curious human learning, so whatever world models LLMs have are not like ours. Critiques benchmarks and doubts job-replacement predictions built on task benchmarks, sympathetically presents the view of AI as a cultural and social technology, and says society must decide collectively what AI should be used for. Full essay inspected.

yalereview.org
Do half of AI researchers believe that there’s a 10% chance AI will kill us all?

Older fact-check she relinked in September 2026. Shows the widely repeated claim rests on one question from the 2022 AI Impacts survey answered by 162 respondents, with a vague question lacking any time horizon, a small sample, possible response bias, unclear expertise and enormous variance. Concludes the media claim is not well supported. A critique of evidence, not her own estimate. Full post inspected.

aiguide.substack.com
Why do people keep saying the models are uncontrollable?

Bluesky post rejecting the description of current models as an uncontrollable alien intelligence: she says any of them could be put in an unhackable sandbox, which exists, and any company could shut any model off at any time. A claim about present systems and company choices, not about every possible future system. The quoted phrase is another author’s. Post text inspected via the public Bluesky API.

bsky.app
People are choosing how to build AI

Replying to a New York Times reporter, she says AI is not evolving on its own: people choose how to build, train and run it, and perhaps the wrong people are making those choices. Emphasizes human agency and responsibility; not a specific governance proposal. Post text inspected via the public Bluesky API.

bsky.app
On Evaluating Cognitive Capabilities in Machines (and Other “Alien” Intelligences)

Write-up of her NeurIPS 2025 keynote. Argues benchmark performance rarely predicts real-world capability because of data contamination, approximate retrieval, shortcuts, missing tests of consistency, robustness and generalization, weak construct validity and anthropomorphic assumptions. Proposes principles from developmental and comparative psychology: guard against anthropomorphic bias, design control experiments, test novel variations, and probe mechanisms, using her analogy and ARC studies as examples. A methodological program, not a forecast. Most of the post inspected.

aiguide.substack.com
Reflections on AI from Melanie Mitchell, thinking human

Says she is not an AI hater, works in AI and finds it fascinating, but worries about current downsides foreseen by Joseph Weizenbaum, including anthropomorphism, misplaced trust and outsourcing cognition. Says science fiction primes people to take extreme scenarios more seriously than they should and that the polarized field shows how uncertain things are. Thinks LLMs do not yet have the world models needed for novelty, is agnostic on whether embodiment is required, and says ARC lost usefulness once it became a target. Riley’s naming of Hinton and Yudkowsky is his. Full interview inspected.

buildcognitiveresonance.substack.com
Magical Thinking on AI

Response to Thomas Friedman’s columns. Supports US–China cooperation on AI safety and regulation of current and likely harms such as deepfakes, bias, misinformation, surveillance and lost privacy. Calls claims of imminent superintelligence with agency of its own magical thinking, explaining “emergent” language and scheming stories through training data and role-play. Calls “only AI can regulate AI” remarkably bad advice and doubts any AI can reliably adjudicate moral principles. Full post inspected; slightly older than her 2026 sources.

aiguide.substack.com
Onde você se situa?
Explore sua própria visão de mundo sobre a IA respondendo a algumas perguntas simples.
Mapeie sua própria visão de mundo

Onde você se situa?

Mapear minha visão de mundo