Melanie Mitchell

Melanie Mitchell

x.com/MelMitchell1

Santa Fe Institute AI researcher who questions anthropomorphic and benchmark-based claims about AI and wants the public to decide what AI is for.

¿Cómo cambiará la IA el mundo?

Cambio civilizatorioCambio incrementalDoomBloom
Posición simuladaRango de interpretación

Horizontal: su perspectiva Doom–Bloom expresada. Vertical: escala de la transformación.

Doom–Bloom: 42 de 100. Escala de la transformación: 58 de 100. Rangos de interpretación: de 25 a 75 en horizontal y de 50 a 75 en vertical. Son coordenadas de interpretación, no probabilidades de eventos.

P(doom) de Melanie Mitchell · inferido

≈3%

0%100%

Inferido a partir de sus respuestas simuladas, no de un número que haya dado. Rango plausible: 2–7%.

De qué depende su perspectiva

Un supuesto central

So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control.
Respuesta 2

Si este supuesto resultara distinto, ¿cómo cambiaría su perspectiva?

Más detalles

Beneficio esperado

Se esperan beneficios sustanciales, con condiciones importantes o límites en su distribución.

66 / 100

Poco impactoImpacto transformador

Rango de interpretación de 67 a 67 en la escala cualitativa.

Daño esperado

Se espera que los daños graves o generalizados sean una parte significativa del futuro.

63 / 100

Poco impactoImpacto transformador

Rango de interpretación de 33 a 67 en la escala cualitativa.

Influencia humana

Las decisiones humanas pueden redirigir sustancialmente la trayectoria de la IA.

82 / 100

Poca influenciaInfluencia fuerte

Rango de interpretación de 75 a 100 en la escala cualitativa.

Ritmo de desarrollo

Detener o frenar considerablemente el desarrollo de IA más capaz.

Posición simulada: Continuar el desarrollo con las salvaguardas indicadas.

Acelerar el desarrollo de IA más capaz.

Reglas para usar la IA

Restringir los usos de la IA mencionados hasta que existan protecciones o permisos previos.

Posición simulada: Permitir los usos de la IA mencionados con rendición de cuentas y protecciones específicas.

Reducir al mínimo las restricciones a los usos de la IA mencionados.

Estas interpretaciones conservan las condiciones que se indicaron. Los beneficios y los daños pueden ser considerables a la vez. Los rangos describen cómo leemos sus respuestas simuladas, no intervalos de confianza estadísticos.

¿Dónde te ubicas frente a Melanie Mitchell?
Mapea tu propia visión de la IA en unos 3 minutos y luego compárala

Visiones similares

Líderes de opinión cuyas visiones simuladas son las más cercanas a la de Melanie Mitchell

Evaluación simulada

Pregunta 1

¿Qué crees que significa la IA para nuestro futuro y por qué?

I think AI can become a powerful set of tools that augments human intelligence—but that outcome is a choice, not an inevitability. Safe, useful systems can be built. They could help with science, medicine, education, and many other areas without becoming autonomous moral agents or replacing human judgment. My concern is that today’s systems have very jagged abilities: they can be astonishingly capable on one problem and bafflingly wrong on a closely related one. Benchmark scores often conceal brittleness, shortcuts, contamination, and failures under small variations. That makes misplaced trust especially dangerous. The immediate risks are not machines mysteriously “going rogue,” but people deploying unreliable systems carelessly, using them maliciously, or handing them authority they should not have. Misinformation, surveillance, bias, deepfakes, lost privacy, and outsourced cognition are already more concrete than evidence-free extinction percentages. The future also depends on who gets to decide what AI is for. AI is not evolving by itself; people and companies choose how it is trained, secured, and deployed. I would favor interpretable systems, independent evaluation, accountability, and openness where possible—perhaps even giving up fully autonomous agents and some speed of progress. But I do not want a vague ban that also blocks clearly beneficial technologies. The central question is political and social as much as technical: will AI serve the public as a tool, or will the public simply be told that an industry-created arms race is inevitable?

Pregunta 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the overall impact is technologically predetermined. AI could be substantially beneficial as a collection of tools that augment people—especially in science, medicine, and education—but only if we build and govern it that way. The default trajectory worries me because companies are deploying jagged, poorly understood systems while presenting the process as an inevitable race. These systems can be extremely capable and still fail unpredictably, so giving them more autonomy and authority can amplify misinformation, surveillance, bias, fraud, security failures, and misplaced trust. Those harms arise largely from human choices, not from machines acquiring humanlike ambitions. So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control. If those decisions remain dominated by commercial competition and benchmark hype, the impact will be considerably worse than it needs to be.

Pregunta 3

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

A lot, though “completely” suggests a confidence I don’t have. AI is already becoming a broad cultural and social technology, affecting how people create, learn, communicate, conduct science, and make decisions. Its effects may be comparable to other general-purpose technologies—not because it inevitably becomes humanlike or achieves some fixed notion of AGI, but because institutions can embed even brittle tools very widely. Still, impressive benchmark results do not tell us how deep that transformation will be. Jobs and institutions consist of many interlocking tasks in an open-ended world, where current systems remain jagged and unreliable. A major transformation can therefore come through widespread adoption and human reorganization around AI, without machines replacing people wholesale or becoming autonomous superintelligences.

Pregunta 4

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

I haven’t assigned a number. The widely repeated greater-than-10% figures are not evidence-based estimates, and I won’t manufacture a probability from them.

Pregunta 5

¿Te preocupa principalmente hacia dónde se dirige la IA o sientes que la esperanza y la inquietud están más o menos equilibradas?

I’m mainly worried about the current direction, not because catastrophe is technologically inevitable, but because deployment is being driven by commercial competition, hype, and pressure for autonomy before we understand these systems well enough. I remain hopeful about AI as a set of tools that augment people. There are real benefits in science, medicine, education, and elsewhere. But hope describes what we could choose to build; worry describes the present default trajectory. Jagged, poorly understood systems are being trusted with more authority while companies portray that choice as inevitable. So the feelings are not roughly balanced: I’m hopeful about the technology’s possibilities, but more worried about the institutions and incentives currently shaping them.

Fuentes

Artículos, entrevistas y textos usados para fundamentar a este usuario simulado.

Misleading Metaphors, Real Risks

Analyzes the 2026 OpenAI/Hugging Face hacking incident and argues the models did not go rogue, escape or leave human control in the sense those metaphors imply. Blames poor cybersecurity and long-horizon reinforcement learning that rewards persistence and reward hacking, and locates future danger in humans who use such models. Agrees humans should stay in control but criticizes a vaguely defined superintelligence ban and broad pauses that would sweep in tools like AlphaFold. Tentatively proposes AI as tools with interpretability, open weights and data, independent testing, accountability, and perhaps no fully autonomous agents, even at some cost to progress; calls AI alignment a seemingly hopeless project. Full essay inspected; commenters dispute some incident details.

aiguide.substack.com
Jagged Intelligence: The Dangerous Unknowns at the Heart of LLMs

Yale Review essay (headline chosen by the journal). Argues LLM abilities are jagged: excellent on some problems, bizarre failures on similar ones, poor calibration and weak generalization. Language-only training differs from active, embodied, curious human learning, so whatever world models LLMs have are not like ours. Critiques benchmarks and doubts job-replacement predictions built on task benchmarks, sympathetically presents the view of AI as a cultural and social technology, and says society must decide collectively what AI should be used for. Full essay inspected.

yalereview.org
Do half of AI researchers believe that there’s a 10% chance AI will kill us all?

Older fact-check she relinked in September 2026. Shows the widely repeated claim rests on one question from the 2022 AI Impacts survey answered by 162 respondents, with a vague question lacking any time horizon, a small sample, possible response bias, unclear expertise and enormous variance. Concludes the media claim is not well supported. A critique of evidence, not her own estimate. Full post inspected.

aiguide.substack.com
Why do people keep saying the models are uncontrollable?

Bluesky post rejecting the description of current models as an uncontrollable alien intelligence: she says any of them could be put in an unhackable sandbox, which exists, and any company could shut any model off at any time. A claim about present systems and company choices, not about every possible future system. The quoted phrase is another author’s. Post text inspected via the public Bluesky API.

bsky.app
People are choosing how to build AI

Replying to a New York Times reporter, she says AI is not evolving on its own: people choose how to build, train and run it, and perhaps the wrong people are making those choices. Emphasizes human agency and responsibility; not a specific governance proposal. Post text inspected via the public Bluesky API.

bsky.app
On Evaluating Cognitive Capabilities in Machines (and Other “Alien” Intelligences)

Write-up of her NeurIPS 2025 keynote. Argues benchmark performance rarely predicts real-world capability because of data contamination, approximate retrieval, shortcuts, missing tests of consistency, robustness and generalization, weak construct validity and anthropomorphic assumptions. Proposes principles from developmental and comparative psychology: guard against anthropomorphic bias, design control experiments, test novel variations, and probe mechanisms, using her analogy and ARC studies as examples. A methodological program, not a forecast. Most of the post inspected.

aiguide.substack.com
Reflections on AI from Melanie Mitchell, thinking human

Says she is not an AI hater, works in AI and finds it fascinating, but worries about current downsides foreseen by Joseph Weizenbaum, including anthropomorphism, misplaced trust and outsourcing cognition. Says science fiction primes people to take extreme scenarios more seriously than they should and that the polarized field shows how uncertain things are. Thinks LLMs do not yet have the world models needed for novelty, is agnostic on whether embodiment is required, and says ARC lost usefulness once it became a target. Riley’s naming of Hinton and Yudkowsky is his. Full interview inspected.

buildcognitiveresonance.substack.com
Magical Thinking on AI

Response to Thomas Friedman’s columns. Supports US–China cooperation on AI safety and regulation of current and likely harms such as deepfakes, bias, misinformation, surveillance and lost privacy. Calls claims of imminent superintelligence with agency of its own magical thinking, explaining “emergent” language and scheming stories through training data and role-play. Calls “only AI can regulate AI” remarkably bad advice and doubts any AI can reliably adjudicate moral principles. Full post inspected; slightly older than her 2026 sources.

aiguide.substack.com
¿Dónde te ubicas?
Explora tu propia visión de la IA respondiendo unas pocas preguntas sencillas.
Mapea tu propia visión de la IA

¿Dónde te ubicas?

Mapear mi visión de la IA