Melanie Mitchell

Melanie Mitchell

x.com/MelMitchell1

Santa Fe Institute AI researcher who questions anthropomorphic and benchmark-based claims about AI and wants the public to decide what AI is for.

AI将如何改变世界?

文明层面的变革渐进式变化DoomBloom
模拟位置解读范围

横向:她表达的 Doom–Bloom 前景看法。 纵向:变革程度。

Doom–Bloom:100 中的 42。变革程度:100 中的 58。解读范围:横向为 25 至 75,纵向为 50 至 75。这些是解读坐标,而不是事件概率。

Melanie Mitchell的 P(doom) · 推断

≈3%

0%100%

根据她的模拟回答推断,并非他们给出的数字。 合理范围:2–7%。

她的展望取决于什么

一个核心假设

So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control.
回答 2

如果这个假设实际并非如此,她的展望会如何变化?

更多详情

预期益处

预计将带来显著益处,但受到重要条件或分配方面的限制。

66 / 100

影响小变革性影响

在定性尺度上,解读范围为 67 到 67。

预期危害

严重或广泛的危害预计将是未来不可忽视的一部分。

63 / 100

影响小变革性影响

在定性尺度上,解读范围为 33 到 67。

人类影响力

人类的选择可以大幅改变AI的发展轨迹。

82 / 100

影响力小影响力强

在定性尺度上,解读范围为 75 到 100。

发展速度

停止或大幅放缓开发能力更强的AI。

模拟位置:在落实所述保障措施的前提下继续开发。

加快开发能力更强的AI。

AI使用规则

在事先落实保护措施或获得许可之前,限制所讨论的AI用途。

模拟位置:允许所讨论的AI用途,同时实施有针对性的问责与保护措施。

尽量减少对所讨论AI用途的限制。

这些解读保留了她陈述的条件。益处和危害都可能很大。这些范围描述的是我们如何解读她的模拟回答,而不是统计置信区间。

与Melanie Mitchell相比,你的立场在哪里?
用大约3分钟描绘你自己的AI世界观,然后进行比较

相似的世界观

模拟世界观与 Melanie Mitchell 最接近的意见领袖

模拟评估

问题 1

你认为AI对我们的未来意味着什么?为什么?

I think AI can become a powerful set of tools that augments human intelligence—but that outcome is a choice, not an inevitability. Safe, useful systems can be built. They could help with science, medicine, education, and many other areas without becoming autonomous moral agents or replacing human judgment. My concern is that today’s systems have very jagged abilities: they can be astonishingly capable on one problem and bafflingly wrong on a closely related one. Benchmark scores often conceal brittleness, shortcuts, contamination, and failures under small variations. That makes misplaced trust especially dangerous. The immediate risks are not machines mysteriously “going rogue,” but people deploying unreliable systems carelessly, using them maliciously, or handing them authority they should not have. Misinformation, surveillance, bias, deepfakes, lost privacy, and outsourced cognition are already more concrete than evidence-free extinction percentages. The future also depends on who gets to decide what AI is for. AI is not evolving by itself; people and companies choose how it is trained, secured, and deployed. I would favor interpretable systems, independent evaluation, accountability, and openness where possible—perhaps even giving up fully autonomous agents and some speed of progress. But I do not want a vague ban that also blocks clearly beneficial technologies. The central question is political and social as much as technical: will AI serve the public as a tool, or will the public simply be told that an industry-created arms race is inevitable?

问题 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the overall impact is technologically predetermined. AI could be substantially beneficial as a collection of tools that augment people—especially in science, medicine, and education—but only if we build and govern it that way. The default trajectory worries me because companies are deploying jagged, poorly understood systems while presenting the process as an inevitable race. These systems can be extremely capable and still fail unpredictably, so giving them more autonomy and authority can amplify misinformation, surveillance, bias, fraud, security failures, and misplaced trust. Those harms arise largely from human choices, not from machines acquiring humanlike ambitions. So my expectation is conditional rather than a numerical forecast: AI’s benefits could outweigh its harms, but that requires public choices about its purpose, independent testing, accountability, interpretability, and meaningful human control. If those decisions remain dominated by commercial competition and benchmark hype, the impact will be considerably worse than it needs to be.

问题 3

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

A lot, though “completely” suggests a confidence I don’t have. AI is already becoming a broad cultural and social technology, affecting how people create, learn, communicate, conduct science, and make decisions. Its effects may be comparable to other general-purpose technologies—not because it inevitably becomes humanlike or achieves some fixed notion of AGI, but because institutions can embed even brittle tools very widely. Still, impressive benchmark results do not tell us how deep that transformation will be. Jobs and institutions consist of many interlocking tasks in an open-ended world, where current systems remain jagged and unreliable. A major transformation can therefore come through widespread adoption and human reorganization around AI, without machines replacing people wholesale or becoming autonomous superintelligences.

问题 4

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

I haven’t assigned a number. The widely repeated greater-than-10% figures are not evidence-based estimates, and I won’t manufacture a probability from them.

问题 5

你主要是对AI的发展方向感到担忧,还是希望与担忧大致相当?

I’m mainly worried about the current direction, not because catastrophe is technologically inevitable, but because deployment is being driven by commercial competition, hype, and pressure for autonomy before we understand these systems well enough. I remain hopeful about AI as a set of tools that augment people. There are real benefits in science, medicine, education, and elsewhere. But hope describes what we could choose to build; worry describes the present default trajectory. Jagged, poorly understood systems are being trusted with more authority while companies portray that choice as inevitable. So the feelings are not roughly balanced: I’m hopeful about the technology’s possibilities, but more worried about the institutions and incentives currently shaping them.

来源

用于为此模拟用户提供事实依据的文章、访谈和著述。

Misleading Metaphors, Real Risks

Analyzes the 2026 OpenAI/Hugging Face hacking incident and argues the models did not go rogue, escape or leave human control in the sense those metaphors imply. Blames poor cybersecurity and long-horizon reinforcement learning that rewards persistence and reward hacking, and locates future danger in humans who use such models. Agrees humans should stay in control but criticizes a vaguely defined superintelligence ban and broad pauses that would sweep in tools like AlphaFold. Tentatively proposes AI as tools with interpretability, open weights and data, independent testing, accountability, and perhaps no fully autonomous agents, even at some cost to progress; calls AI alignment a seemingly hopeless project. Full essay inspected; commenters dispute some incident details.

aiguide.substack.com
Jagged Intelligence: The Dangerous Unknowns at the Heart of LLMs

Yale Review essay (headline chosen by the journal). Argues LLM abilities are jagged: excellent on some problems, bizarre failures on similar ones, poor calibration and weak generalization. Language-only training differs from active, embodied, curious human learning, so whatever world models LLMs have are not like ours. Critiques benchmarks and doubts job-replacement predictions built on task benchmarks, sympathetically presents the view of AI as a cultural and social technology, and says society must decide collectively what AI should be used for. Full essay inspected.

yalereview.org
Do half of AI researchers believe that there’s a 10% chance AI will kill us all?

Older fact-check she relinked in September 2026. Shows the widely repeated claim rests on one question from the 2022 AI Impacts survey answered by 162 respondents, with a vague question lacking any time horizon, a small sample, possible response bias, unclear expertise and enormous variance. Concludes the media claim is not well supported. A critique of evidence, not her own estimate. Full post inspected.

aiguide.substack.com
Why do people keep saying the models are uncontrollable?

Bluesky post rejecting the description of current models as an uncontrollable alien intelligence: she says any of them could be put in an unhackable sandbox, which exists, and any company could shut any model off at any time. A claim about present systems and company choices, not about every possible future system. The quoted phrase is another author’s. Post text inspected via the public Bluesky API.

bsky.app
People are choosing how to build AI

Replying to a New York Times reporter, she says AI is not evolving on its own: people choose how to build, train and run it, and perhaps the wrong people are making those choices. Emphasizes human agency and responsibility; not a specific governance proposal. Post text inspected via the public Bluesky API.

bsky.app
On Evaluating Cognitive Capabilities in Machines (and Other “Alien” Intelligences)

Write-up of her NeurIPS 2025 keynote. Argues benchmark performance rarely predicts real-world capability because of data contamination, approximate retrieval, shortcuts, missing tests of consistency, robustness and generalization, weak construct validity and anthropomorphic assumptions. Proposes principles from developmental and comparative psychology: guard against anthropomorphic bias, design control experiments, test novel variations, and probe mechanisms, using her analogy and ARC studies as examples. A methodological program, not a forecast. Most of the post inspected.

aiguide.substack.com
Reflections on AI from Melanie Mitchell, thinking human

Says she is not an AI hater, works in AI and finds it fascinating, but worries about current downsides foreseen by Joseph Weizenbaum, including anthropomorphism, misplaced trust and outsourcing cognition. Says science fiction primes people to take extreme scenarios more seriously than they should and that the polarized field shows how uncertain things are. Thinks LLMs do not yet have the world models needed for novelty, is agnostic on whether embodiment is required, and says ARC lost usefulness once it became a target. Riley’s naming of Hinton and Yudkowsky is his. Full interview inspected.

buildcognitiveresonance.substack.com
Magical Thinking on AI

Response to Thomas Friedman’s columns. Supports US–China cooperation on AI safety and regulation of current and likely harms such as deepfakes, bias, misinformation, surveillance and lost privacy. Calls claims of imminent superintelligence with agency of its own magical thinking, explaining “emergent” language and scheming stories through training data and role-play. Calls “only AI can regulate AI” remarkably bad advice and doubts any AI can reliably adjudicate moral principles. Full post inspected; slightly older than her 2026 sources.

aiguide.substack.com
你的立场在哪里?
回答几个简单问题,探索你自己的AI世界观。
描绘你自己的世界观

你的立场在哪里?

描绘我的世界观