Researcher building programmable, optimized language-model systems.

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: their expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 72 out of 100. Scale of transformation: 41 out of 100. Interpretation ranges: 50 to 75 horizontally, 8 to 67 vertically. These are interpretation coordinates, not event probabilities.

Omar Khattab’s estimated P(doom)

Not specified

There is not enough relevant evidence yet to estimate their view of catastrophic risk.

What their outlook hinges on

A central assumption

The balance will depend less on isolated model behavior than on deployment: task decomposition, verification, context management, and optimization of the complete program.
Answer 2

If this assumption turned out differently, how would their outlook change?

An unresolved question

How large the net impact becomes, or how quickly, is not something I would quantify confidently.
Answer 2

What would help them distinguish the plausible outcomes here?

What could change their mind

The biggest update would come from strong evidence about learned task decomposition.
Answer 3

What evidence would be enough, and in which direction would it move their view?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

67 / 100

Little impactTransformative impact

Interpretation range 67 to 67 on the qualitative scale.

Expected harm

Manageable or localized harms are expected.

33 / 100

Little impactTransformative impact

Interpretation range 33 to 33 on the qualitative scale.

Demonstrated reasoning

Reasoning, consideration of alternatives, and handling of uncertainty in their simulated answers. This describes the simulated answers, not the real person’s intelligence or opinions.

98 / 100

Little demonstratedWell developed

Interpretation range 95 to 100 on the qualitative scale.

Human influence

Human choices can substantially redirect the AI trajectory.

67 / 100

Little influenceStrong influence

Interpretation range 41 to 84 on the qualitative scale.

These interpretations keep their stated conditions. Benefits and harms can both be substantial. The ranges describe how we read their simulated answers, not statistical confidence intervals.

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I expect AI to create substantial value, but not simply because frontier models become uniformly superhuman. Today’s models are remarkably knowledgeable and useful, yet still brittle on broad, multidimensional work: they struggle to adapt reliably across long tasks, changing requirements, feedback, and interacting constraints. More narrow, verifiable successes will arrive, but those should not be mistaken for broad competence. The more interesting possibility is that we are systematically underusing the capabilities already present. The “mismanaged geniuses” hypothesis is that much of the limitation lies in the surrounding scaffolds: how tasks are decomposed, context is managed, intermediate results are checked, and model calls are composed. If systems can learn better decompositions rather than relying on brittle hand-written prompts, they may become much stronger at long-horizon work and scientific applications. That is an ambitious research hypothesis, not an established conclusion. So I think the future depends heavily on treating AI as programmable systems rather than isolated chat models. We should optimize complete programs against measurable objectives and evaluate safety, factuality, consistency, cost, and usefulness at that same system level. Sometimes a small specialized retrieval model will beat a much larger general model on the actual task. The central question is therefore not only how capable the next model is, but how effectively—and responsibly—we organize models into systems that can do real work.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect a substantial positive impact, driven by daily usefulness and better systems for retrieval, analysis, and complex work. But I would not equate that with models becoming broadly superhuman or reliably autonomous. Current systems remain brittle, and impressive performance on narrow, verifiable tasks can conceal failures under changing requirements or long-horizon constraints. The balance will depend less on isolated model behavior than on deployment: task decomposition, verification, context management, and optimization of the complete program. Those choices also determine many harms—factual errors, inconsistency, unsafe outputs, wasted resources, and misplaced trust. If we evaluate and optimize these properties at the system level, AI can create much more value than prompt-driven deployments suggest. How large the net impact becomes, or how quickly, is not something I would quantify confidently.

Question 3

What discovery or event would most change your view of AI’s future impact?

The biggest update would come from strong evidence about learned task decomposition. If systems could reliably discover how to break unfamiliar, long-horizon work into useful subtasks, manage context, incorporate feedback, and verify intermediate results across many domains, I would become substantially more optimistic about broad scientific and economic impact. That would support the hypothesis that today’s models are often limited by poor scaffolding rather than missing core capability. The opposite result would matter just as much: repeated, careful failures showing that better programs and optimization do not overcome brittleness outside narrow, verifiable tasks. If elaborate systems still failed to adapt to changing requirements and interacting constraints, that would weaken the “mismanaged geniuses” hypothesis and suggest that major gains require fundamentally more capable models, not merely better orchestration. In either direction, I would care more about robust performance on real, multidimensional work than another benchmark record or striking narrow demonstration.

Sources

Articles, interviews, and writings used to ground this simulated user.

Where do you land?
Explore your own AI worldview by answering a few simple questions.
Map your own worldview