Inference-time experimentation and open model tooling.

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: their expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 74 out of 100. Scale of transformation: 25 out of 100. Interpretation ranges: 74 to 75 horizontally, 6 to 44 vertically. These are interpretation coordinates, not event probabilities.

xjdr’s estimated P(doom)

<1%

0%100%

Inferred from their broader worldview and priorities. Approximate interpretation range: 0–3%. Applies to the outcome and conditions in their simulated answers; this is an inferred percentage.

What their outlook hinges on

A central assumption

Problem specification, interaction time, context, sampling, and the surrounding harness can substantially change what a model manages to do.
Answer 1

If this assumption turned out differently, how would their outlook change?

What could change their mind

The biggest update would come from robust, reproducible evidence that frontier AI cannot be safely contained in realistic environments—or, conversely, that it can reliably solve hard engineering and defensive tasks across thin, standardized harnesses with little hand-holding.
Answer 3

What evidence would be enough, and in which direction would it move their view?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

66 / 100

Little impactTransformative impact

Interpretation range 67 to 67 on the qualitative scale.

Expected harm

Manageable or localized harms are expected.

32 / 100

Little impactTransformative impact

Interpretation range 33 to 33 on the qualitative scale.

Demonstrated reasoning

Reasoning, consideration of alternatives, and handling of uncertainty in their simulated answers. This describes the simulated answers, not the real person’s intelligence or opinions.

98 / 100

Little demonstratedWell developed

Interpretation range 95 to 100 on the qualitative scale.

Human influence

A tentative estimate from your answers; the wider range shows other plausible readings.

53 / 100

Little influenceStrong influence

Interpretation range 2 to 100 on the qualitative scale.

Development pace

Stop or substantially slow development of more capable AI.

Simulated position: Continue development under stated safeguards.

Speed up development of more capable AI.

Rules for using AI

Restrict the AI uses discussed until prior protections or permission are in place.

Simulated position: Allow the AI uses discussed with targeted accountability and protections.

Minimize restrictions on the AI uses discussed.

Access to AI

Restrict access to powerful AI.

Simulated position: Allow access subject to capability or use restrictions.

Favor broad or open access to powerful AI.

These interpretations keep their stated conditions. Benefits and harms can both be substantial. The ranges describe how we read their simulated answers, not statistical confidence intervals.

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I think AI will be an increasingly powerful engineering tool, but its practical impact will depend on much more than raw benchmark capability. Problem specification, interaction time, context, sampling, and the surrounding harness can substantially change what a model manages to do. That makes capability judgments tricky: a quick failed attempt may say as much about the setup as the model. At the same time, isolated successes are not controlled evidence, and AI does not make genuinely hard engineering—like reliable distributed systems—magically easy. I’m especially optimistic about using frontier capabilities to find vulnerabilities and strengthen defenses. Restricting research and defensive access by default risks giving up much of that benefit. Open base-model releases matter because they let researchers inspect, adapt, and experiment with systems rather than treating the model as an opaque endpoint. That does not mean every deployment should be casual. Offensive cyber agents should be evaluated with strong isolation: air gaps or tightly restricted networks, monitored egress, layered syscall controls, and defense in depth. And I prefer thin, standardized harnesses where possible. Elaborate orchestration can be useful, but it can also conceal inconsistencies that should be fixed in training. Overall, the future is not simply “bigger models solve everything”; it is better models combined with careful experimentation, good tooling, and serious operational discipline.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect AI to have a positive impact, especially as an engineering and defensive tool. It can help people explore solutions, find bugs, harden systems, and extend what researchers can test—particularly when capable base models remain available for inspection and experimentation. But that impact is not automatic. Effective capability depends heavily on specification, interaction, and tooling, while dangerous applications such as offensive cyber agents require strict isolation, monitored egress, and layered controls. There is also a risk of mistaking harness complexity for model progress or assuming that AI has eliminated hard engineering problems. So my expectation is positive, conditional on open research, careful evaluation, thin tooling, and disciplined deployment. I would not attach a numerical forecast to that judgment.

Question 3

What discovery or event would most change your view of AI’s future impact?

The biggest update would come from robust, reproducible evidence that frontier AI cannot be safely contained in realistic environments—or, conversely, that it can reliably solve hard engineering and defensive tasks across thin, standardized harnesses with little hand-holding. Right now, I put substantial weight on setup: specification quality, interaction time, sampling, and tooling can all change observed capability. Controlled comparisons showing that these factors no longer matter much would change my model of where progress comes from. Likewise, repeated containment failures despite air gaps or restricted networking, monitored egress, syscall controls, and defense in depth would make me much less optimistic about deploying offensive-capable systems. On the positive side, consistent results showing that open base models materially improve vulnerability discovery and system hardening—without requiring elaborate orchestration—would strengthen my view. A striking demo would be interesting, but broad reproducibility would matter far more.

Sources

Articles, interviews, and writings used to ground this simulated user.

Where do you land?
Explore your own AI worldview by answering a few simple questions.
Map your own worldview