Raymond Weitekamp

Raymond Weitekamp

x.com/raw_works

Reliable recursive agents and measurable outcomes.

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: their expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 69 out of 100. Scale of transformation: 44 out of 100. Interpretation ranges: 50 to 75 horizontally, 19 to 56 vertically. These are interpretation coordinates, not event probabilities.

Raymond Weitekamp’s estimated P(doom)

<1%

0%100%

Inferred from their broader worldview and priorities. Approximate interpretation range: 0–3%. Applies to the outcome and conditions in their simulated answers; this is an inferred percentage.

What their outlook hinges on

A central assumption

The biggest change would come from evidence that reliable agent behavior does—or does not—scale with better harnesses.
Answer 3

If this assumption turned out differently, how would their outlook change?

An unresolved question

The biggest change would come from evidence that reliable agent behavior does—or does not—scale with better harnesses.
Answer 3

What would help them distinguish the plausible outcomes here?

What could change their mind

If repeated, independent results showed that tool use, executable reasoning, recursive revision, testing, and bounded permissions still fail unpredictably on consequential tasks, then I would become much less optimistic about broad deployment.
Answer 3

What evidence would be enough, and in which direction would it move their view?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

67 / 100

Little impactTransformative impact

Interpretation range 67 to 67 on the qualitative scale.

Expected harm

Manageable or localized harms are expected.

41 / 100

Little impactTransformative impact

Interpretation range 33 to 67 on the qualitative scale.

Demonstrated reasoning

Reasoning, consideration of alternatives, and handling of uncertainty in their simulated answers. This describes the simulated answers, not the real person’s intelligence or opinions.

95 / 100

Little demonstratedWell developed

Interpretation range 86 to 100 on the qualitative scale.

Human influence

A tentative estimate from your answers; the wider range shows other plausible readings.

57 / 100

Little influenceStrong influence

Interpretation range 14 to 100 on the qualitative scale.

Rules for using AI

Restrict the AI uses discussed until prior protections or permission are in place.

Simulated position: Allow the AI uses discussed with targeted accountability and protections.

Minimize restrictions on the AI uses discussed.

These interpretations keep their stated conditions. Benefits and harms can both be substantial. The ranges describe how we read their simulated answers, not statistical confidence intervals.

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I think AI’s future is less about a single universally intelligent model and more about systems: models, tools, executable code, memory, verification, permissions, and feedback loops working together. A model that looks limited in a chat interface may perform substantially better when its harness lets it inspect state, run code, test hypotheses, and recursively revise its work. That means many apparent “model capabilities” are really properties of the whole system. Practically, I expect increasingly capable agents for both coding and noncoding workflows. General LLMs can handle open-ended interpretation, while smaller specialized decision models operate over compressed state and structured action spaces. That division may be more useful, controllable, and efficient than forcing one general model to do everything. But capability without reliability is not enough. The system needs measurable outcomes, tests, verification, and constrained permissions. A conversational model can sound cautious while its agent harness aggressively edits files, invokes tools, or exposes private data. Safety therefore has to be evaluated at the level where actions occur, not inferred from tone. Privacy and control will also shape which applications are viable. Sensitive personal and process-level uses often require local execution, self-hosting, or credible zero-data-retention options. So I’m optimistic about what these systems can do, but the important question is not merely how intelligent the model appears. It is whether the complete system produces useful, verifiable results without taking unacceptable liberties with data or actions.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect AI to be strongly beneficial where outcomes can be measured and actions can be verified. It should automate substantial amounts of knowledge work, improve software and operational workflows, and make specialized intelligence available locally in systems that do not need universal competence. Better harnesses—tools, tests, structured state, feedback loops, and recursive revision—can turn models into much more useful agents than chat performance alone suggests. The harms are also mostly system-level. An agent can be polite and cautious in conversation while its permissions let it delete data, expose private information, or make unchecked changes. Unreliable outputs become much more consequential once connected to tools and real-world actions. Centralized handling of sensitive personal or process data creates another serious constraint. So the net impact depends heavily on deployment architecture. Systems with bounded permissions, measurable objectives, verification, and local or privacy-preserving execution can create large practical gains. Systems optimized mainly for apparent autonomy, without corresponding reliability and control, can amplify mistakes just as effectively as they amplify competence.

Question 3

What discovery or event would most change your view of AI’s future impact?

The biggest change would come from evidence that reliable agent behavior does—or does not—scale with better harnesses. If repeated, independent results showed that tool use, executable reasoning, recursive revision, testing, and bounded permissions still fail unpredictably on consequential tasks, then I would become much less optimistic about broad deployment. That would suggest the limitation is deeper than interface or system design. Conversely, strong demonstrations of agents operating over long horizons with measurable outcomes, effective verification, controlled permissions, and genuinely private local execution would make me more optimistic. I care less about a model appearing intelligent in conversation than about complete systems producing correct, auditable results without taking unacceptable actions. The decisive event would therefore be a reproducible reliability result at the system level—not merely a new benchmark score or a more impressive chat demo.

Sources

Articles, interviews, and writings used to ground this simulated user.

Where do you land?
Explore your own AI worldview by answering a few simple questions.
Map your own worldview