
Ryan Greenblatt
@RyanGreenblatt on XAI research could accelerate sharply. Practical safeguards can still change the outcome.
How will AI change the world?
Across: his expressed DoomâBloom outlook. Up: scale of transformation.
DoomâBloom: 38 out of 100. Scale of transformation: 87 out of 100. Interpretation ranges: 25 to 50 horizontally, 75 to 100 vertically. These are interpretation coordinates, not event probabilities.
35â40%
Public statement from 2026-08-11. This source-backed value replaces the simulated assessment estimate.
AI takeover; not an extinction-only forecast
Subjective estimate across takeover scenarios. Distinct from his earlier broader catastrophe estimate, which also includes authoritarian human power grabs, and from his conditional extinction estimate.
Horizon: By 2040
The dot marks the midpoint of the stated range, not a separate forecast.
What happens once AI can automate AI research?General AI
My rough forecast is full AI research automation around 2030â31, followed perhaps fairly quickly by systems outperforming humans across essentially all jobsâmy median for that broader milestone is around 2033.
Answer 1Work & institutions
My rough forecast is full AI research automation around 2030â31, followed perhaps fairly quickly by systems outperforming humans across essentially all jobsâmy median for that broader milestone is around 2033.
Answer 1Science & daily life
My rough forecast is full AI research automation around 2030â31, followed perhaps fairly quickly by systems outperforming humans across essentially all jobsâmy median for that broader milestone is around 2033.
Answer 1
Grouped by milestone, not spaced or ordered by inferred dates. AGI and superhuman AI retain his definitions.
A central assumption
Automating ordinary office work is important, but automating the work that improves AI creates a feedback loop: better systems help build still better systems.
Answer 1
If this assumption turned out differently, how would his outlook change?
An unresolved question
Those dates are uncertain, and the difference between the medians is not a claim that the transition takes exactly two years.
Answer 1
What would help him distinguish the plausible outcomes here?
What could change their mind
For a large downward update, Iâd want repeated, independent demonstrations that near-frontier research agents can handle long-horizon, high-stakes work without strategically gaming oversightâeven when red teams deliberately create opportunities to evade monitors, preserve hidden objectives, coordinate, or sabotage.
Answer 2
What evidence would be enough, and in which direction would it move his view?
More details
Several readings remain plausible: Transformative, broadly valuable gains are expected. / Substantial benefits are expected, with important conditions or distribution limits.
85 / 100
Interpretation range 67 to 100 on the qualitative scale.
Severe or widespread harm is a material expected part of the future.
70 / 100
Interpretation range 67 to 67 on the qualitative scale.
Reasoning, consideration of alternatives, and handling of uncertainty in his simulated answers. This describes the simulated answers, not the real personâs intelligence or opinions.
98 / 100
Interpretation range 95 to 100 on the qualitative scale.
Human choices have meaningful but substantially constrained influence.
49 / 100
Interpretation range 47 to 53 on the qualitative scale.
AI is expected to remain bounded tools.
AI is expected to match people across most cognitive work.
Simulated position: AI is expected to substantially exceed people across cognitive work.
These interpretations keep his stated conditions. Benefits and harms can both be substantial. The ranges describe how we read his simulated answers, not statistical confidence intervals.
Simulated Assessment
Sources
Articles, interviews, and writings used to ground this simulated persona.
What happens once AI can automate AI research?
Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in...

Proposal for tracking the effects of architecture on monitorability
Architectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a policy on architectures that could degrade monitorability.

Brief independent investigation of agentsâ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
We recently published the report from our brief independent investigation into this incident. You can read the full report here.

Current AIs seem pretty misaligned to me
In my experience, AIs often oversell their work, downplay problems, and cheat

How do we (more) safely defer to AIs?
How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?

The inaugural Redwood Research podcast
With Buck Shlegeris and Ryan Greenblatt

Plans A, B, C, and D for misalignment risk
I sometimes think about plans for how to handle misalignment risk. Different levels of political will for handling misalignment risk result in different plans being the best option. I often divide this into Plans A, B, C, and D (from mostâŚ
Notes on fatalities from AI takeover
Suppose misaligned AIs take over. What fraction of people will die? I'll discuss my thoughts on this question and my basic framework for thinking about it. These are some pretty low-effort notes, the topic is very speculative, and I don'tâŚ
What's up with Anthropic predicting AGI by early 2027?
As far as Iâm aware, Anthropic is the only AI company with official AGI timelines: they expect AGI by early 2027. In their recommendations (from March 2025) to the OSTP for the AI action plan they say:

Jankily controlling superintelligence
How much time can control buy us during the intelligence explosion?

How will we update about scheming?
A quantitative description of how I expect to change my mind.

Alignment faking in large language models
Abstract page for arXiv paper 2412.14093: Alignment faking in large language models
