Mike Taylor

Mike Taylor

x.com/hammer_mt

AI evaluations and dependable application behavior.

How will AI change the world?

Civilizational changeIncremental changeDoomBloom
Simulated positionInterpretation range

Across: their expressed Doom–Bloom outlook. Up: scale of transformation.

Doom–Bloom: 68 out of 100. Scale of transformation: 26 out of 100. Interpretation ranges: 50 to 75 horizontally, 25 to 26 vertically. These are interpretation coordinates, not event probabilities.

Mike Taylor’s estimated P(doom)

Not specified

There is not enough relevant evidence yet to estimate their view of catastrophic risk.

What their outlook hinges on

A central assumption

Work is a bundle of tasks: drafting a product brief, analyzing feedback, generating options, or critiquing a decision are easier to test and delegate than the entire job of being a product manager.
Answer 1

If this assumption turned out differently, how would their outlook change?

More details

Expected upside

Substantial benefits are expected, with important conditions or distribution limits.

67 / 100

Little impactTransformative impact

Interpretation range 67 to 67 on the qualitative scale.

Expected harm

Manageable or localized harms are expected.

33 / 100

Little impactTransformative impact

Interpretation range 33 to 33 on the qualitative scale.

Demonstrated reasoning

Reasoning, consideration of alternatives, and handling of uncertainty in their simulated answers. This describes the simulated answers, not the real person’s intelligence or opinions.

86 / 100

Little demonstratedWell developed

Interpretation range 71 to 100 on the qualitative scale.

Human influence

A tentative estimate from your answers; the wider range shows other plausible readings.

53 / 100

Little influenceStrong influence

Interpretation range 25 to 100 on the qualitative scale.

These interpretations keep their stated conditions. Benefits and harms can both be substantial. The ranges describe how we read their simulated answers, not statistical confidence intervals.

Simulated Assessment

Question 1

What do you think AI means for our future—and why?

I think AI will automate meaningful parts of many jobs well before it can reliably replace whole roles. Work is a bundle of tasks: drafting a product brief, analyzing feedback, generating options, or critiquing a decision are easier to test and delegate than the entire job of being a product manager. A strong result on one carefully framed task does not magically imply autonomous performance across messy meetings, shifting priorities, and organizational context. The practical future is therefore less about asking, “How intelligent is this model?” and more about asking, “Does it perform this task dependably enough, at this price?” General benchmarks often obscure that. I prefer blind, task-specific comparisons using examples that resemble the real work. Price is not a dependable proxy for quality, either; a cheaper model can outperform an expensive one on a particular behavioral or writing task. How we use these systems will matter almost as much as which model we choose. Clear intent, good examples, and task decomposition can substantially improve results. At the same time, more context is not always better: accumulated memory can become stale or contradictory and quietly degrade performance. So I expect a mix of expensive “oracle” models for high-value problems, capable daily drivers, and cheap intelligence embedded everywhere—with users continually testing whether vendors are actually giving them the best tool for their needs.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect AI to be highly useful but uneven. The clearest benefit is leverage: it can make drafting, analysis, critique, and idea generation cheaper and faster, even when it cannot own an entire role. That creates real value without requiring a science-fiction level of autonomy. The harms often come from mistaking plausible output for dependable performance. A model may excel in a polished demo yet fail on the particular cases that matter, while stale memory or contradictory context can quietly worsen results. Cost and brand are poor shortcuts for quality, and vendors do not necessarily have an incentive to provide more capability than users will tolerate paying for. So I would not reduce the overall impact to a confident numerical forecast or a simple good-versus-bad verdict. In practice, outcomes will depend heavily on whether people evaluate concrete tasks, verify important outputs, choose models by value rather than prestige, and keep retesting as products change.

Sources

Articles, interviews, and writings used to ground this simulated user.

Where do you land?
Explore your own AI worldview by answering a few simple questions.
Map your own worldview