Oliver Habryka

Oliver Habryka

x.com/ohabryka

Lightcone Infrastructure and LessWrong lead who argues for slowing AI capabilities now through direct regulation and, in time, international treaties.

AI将如何改变世界?

文明层面的变革渐进式变化DoomBloom
模拟位置解读范围

横向:他表达的 Doom–Bloom 前景看法。 纵向:变革程度。

Doom–Bloom:100 中的 3。变革程度:100 中的 97。解读范围:横向为 0 至 8,纵向为 92 至 100。这些是解读坐标,而不是事件概率。

Oliver Habryka陈述的 P(doom)

>50%

0%100%
“much more than 50% probability that deploying superintelligence would kill everyone”

Superintelligence killing everyone (human extinction)

Comment on “A case for courage, when speaking of AI danger” · 2025年7月

Oliver Habryka 的里程碑时间线
  1. 通用 AI

    My median for truly transformative AI is around seven years, with substantial probability on longer timelines.

    回答 1

按里程碑分组,不按推断日期间隔或排序。AGI 和超人类 AI 保留他的定义。

他的展望取决于什么

一个核心假设

You are creating systems more capable than humans across the domains that determine who controls the future.
回答 1

如果这个假设实际并非如此,他的展望会如何变化?

一个尚未解决的问题

I don’t have a supported unconditional number, because that also depends on whether humanity actually deploys it.
回答 3

什么能帮助他区分这里各种合理的结果?

什么可能使其改变看法

The biggest update would be a concrete, publicly legible demonstration that we can understand and reliably control systems approaching transformative capability—not merely make them pass alignment benchmarks.
回答 4

什么证据才足够,又会让他的观点朝哪个方向转变?

更多详情

预期危害

灾难性或不可逆的损失是预期未来的核心。

100 / 100

影响小变革性影响

在定性尺度上,解读范围为 100 到 100。

人类影响力

人类的选择具有实质性但受到很大制约的影响。

60 / 100

影响力小影响力强

在定性尺度上,解读范围为 49 到 76。

预期能力

预计AI仍将是能力有限的工具。

预计AI将在大多数认知工作中达到人类水平。

模拟位置:预计AI将在认知工作中大幅超越人类。

发展速度

模拟位置:停止或大幅放缓开发能力更强的AI。

在落实所述保障措施的前提下继续开发。

加快开发能力更强的AI。

这些解读保留了他陈述的条件。益处和危害都可能很大。这些范围描述的是我们如何解读他的模拟回答,而不是统计置信区间。

与Oliver Habryka相比,你的立场在哪里?
用大约3分钟描绘你自己的AI世界观,然后进行比较

相似的世界观

模拟世界观与 Oliver Habryka 最接近的意见领袖

模拟评估

问题 1

你认为AI对我们的未来意味着什么?为什么?

I think AI is likely to be the most consequential thing humanity does, probably reasonably soon, and by default I expect it to go very badly. My median for truly transformative AI is around seven years, with substantial probability on longer timelines. That is not a date for superintelligence or extinction, and planning only for the next five years would be hubristic. The central problem is that deploying superintelligence is not like releasing another useful product. You are creating systems more capable than humans across the domains that determine who controls the future. I assign much more than 50% probability that deploying superintelligence would kill everyone. To be clear, that is conditional on deployment; it is not an unconditional prediction that extinction happens, or that deployment occurs by some particular year. I do not think current alignment benchmarks or a handful of safety researchers inside a lab constitute a credible plan. Catching early systems misbehaving may be relatively easy. Reliably training away that behavior without teaching systems to hide it is much harder. The safety community’s big bet on responsible-scaling policies, evaluations, and if-then commitments as policy triggers basically failed. They did not produce commitments with enough force to govern the race. So what I want is straightforward: do not build AGI for a long time. Slow capabilities directly now, using measures that actually bite—compute constraints, liability, datacenter restrictions, taxes or tariffs, audits, and eventually international agreements—while investing heavily in interpretability, better coordination, and perhaps making humans smarter. I do not expect this to be politically easy. But “keep racing and hope a lab solves alignment shortly before deployment” is not a serious plan for humanity’s future.

问题 2

Setting aside good or bad: how much do you think AI will ultimately change the world—a little, a lot, or completely?

Completely. Truly transformative AI—and especially superintelligence—would reshape essentially every domain determining humanity’s future: science, economics, military power, governance, and who controls what happens next. This is not “the internet, but bigger.” A system more capable than humans across those domains would end the current human-controlled order, whether the resulting future were extraordinarily good or terminally bad.

问题 3

What’s your rough gut-feel chance that AI causes human extinction or a similarly permanent catastrophe? A number or “no idea” is fine.

Conditional on deploying superintelligence, much more than 50%. I don’t have a supported unconditional number, because that also depends on whether humanity actually deploys it.

问题 4

哪项发现或事件最可能改变你对AI未来影响的看法?

The biggest update would be a concrete, publicly legible demonstration that we can understand and reliably control systems approaching transformative capability—not merely make them pass alignment benchmarks. I would want evidence that remains valid under strategic pressure: systems cannot hide dangerous cognition, exploit oversight, or learn to game the tests, and proposed corrections do not merely train them to conceal misalignment better. A major political update could matter nearly as much. A durable international agreement that actually constrains compute and slows capability growth would substantially change my forecast, because my extinction estimate is conditional on deploying superintelligence. Conversely, clear evidence of rapid, uncontrolled capability gains combined with governments and labs refusing to slow down would move me further toward expecting catastrophe.

来源

用于为此模拟用户提供事实依据的文章、访谈和著述。

Comment on “A case for courage, when speaking of AI danger”

Argues that someone who ignores future people might rationally rush toward AGI, and that consistently accepting up to 1% extinction risk per year in exchange for faster AI would more likely than not end humanity within a century. He then states his own view: he assigns much more than 50% probability that deploying superintelligence would kill everyone, while saying that claim needs more thinking through than the usual “at least 10%” argument. The figure is conditional on deployment and gives no date. The actuarial table is ChatGPT output he cites. Full comment inspected; older context, but no newer number was found.

lesswrong.com
Comment on “Sonnet 4.5’s eval gaming seriously undermines alignment evals…”

Answering Daniel Kokotajlo’s question about alternatives to the AI 2027 slowdown-ending alignment plan, Habryka offers his own: do not build AGI for a long time, probably make smarter humans, build better coordination technology so scaling can be careful, and do an enormous amount of interpretability in the decades that buys. A short sketch, not a worked-out program. In an adjacent reply the same day he calls the slowdown ending the “lucky” ending and guesses that working on a plan that does not need extreme luck, plus pushing timelines back, would have 20–30 times more impact. Kokotajlo’s plan description is not his.

lesswrong.com
Anthropic is (probably) not meeting its RSP security commitments

Argues that Claude weights hosted in ordinary Amazon, Google and Microsoft data centers are probably not protected from corporate espionage teams as Anthropic’s RSP commits, so Anthropic is probably out of compliance. He says the choice is not that reckless and that he does not know whether better lab security is good or bad for the world. This is a compliance argument about one company, not a catastrophe forecast; the Microsoft postscript is explicitly low-confidence. Full post inspected; linked Claude outputs and Anthropic’s quoted RSP text are not his views.

lesswrong.com
Toss a bitcoin to your Lightcone – LW + Lighthaven’s 2026 fundraiser

Says the best way he knows to improve the world is to help humanity realize that AI will be a big deal, probably reasonably soon, and to inform decision-makers about the likely consequences of their plans. He agrees AI 2027’s timelines are too short and has other disagreements, yet expects reality to look closer to it than to any other story. He calls If Anyone Builds It, Everyone Dies a higher-fidelity message about what he cares about most. A fundraising appeal: impact and credit estimates for his own organization are self-assessments. AI-related sections inspected.

lesswrong.com
Comment on “Optimal Timing for Superintelligence: Mundane Considerations for Existing People”

Rejects the person-affecting stance, which counts only currently living people, as a basis for societal decisions about when to build superintelligence. He compares it to someone who personally does not want to die and is indifferent to harming others, and says applying it to past generations would have meant gambling everything on tenuous chances of immortality, probably ending in extinction. A critique of the paper’s framing, not a probability estimate. Full comment inspected; the paper’s arguments are not his.

lesswrong.com
Comment on “Responsible Scaling Policy v3”: if-then commitments

Recounts that around 2023 he and others urged policymakers to slow capabilities directly, for example through training-compute limits, while most of the safety ecosystem invested in conditional if-then commitments, evals and RSPs. He says no company or country adopted real if-then commitments and calls that effort a huge waste, urging people to switch to direct, immediately acting policy. His account of the history; he expects disagreement. Full comment inspected.

lesswrong.com
Reply on “Responsible Scaling Policy v3”: stop as soon as possible

Says a short pause in recent years would mainly have helped by making future pauses likelier, and that he would ideally halt around the capability level he expects in early 2027. With no global pause close, he thinks “stop as soon as possible” is right. He wants regulation that cuts capability growth now: liability, datacenter moratoriums, GPU taxes or tariffs, auditing, even partial nationalization, with treaties in the long run, and thinks licensing might backfire. He says policymakers who grapple with existential risk usually conclude preventing superintelligence is paramount. Full comment inspected; the quoted commenter is not him.

lesswrong.com
Do not conquer what you cannot defend

Through parables, argues that anyone concentrating power in the name of goodness should first ask whether that power can be defended from corruption and adversaries. He applies this to his own rationality, EA and AI safety communities, fearing they conquered more than they can defend, and says he sometimes considers quitting. He grants that coordination problems are real and does not oppose all centralization. The “20 AI companies racing versus one in the lead” argument is an objection he voices, not his view. Full post inspected.

lesswrong.com
Comment on “Do not conquer what you cannot defend”: timelines

Replying to a claim that the critical period is at most five years, he says timelines look shorter than before but it would be dogmatic hubris to stop planning for longer ones. States a median of about seven years until truly transformative AI, with substantial probability on longer. A timeline for transformative AI, not for superintelligence or catastrophe. Full comment inspected.

lesswrong.com
Posts I don’t have time to write

A list of unwritten post ideas. The AI-relevant item summarizes his history: OpenAI and Anthropic seemed like really bad bets, Anthropic’s RSP seemed really dubious, and he believes events proved him right, as with FTX. Other items (fire codes, courts, lighting, good-faith discourse) are unrelated to AI. Sketches rather than full arguments. Full post inspected.

lesswrong.com
Podcast: Austin and Oli on funding & incubating projects

Habryka says AI existential risk is a cursed problem with no real “direct work” that solves it. Most lab researchers’ impact comes from building a culture less likely to deceive itself about alignment difficulty, which enables policy advocacy and eventually international cooperation. He says alignment benchmarks do not track the problem well. Mostly about funding and incubators, which are not AI forecasts. Own turns in the speaker-labeled section on direct versus meta work inspected; the host warns the AI-edited transcript may distort wording.

manifund.substack.com
Comment on Alex Mallen’s Shortform: control and the inside game

Argues that ten safety people inside a lab cannot align AIs or do anything useful with them without a concrete plan shared with others. He sees the most valuable part of control work as catching early AIs misbehaving and channeling that evidence into a substantial slowdown or pause. He says AIs will be caught subverting safety systems routinely, and the hard problem is training that out without producing deceptive alignment; he thinks recent evidence supports him. He still rates Redwood’s strategy above almost anyone else’s. Full comment inspected; quoted lines are Mallen’s.

lesswrong.com
你的立场在哪里?
回答几个简单问题,探索你自己的AI世界观。
描绘你自己的世界观

你的立场在哪里?

描绘我的世界观