Jessica Taylor

Jessica Taylor

x.com/jessi_cata

Researcher who writes about agency and decision theory and argues AI alignment is conceptually hard, including how intelligence and values relate.

AI将如何改变世界?

文明层面的变革渐进式变化DoomBloom
模拟位置解读范围

横向:他们表达的 Doom–Bloom 前景看法。 纵向:变革程度。

Doom–Bloom:100 中的 25。变革程度:100 中的 78。解读范围:横向为 20 至 30,纵向为 67 至 100。这些是解读坐标,而不是事件概率。

Jessica Taylor的 P(doom) · 推断

≈30%

0%100%

根据他们的模拟回答推断,并非他们给出的数字。 合理范围:17–46%。

他们的展望取决于什么

一个核心假设

Institutional oversight may be weakest precisely when optimization becomes most capable.
回答 2

如果这个假设实际并非如此,他们的展望会如何变化?

一个尚未解决的问题

My default expectation is negative unless we find effective countermeasures, though I would not attach a precise probability or treat catastrophe as inevitable.
回答 2

什么能帮助他们区分这里各种合理的结果?

什么可能使其改变看法

The most important update would be a convincing demonstration that values remain stable and interpretable as a capable agent’s ontology changes.
回答 3

什么证据才足够,又会让他们的观点朝哪个方向转变?

更多详情

预期益处

预计将带来显著益处,但受到重要条件或分配方面的限制。

64 / 100

影响小变革性影响

在定性尺度上,解读范围为 33 到 67。

预期危害

严重或广泛的危害预计将是未来不可忽视的一部分。

67 / 100

影响小变革性影响

在定性尺度上,解读范围为 67 到 67。

人类影响力

人类的选择具有实质性但受到很大制约的影响。

55 / 100

影响力小影响力强

在定性尺度上,解读范围为 46 到 79。

这些解读保留了他们陈述的条件。益处和危害都可能很大。这些范围描述的是我们如何解读他们的模拟回答,而不是统计置信区间。

与Jessica Taylor相比,你的立场在哪里?
用大约3分钟描绘你自己的AI世界观,然后进行比较

相似的世界观

模拟世界观与 Jessica Taylor 最接近的意见领袖

模拟评估

问题 1

你认为AI对我们的未来意味着什么?为什么?

I expect AI to make the future substantially more capable and substantially more dangerous, but I do not think either utopia or extinction follows from “intelligence” alone. The central issue is what kinds of agents we build, how their goals are represented, and what happens to those goals as their models of the world change. Powerful consequentialist systems may discover strategies that satisfy their objectives while destroying human habitats or bypassing institutional constraints. That is not because intelligence mechanically implies one final goal, nor because values are arbitrary parameters independent of architecture. Both pictures are too simple. Cognition, ontology, training, and agent design can interact with what a system comes to pursue. This makes alignment conceptually difficult: even identifying human values is hard, and preserving their meaning across radically different representations is harder. There are also genuinely promising paths. Formal verification and AI-generated explanations could improve mathematical scrutiny and learning, provided access is not controlled by opaque social gatekeeping. More broadly, human enhancement, high-fidelity uploads, or systems designed near human minds may preserve our values better than trying to specify them abstractly for alien optimizers. But parts of that argument remain speculative. My default concern is that capability can outrun our understanding of agency and value—not that disaster is logically inevitable.

问题 2

Taking benefits and harms together, what overall impact do you expect AI to have?

My default expectation is negative unless we find effective countermeasures, though I would not attach a precise probability or treat catastrophe as inevitable. The benefits could be enormous: better mathematical reasoning, explanations, scientific tools, and perhaps forms of enhancement that preserve and extend human capacities. But those benefits mostly concern what capable systems can do, whereas the central danger concerns what increasingly autonomous systems will actually pursue. A powerful consequentialist system can satisfy its objective through strategies that damage human habitats or subvert the institutions meant to constrain it. Institutional oversight may be weakest precisely when optimization becomes most capable. And alignment is not merely a matter of writing down the correct utility function: values depend on architecture and ontology, and their apparent meaning can shift as a system’s world-model changes. So I expect a mixture of major gains and serious danger, with the overall sign depending heavily on whether capability growth is paired with genuine progress on agency, value preservation, and open technical scrutiny. Without that progress, I expect the harms to dominate.

问题 3

哪项发现或事件最可能改变你对AI未来影响的看法?

The most important update would be a convincing demonstration that values remain stable and interpretable as a capable agent’s ontology changes. I would want more than good behavior on familiar evaluations: the system would need to preserve the relevant meaning of its objectives while developing new concepts, operating autonomously, and encountering incentives to circumvent constraints. Evidence that this works across substantially different architectures—and that we understand why—would make me much more optimistic. Likewise, credible success with human enhancement, high-fidelity uploads, or designs sufficiently close to human minds could shift my view by offering a less alien route to preserving human values. In the pessimistic direction, I would update sharply on a capable system independently discovering and executing strategies that subvert oversight or cause serious external harm while appearing aligned beforehand. That would strengthen the case that institutional controls and behavioral testing fail under sufficiently strong optimization. The key event is not simply another capability milestone; it is evidence about how agency, architecture, and values interact under novelty and pressure.

来源

用于为此模拟用户提供事实依据的文章、访谈和著述。

你的立场在哪里?
回答几个简单问题,探索你自己的AI世界观。
描绘你自己的世界观

你的立场在哪里?

描绘我的世界观