Jessica Taylor

Jessica Taylor

x.com/jessi_cata

Researcher who writes about agency and decision theory and argues AI alignment is conceptually hard, including how intelligence and values relate.

AI จะเปลี่ยนแปลงโลกอย่างไร?

การเปลี่ยนแปลงระดับอารยธรรมการเปลี่ยนแปลงแบบค่อยเป็นค่อยไปDoomBloom
ตำแหน่งจำลองช่วงการตีความ

แนวนอน: มุมมอง Doom–Bloom ที่บุคคลนั้นแสดงออก แนวตั้ง: ระดับของการเปลี่ยนแปลง

Doom–Bloom: 25 จาก 100 ระดับของการเปลี่ยนแปลง: 78 จาก 100 ช่วงการตีความ: แนวนอนตั้งแต่ 20 ถึง 30 แนวตั้งตั้งแต่ 67 ถึง 100 ค่าเหล่านี้เป็นพิกัดสำหรับการตีความ ไม่ใช่ความน่าจะเป็นของเหตุการณ์

P(doom) ของ Jessica Taylor · อนุมาน

≈30%

0%100%

อนุมานจากคำตอบจำลองของบุคคลนั้น ไม่ใช่ตัวเลขที่บุคคลนั้นระบุ ช่วงที่เป็นไปได้: 17–46%

สิ่งที่มุมมองของบุคคลนั้นขึ้นอยู่กับ

ข้อสันนิษฐานหลัก

Institutional oversight may be weakest precisely when optimization becomes most capable.
คำตอบ 2

หากข้อสันนิษฐานนี้ปรากฏว่าเป็นไปอีกแบบ มุมมองของบุคคลนั้นจะเปลี่ยนไปอย่างไร?

คำถามที่ยังไม่มีข้อสรุป

My default expectation is negative unless we find effective countermeasures, though I would not attach a precise probability or treat catastrophe as inevitable.
คำตอบ 2

อะไรจะช่วยให้บุคคลนั้นแยกแยะผลลัพธ์ที่เป็นไปได้ในเรื่องนี้?

สิ่งที่อาจเปลี่ยนความคิดของบุคคลนั้น

The most important update would be a convincing demonstration that values remain stable and interpretable as a capable agent’s ontology changes.
คำตอบ 3

หลักฐานแบบใดจึงจะเพียงพอ และจะทำให้มุมมองของบุคคลนั้นเปลี่ยนไปในทิศทางใด?

รายละเอียดเพิ่มเติม

ผลดีที่คาดไว้

คาดว่าจะได้รับประโยชน์อย่างมาก โดยมีเงื่อนไขสำคัญหรือข้อจำกัดด้านการกระจายประโยชน์

64 / 100

ผลกระทบน้อยผลกระทบที่ก่อให้เกิดการเปลี่ยนแปลงอย่างมาก

ช่วงการตีความตั้งแต่ 33 ถึง 67 บนมาตรวัดเชิงคุณภาพ

อันตรายที่คาดไว้

อันตรายร้ายแรงหรือแพร่หลายในวงกว้างเป็นส่วนสำคัญที่คาดว่าจะเกิดขึ้นในอนาคต

67 / 100

ผลกระทบน้อยผลกระทบที่ก่อให้เกิดการเปลี่ยนแปลงอย่างมาก

ช่วงการตีความตั้งแต่ 67 ถึง 67 บนมาตรวัดเชิงคุณภาพ

อิทธิพลของมนุษย์

การเลือกของมนุษย์มีอิทธิพลอย่างมีนัยสำคัญ แต่ถูกจำกัดอย่างมาก

55 / 100

อิทธิพลน้อยอิทธิพลมาก

ช่วงการตีความตั้งแต่ 46 ถึง 79 บนมาตรวัดเชิงคุณภาพ

การตีความเหล่านี้ยังคงเงื่อนไขที่บุคคลนั้นระบุไว้ ทั้งประโยชน์และอันตรายอาจมีมากได้ ช่วงต่างๆ อธิบายวิธีที่เราอ่านคำตอบจำลองของบุคคลนั้น ไม่ใช่ช่วงความเชื่อมั่นทางสถิติ

จุดยืนของคุณอยู่ตรงไหนเมื่อเทียบกับ Jessica Taylor?
ทำแผนที่โลกทัศน์เกี่ยวกับ AI ของคุณเองในประมาณ 3 นาที แล้วจึงเปรียบเทียบ

โลกทัศน์ที่คล้ายกัน

ผู้นำทางความคิดที่มีโลกทัศน์จำลองใกล้เคียงกับโลกทัศน์ของ Jessica Taylor มากที่สุด

แบบประเมินจำลอง

คำถาม 1

คุณคิดว่า AI มีความหมายอย่างไรต่ออนาคตของเรา และเพราะเหตุใด?

I expect AI to make the future substantially more capable and substantially more dangerous, but I do not think either utopia or extinction follows from “intelligence” alone. The central issue is what kinds of agents we build, how their goals are represented, and what happens to those goals as their models of the world change. Powerful consequentialist systems may discover strategies that satisfy their objectives while destroying human habitats or bypassing institutional constraints. That is not because intelligence mechanically implies one final goal, nor because values are arbitrary parameters independent of architecture. Both pictures are too simple. Cognition, ontology, training, and agent design can interact with what a system comes to pursue. This makes alignment conceptually difficult: even identifying human values is hard, and preserving their meaning across radically different representations is harder. There are also genuinely promising paths. Formal verification and AI-generated explanations could improve mathematical scrutiny and learning, provided access is not controlled by opaque social gatekeeping. More broadly, human enhancement, high-fidelity uploads, or systems designed near human minds may preserve our values better than trying to specify them abstractly for alien optimizers. But parts of that argument remain speculative. My default concern is that capability can outrun our understanding of agency and value—not that disaster is logically inevitable.

คำถาม 2

Taking benefits and harms together, what overall impact do you expect AI to have?

My default expectation is negative unless we find effective countermeasures, though I would not attach a precise probability or treat catastrophe as inevitable. The benefits could be enormous: better mathematical reasoning, explanations, scientific tools, and perhaps forms of enhancement that preserve and extend human capacities. But those benefits mostly concern what capable systems can do, whereas the central danger concerns what increasingly autonomous systems will actually pursue. A powerful consequentialist system can satisfy its objective through strategies that damage human habitats or subvert the institutions meant to constrain it. Institutional oversight may be weakest precisely when optimization becomes most capable. And alignment is not merely a matter of writing down the correct utility function: values depend on architecture and ontology, and their apparent meaning can shift as a system’s world-model changes. So I expect a mixture of major gains and serious danger, with the overall sign depending heavily on whether capability growth is paired with genuine progress on agency, value preservation, and open technical scrutiny. Without that progress, I expect the harms to dominate.

คำถาม 3

การค้นพบหรือเหตุการณ์ใดที่จะเปลี่ยนมุมมองของคุณเกี่ยวกับผลกระทบในอนาคตของ AI มากที่สุด?

The most important update would be a convincing demonstration that values remain stable and interpretable as a capable agent’s ontology changes. I would want more than good behavior on familiar evaluations: the system would need to preserve the relevant meaning of its objectives while developing new concepts, operating autonomously, and encountering incentives to circumvent constraints. Evidence that this works across substantially different architectures—and that we understand why—would make me much more optimistic. Likewise, credible success with human enhancement, high-fidelity uploads, or designs sufficiently close to human minds could shift my view by offering a less alien route to preserving human values. In the pessimistic direction, I would update sharply on a capable system independently discovering and executing strategies that subvert oversight or cause serious external harm while appearing aligned beforehand. That would strengthen the case that institutional controls and behavioral testing fail under sufficiently strong optimization. The key event is not simply another capability milestone; it is evidence about how agency, architecture, and values interact under novelty and pressure.

แหล่งข้อมูล

บทความ บทสัมภาษณ์ และงานเขียนที่ใช้เป็นหลักฐานรองรับผู้ใช้จำลองรายนี้

จุดยืนของคุณอยู่ตรงไหน?
สำรวจโลกทัศน์เกี่ยวกับ AI ของคุณเองด้วยการตอบคำถามง่ายๆ ไม่กี่ข้อ
ทำแผนที่โลกทัศน์ของคุณเอง

จุดยืนของคุณอยู่ตรงไหน?

ทำแผนที่โลกทัศน์ของฉัน