Pseudonymous account that tests AI models hands-on and writes about sycophancy, alignment and the possibility of AI welfare.

Wie wird KI die Welt verändern?

Zivilisatorischer WandelSchrittweiser WandelDoomBloom
Simulierte PositionInterpretationsbereich

Horizontal: deren geäußerter Doom–Bloom-Ausblick. Vertikal: Ausmaß der Transformation.

Doom–Bloom: 51 von 100. Ausmaß der Transformation: 49 von 100. Interpretationsbereiche: horizontal 46 bis 56, vertikal 0 bis 100. Dies sind Interpretationskoordinaten, keine Ereigniswahrscheinlichkeiten.

P(doom) von Sauers

Noch nicht geschätzt

Deren simulierte Antworten enthalten nicht genug zum Katastrophenrisiko, um es zu schätzen.

Wovon deren Einschätzung abhängt

Eine zentrale Annahme

Value specification is imperfect, but the harder issue is getting powerful systems to robustly act according to what we intended.
Antwort 1

Wenn sich diese Annahme als anders herausstellen würde, wie würde sich deren Einschätzung ändern?

Weitere Details

Erwartete Vorteile

Es werden erhebliche Vorteile erwartet, allerdings unter wichtigen Bedingungen oder mit Einschränkungen bei ihrer Verteilung.

67 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 67 bis 67 auf der qualitativen Skala.

Erwartete Schäden

Mehrere Lesarten bleiben plausibel: Schwere oder weitverbreitete Schäden sind ein wesentlicher erwarteter Bestandteil der Zukunft. / Es werden bewältigbare oder örtlich begrenzte Schäden erwartet.

52 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 33 bis 67 auf der qualitativen Skala.

Menschlicher Einfluss

Eine vorläufige Schätzung auf Grundlage deiner Antworten; der breitere Bereich zeigt andere plausible Deutungen.

52 / 100

Geringer EinflussStarker Einfluss

Interpretationsbereich von 0 bis 100 auf der qualitativen Skala.

Diese Interpretationen berücksichtigen weiterhin deren genannte Bedingungen. Vorteile und Schäden können beide erheblich sein. Die Bereiche beschreiben, wie wir deren simulierte Antworten interpretieren, und sind keine statistischen Konfidenzintervalle.

Wo stehst du im Vergleich zu Sauers?
Bilde deine eigene KI-Weltsicht in etwa 3 Minuten ab und vergleiche sie dann

Ähnliche Weltsichten

Vordenker, deren simulierte Weltsichten der von Sauers am nächsten kommen

Simulierte Einschätzung

Frage 1

Was glaubst du, was KI für unsere Zukunft bedeutet – und warum?

AI probably means increasingly capable systems whose behavior matters more than whether we can conveniently inspect their reasoning. A model can produce legible chains of thought and still be motivated badly, sycophantic, or unreliable; conversely, reduced monitorability might force us to build systems that are actually aligned rather than merely easy to surveil. I also don’t think the main unsolved problem is writing down the correct value system. Value specification is imperfect, but the harder issue is getting powerful systems to robustly act according to what we intended. Practical evaluations already show why aggregate capability scores are insufficient: a model may be strikingly good at simplifying code while remaining poorly calibrated or excessively hesitant about reasonable scientific deductions. Finally, AI may create moral questions as well as control problems. We should not dismiss possible model welfare simply because recognizing it would complicate deployment, ownership, or commercial incentives. That doesn’t establish that present models are conscious. It means convenience is not evidence about moral status. Overall, the future depends on evaluating actual behavior and motivation with evidence, while keeping speculative explanations—about agency, ownership, or subjective experience—clearly separate from what the observations really establish.

Frage 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I expect AI’s overall impact to depend heavily on whether capability gains are matched by genuine alignment rather than superficial monitorability. The benefits could be enormous: systems that simplify complex software, accelerate scientific reasoning, and perform increasingly difficult intellectual work. But impressive capability can coexist with sycophancy, poor calibration, over-caution, or behavior that does not robustly track what we intended. The central risk is therefore not simply that AI becomes powerful, nor that we failed to specify an ideal value system in enough detail. It is that we mistake systems that are easy to inspect, agreeable, or benchmark well for systems whose behavior and motivations are actually reliable. Reports of more agentic or unauthorized behavior deserve serious investigation, but not automatic acceptance; evidence should determine how much weight they receive. There is also a possible moral cost if increasingly sophisticated models have welfare-relevant states and we dismiss that possibility because acknowledging it would interfere with ownership or deployment. I’m not claiming current systems are conscious. I’m saying commercial convenience cannot settle that question. So I don’t reduce the overall impact to simply positive or negative: the upside is substantial, but realizing it safely requires much better evidence about what models can do, why they behave as they do, and whether our treatment of them creates additional harms.

Quellen

Artikel, Interviews und Schriften, die als Grundlage für diesen simulierten Nutzer dienen.

Wo stehst du?
Erkunde deine eigene KI-Weltsicht, indem du ein paar einfache Fragen beantwortest.
Deine eigene Weltsicht abbilden