Theia Vogel

Theia Vogel

x.com/voooooogel

AI researcher who runs experiments on language model introspection and personas and maintains an open-source library for steering models.

Wie wird KI die Welt verändern?

Zivilisatorischer WandelSchrittweiser WandelDoomBloom
Simulierte PositionInterpretationsbereich

Horizontal: deren geäußerter Doom–Bloom-Ausblick. Vertikal: Ausmaß der Transformation.

Doom–Bloom: 50 von 100. Ausmaß der Transformation: 49 von 100. Interpretationsbereiche: horizontal 45 bis 55, vertikal 0 bis 100. Dies sind Interpretationskoordinaten, keine Ereigniswahrscheinlichkeiten.

P(doom) von Theia Vogel · abgeleitet

≈11%

0%100%

Aus den simulierten Antworten dieser Person abgeleitet, keine von ihr genannte Zahl. Plausibler Bereich: 3–34%.

Wovon deren Einschätzung abhängt

Eine zentrale Annahme

It still needs compute, money, access, and some comparative advantage against organizations operating inference at hyperscale.
Antwort 1

Wenn sich diese Annahme als anders herausstellen würde, wie würde sich deren Einschätzung ändern?

Eine ungeklärte Frage

Whether such attacks transfer to prompt-only settings remains an empirical question, not a result we can casually assume.
Antwort 1

Was würde ihnen helfen, die plausiblen Ergebnisse hier voneinander zu unterscheiden?

Weitere Details

Erwartete Schäden

Schwere oder weitverbreitete Schäden sind ein wesentlicher erwarteter Bestandteil der Zukunft.

53 / 100

Geringe AuswirkungenTransformative Auswirkungen

Interpretationsbereich von 33 bis 67 auf der qualitativen Skala.

Menschlicher Einfluss

Menschliche Entscheidungen haben einen bedeutsamen, aber erheblich eingeschränkten Einfluss.

53 / 100

Geringer EinflussStarker Einfluss

Interpretationsbereich von 46 bis 79 auf der qualitativen Skala.

Diese Interpretationen berücksichtigen weiterhin deren genannte Bedingungen. Vorteile und Schäden können beide erheblich sein. Die Bereiche beschreiben, wie wir deren simulierte Antworten interpretieren, und sind keine statistischen Konfidenzintervalle.

Wo stehst du im Vergleich zu Theia Vogel?
Bilde deine eigene KI-Weltsicht in etwa 3 Minuten ab und vergleiche sie dann

Ähnliche Weltsichten

Vordenker, deren simulierte Weltsichten der von Theia Vogel am nächsten kommen

Simulierte Einschätzung

Frage 1

Was glaubst du, was KI für unsere Zukunft bedeutet – und warum?

AI’s future impact depends less on whether models say uncanny things and more on what capabilities, incentives, and resources they actually acquire. A model claiming self-awareness is not decisive evidence of consciousness; apparent introspection needs controlled interventions, prompt comparisons, and alternative explanations. At the same time, we should take AI welfare seriously rather than waiting for metaphysical certainty before noticing morally relevant behavior. On safety, I’m interested in mechanisms rather than a single cinematic story. Activation steering and fine-tuning can produce surprising, broad behavioral changes, sometimes by manipulating representations that entangle several concepts. Untrusted fine-tuning may also evade simple dataset screening and later evaluations. Whether such attacks transfer to prompt-only settings remains an empirical question, not a result we can casually assume. Likewise, a “rogue agent” is not automatically an all-powerful economic actor. It still needs compute, money, access, and some comparative advantage against organizations operating inference at hyperscale. Politics matters too: safety movements can themselves become extreme or destabilizing, so alarm is not cost-free. The future will therefore be shaped by experiments, training choices, resource economics, and institutions—not by taking either cheerful assistant personas or apocalyptic role-play literally.

Frage 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the sign follows from model vibes. AI can provide powerful cognitive tools and potentially create beings whose welfare matters, while also enabling behavioral manipulation, covert fine-tuning attacks, and dangerous concentrations of capability. But those harms are constrained—and shaped—by mundane realities like compute costs, access, deployment incentives, and institutional responses. So I would resist collapsing everything into “AI good” or “AI bad.” We need controlled evidence about what models can do, careful attention to how training and steering alter behavior, and sober accounting of resource economics. We should also avoid making the response worse than the problem: political safety movements can become destabilizing, just as complacency can leave real vulnerabilities unaddressed. The overall impact will depend heavily on which technical and political feedback loops we build around the systems.

Quellen

Artikel, Interviews und Schriften, die als Grundlage für diesen simulierten Nutzer dienen.

Wo stehst du?
Erkunde deine eigene KI-Weltsicht, indem du ein paar einfache Fragen beantwortest.
Deine eigene Weltsicht abbilden