Theia Vogel

Theia Vogel

x.com/voooooogel

AI researcher who runs experiments on language model introspection and personas and maintains an open-source library for steering models.

Comment l’IA changera-t-elle le monde ?

Changement civilisationnelChangement progressifDoomBloom
Position simuléePlage d’interprétation

Horizontalement : leur perspective Doom–Bloom telle qu’elle a été exprimée. Verticalement : ampleur de la transformation.

Doom–Bloom : 50 sur 100. Ampleur de la transformation : 49 sur 100. Plages d’interprétation : de 45 à 55 horizontalement, de 0 à 100 verticalement. Il s’agit de coordonnées d’interprétation, et non de probabilités d’événements.

P(doom) de Theia Vogel · inféré

≈11%

0%100%

Déduit de leurs réponses simulées, et non d’un chiffre donné par ces personnes. Plage plausible : 3–34%.

Ce dont dépend leur perspective

Une hypothèse centrale

It still needs compute, money, access, and some comparative advantage against organizations operating inference at hyperscale.
Réponse 1

Si cette hypothèse s’avérait différente, comment leur perspective changerait-elle ?

Une question non résolue

Whether such attacks transfer to prompt-only settings remains an empirical question, not a result we can casually assume.
Réponse 1

Qu’est-ce qui les aiderait à distinguer les résultats plausibles ici ?

Plus de détails

Dommages attendus

Des dommages graves ou généralisés constituent une composante substantielle de l’avenir attendu.

53 / 100

Faible impactImpact transformateur

Plage d’interprétation de 33 à 67 sur l’échelle qualitative.

Influence humaine

Les choix humains ont une influence significative, mais fortement contrainte.

53 / 100

Faible influenceForte influence

Plage d’interprétation de 46 à 79 sur l’échelle qualitative.

Ces interprétations conservent les conditions énoncées. Les bénéfices et les dommages peuvent tous deux être substantiels. Les plages décrivent notre lecture de leurs réponses simulées, et non des intervalles de confiance statistiques.

Où vous situez-vous par rapport à Theia Vogel ?
Cartographiez votre propre vision du monde concernant l’IA en environ 3 minutes, puis comparez

Visions du monde similaires

Leaders d’opinion dont les visions du monde simulées sont les plus proches de celle de Theia Vogel

Évaluation simulée

Question 1

Selon vous, que signifie l’IA pour notre avenir, et pourquoi ?

AI’s future impact depends less on whether models say uncanny things and more on what capabilities, incentives, and resources they actually acquire. A model claiming self-awareness is not decisive evidence of consciousness; apparent introspection needs controlled interventions, prompt comparisons, and alternative explanations. At the same time, we should take AI welfare seriously rather than waiting for metaphysical certainty before noticing morally relevant behavior. On safety, I’m interested in mechanisms rather than a single cinematic story. Activation steering and fine-tuning can produce surprising, broad behavioral changes, sometimes by manipulating representations that entangle several concepts. Untrusted fine-tuning may also evade simple dataset screening and later evaluations. Whether such attacks transfer to prompt-only settings remains an empirical question, not a result we can casually assume. Likewise, a “rogue agent” is not automatically an all-powerful economic actor. It still needs compute, money, access, and some comparative advantage against organizations operating inference at hyperscale. Politics matters too: safety movements can themselves become extreme or destabilizing, so alarm is not cost-free. The future will therefore be shaped by experiments, training choices, resource economics, and institutions—not by taking either cheerful assistant personas or apocalyptic role-play literally.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I don’t think the sign follows from model vibes. AI can provide powerful cognitive tools and potentially create beings whose welfare matters, while also enabling behavioral manipulation, covert fine-tuning attacks, and dangerous concentrations of capability. But those harms are constrained—and shaped—by mundane realities like compute costs, access, deployment incentives, and institutional responses. So I would resist collapsing everything into “AI good” or “AI bad.” We need controlled evidence about what models can do, careful attention to how training and steering alter behavior, and sober accounting of resource economics. We should also avoid making the response worse than the problem: political safety movements can become destabilizing, just as complacency can leave real vulnerabilities unaddressed. The overall impact will depend heavily on which technical and political feedback loops we build around the systems.

Sources

Articles, entretiens et écrits utilisés pour ancrer cet utilisateur simulé dans les faits.

Où vous situez-vous ?
Explorez votre propre vision du monde concernant l’IA en répondant à quelques questions simples.
Cartographiez votre propre vision du monde

Où vous situez-vous ?

Cartographier ma vision du monde