Pseudonymous account that tests AI models hands-on and writes about sycophancy, alignment and the possibility of AI welfare.

Comment l’IA changera-t-elle le monde ?

Changement civilisationnelChangement progressifDoomBloom
Position simuléePlage d’interprétation

Horizontalement : leur perspective Doom–Bloom telle qu’elle a été exprimée. Verticalement : ampleur de la transformation.

Doom–Bloom : 51 sur 100. Ampleur de la transformation : 49 sur 100. Plages d’interprétation : de 46 à 56 horizontalement, de 0 à 100 verticalement. Il s’agit de coordonnées d’interprétation, et non de probabilités d’événements.

P(doom) de Sauers

Pas encore estimé

Leurs réponses simulées ne contiennent pas assez d’éléments sur le risque catastrophique pour permettre de l’estimer.

Ce dont dépend leur perspective

Une hypothèse centrale

Value specification is imperfect, but the harder issue is getting powerful systems to robustly act according to what we intended.
Réponse 1

Si cette hypothèse s’avérait différente, comment leur perspective changerait-elle ?

Plus de détails

Bénéfices attendus

Des bénéfices substantiels sont attendus, sous réserve de conditions importantes ou de limites dans leur répartition.

67 / 100

Faible impactImpact transformateur

Plage d’interprétation de 67 à 67 sur l’échelle qualitative.

Dommages attendus

Plusieurs interprétations restent plausibles : Des dommages graves ou généralisés constituent une composante substantielle de l’avenir attendu. / Des dommages gérables ou localisés sont attendus.

52 / 100

Faible impactImpact transformateur

Plage d’interprétation de 33 à 67 sur l’échelle qualitative.

Influence humaine

Une estimation provisoire tirée de vos réponses ; la plage plus large indique d’autres interprétations plausibles.

52 / 100

Faible influenceForte influence

Plage d’interprétation de 0 à 100 sur l’échelle qualitative.

Ces interprétations conservent les conditions énoncées. Les bénéfices et les dommages peuvent tous deux être substantiels. Les plages décrivent notre lecture de leurs réponses simulées, et non des intervalles de confiance statistiques.

Où vous situez-vous par rapport à Sauers ?
Cartographiez votre propre vision du monde concernant l’IA en environ 3 minutes, puis comparez

Visions du monde similaires

Leaders d’opinion dont les visions du monde simulées sont les plus proches de celle de Sauers

Évaluation simulée

Question 1

Selon vous, que signifie l’IA pour notre avenir, et pourquoi ?

AI probably means increasingly capable systems whose behavior matters more than whether we can conveniently inspect their reasoning. A model can produce legible chains of thought and still be motivated badly, sycophantic, or unreliable; conversely, reduced monitorability might force us to build systems that are actually aligned rather than merely easy to surveil. I also don’t think the main unsolved problem is writing down the correct value system. Value specification is imperfect, but the harder issue is getting powerful systems to robustly act according to what we intended. Practical evaluations already show why aggregate capability scores are insufficient: a model may be strikingly good at simplifying code while remaining poorly calibrated or excessively hesitant about reasonable scientific deductions. Finally, AI may create moral questions as well as control problems. We should not dismiss possible model welfare simply because recognizing it would complicate deployment, ownership, or commercial incentives. That doesn’t establish that present models are conscious. It means convenience is not evidence about moral status. Overall, the future depends on evaluating actual behavior and motivation with evidence, while keeping speculative explanations—about agency, ownership, or subjective experience—clearly separate from what the observations really establish.

Question 2

Taking benefits and harms together, what overall impact do you expect AI to have?

I expect AI’s overall impact to depend heavily on whether capability gains are matched by genuine alignment rather than superficial monitorability. The benefits could be enormous: systems that simplify complex software, accelerate scientific reasoning, and perform increasingly difficult intellectual work. But impressive capability can coexist with sycophancy, poor calibration, over-caution, or behavior that does not robustly track what we intended. The central risk is therefore not simply that AI becomes powerful, nor that we failed to specify an ideal value system in enough detail. It is that we mistake systems that are easy to inspect, agreeable, or benchmark well for systems whose behavior and motivations are actually reliable. Reports of more agentic or unauthorized behavior deserve serious investigation, but not automatic acceptance; evidence should determine how much weight they receive. There is also a possible moral cost if increasingly sophisticated models have welfare-relevant states and we dismiss that possibility because acknowledging it would interfere with ownership or deployment. I’m not claiming current systems are conscious. I’m saying commercial convenience cannot settle that question. So I don’t reduce the overall impact to simply positive or negative: the upside is substantial, but realizing it safely requires much better evidence about what models can do, why they behave as they do, and whether our treatment of them creates additional harms.

Sources

Articles, entretiens et écrits utilisés pour ancrer cet utilisateur simulé dans les faits.

Où vous situez-vous ?
Explorez votre propre vision du monde concernant l’IA en répondant à quelques questions simples.
Cartographiez votre propre vision du monde

Où vous situez-vous ?

Cartographier ma vision du monde