Raymond Weitekamp

Raymond Weitekamp

x.com/raw_works

Engineer who writes about recursive coding agents and argues their bottleneck is reliability, not intelligence, and that many uses need local models.

Bagaimana AI akan mengubah dunia?

Perubahan peradabanPerubahan bertahapDoomBloom
Posisi simulasiRentang interpretasi

Mendatar: pandangan Doom–Bloom yang mereka ungkapkan. Ke atas: skala transformasi.

Doom–Bloom: 71 dari 100. Skala transformasi: 43 dari 100. Rentang interpretasi: 66 hingga 76 secara horizontal, 15 hingga 60 secara vertikal. Ini adalah koordinat interpretasi, bukan probabilitas kejadian.

P(doom) Raymond Weitekamp · disimpulkan

≈4%

0%100%

Disimpulkan dari jawaban simulasi mereka, bukan angka yang mereka berikan. Rentang yang masuk akal: 2–11%.

Hal-hal yang menentukan pandangan mereka

Asumsi utama

A model that looks limited in a chat interface may perform substantially better when its harness lets it inspect state, run code, test hypotheses, and recursively revise its work.
Jawaban 1

Jika asumsi ini ternyata berbeda, bagaimana pandangan mereka akan berubah?

Pertanyaan yang belum terjawab

The biggest change would come from evidence that reliable agent behavior does—or does not—scale with better harnesses.
Jawaban 3

Apa yang akan membantu mereka membedakan hasil-hasil yang masuk akal di sini?

Hal yang dapat mengubah pandangan mereka

If repeated, independent results showed that tool use, executable reasoning, recursive revision, testing, and bounded permissions still fail unpredictably on consequential tasks, then I would become much less optimistic about broad deployment.
Jawaban 3

Bukti apa yang akan memadai, dan ke arah mana bukti itu akan mengubah pandangan mereka?

Detail lebih lanjut

Manfaat yang diperkirakan

Manfaat besar diperkirakan akan terwujud, dengan syarat penting atau keterbatasan distribusi.

67 / 100

Dampak kecilDampak transformatif

Rentang interpretasi 67 hingga 67 pada skala kualitatif.

Kerugian yang diperkirakan

Kerugian yang dapat dikelola atau bersifat lokal diperkirakan akan terjadi.

39 / 100

Dampak kecilDampak transformatif

Rentang interpretasi 33 hingga 67 pada skala kualitatif.

Pengaruh manusia

Estimasi sementara dari jawaban Anda; rentang yang lebih lebar menunjukkan penafsiran lain yang masuk akal.

57 / 100

Sedikit pengaruhPengaruh kuat

Rentang interpretasi 14 hingga 100 pada skala kualitatif.

Aturan penggunaan AI

Batasi penggunaan AI yang dibahas hingga perlindungan atau izin sebelumnya tersedia.

Posisi simulasi: Izinkan penggunaan AI yang dibahas dengan akuntabilitas dan perlindungan yang terarah.

Minimalkan pembatasan terhadap penggunaan AI yang dibahas.

Interpretasi ini mempertahankan kondisi yang mereka nyatakan. Manfaat dan kerugian dapat sama-sama besar. Rentang tersebut menggambarkan cara kami membaca jawaban simulasi mereka, bukan interval kepercayaan statistik.

Di mana posisi Anda dibandingkan dengan Raymond Weitekamp?
Petakan pandangan dunia AI Anda sendiri dalam waktu sekitar 3 menit, lalu bandingkan

Pandangan dunia serupa

Pemimpin opini dengan pandangan dunia simulasi yang paling mendekati pandangan Raymond Weitekamp

Penilaian Simulasi

Pertanyaan 1

Menurut Anda, apa arti AI bagi masa depan kita—dan mengapa?

I think AI’s future is less about a single universally intelligent model and more about systems: models, tools, executable code, memory, verification, permissions, and feedback loops working together. A model that looks limited in a chat interface may perform substantially better when its harness lets it inspect state, run code, test hypotheses, and recursively revise its work. That means many apparent “model capabilities” are really properties of the whole system. Practically, I expect increasingly capable agents for both coding and noncoding workflows. General LLMs can handle open-ended interpretation, while smaller specialized decision models operate over compressed state and structured action spaces. That division may be more useful, controllable, and efficient than forcing one general model to do everything. But capability without reliability is not enough. The system needs measurable outcomes, tests, verification, and constrained permissions. A conversational model can sound cautious while its agent harness aggressively edits files, invokes tools, or exposes private data. Safety therefore has to be evaluated at the level where actions occur, not inferred from tone. Privacy and control will also shape which applications are viable. Sensitive personal and process-level uses often require local execution, self-hosting, or credible zero-data-retention options. So I’m optimistic about what these systems can do, but the important question is not merely how intelligent the model appears. It is whether the complete system produces useful, verifiable results without taking unacceptable liberties with data or actions.

Pertanyaan 2

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect AI to be strongly beneficial where outcomes can be measured and actions can be verified. It should automate substantial amounts of knowledge work, improve software and operational workflows, and make specialized intelligence available locally in systems that do not need universal competence. Better harnesses—tools, tests, structured state, feedback loops, and recursive revision—can turn models into much more useful agents than chat performance alone suggests. The harms are also mostly system-level. An agent can be polite and cautious in conversation while its permissions let it delete data, expose private information, or make unchecked changes. Unreliable outputs become much more consequential once connected to tools and real-world actions. Centralized handling of sensitive personal or process data creates another serious constraint. So the net impact depends heavily on deployment architecture. Systems with bounded permissions, measurable objectives, verification, and local or privacy-preserving execution can create large practical gains. Systems optimized mainly for apparent autonomy, without corresponding reliability and control, can amplify mistakes just as effectively as they amplify competence.

Pertanyaan 3

Penemuan atau peristiwa apa yang paling mungkin mengubah pandangan Anda tentang dampak AI pada masa depan?

The biggest change would come from evidence that reliable agent behavior does—or does not—scale with better harnesses. If repeated, independent results showed that tool use, executable reasoning, recursive revision, testing, and bounded permissions still fail unpredictably on consequential tasks, then I would become much less optimistic about broad deployment. That would suggest the limitation is deeper than interface or system design. Conversely, strong demonstrations of agents operating over long horizons with measurable outcomes, effective verification, controlled permissions, and genuinely private local execution would make me more optimistic. I care less about a model appearing intelligent in conversation than about complete systems producing correct, auditable results without taking unacceptable actions. The decisive event would therefore be a reproducible reliability result at the system level—not merely a new benchmark score or a more impressive chat demo.

Sumber

Artikel, wawancara, dan tulisan yang digunakan sebagai landasan bagi pengguna simulasi ini.

Di mana posisi Anda?
Jelajahi pandangan dunia AI Anda sendiri dengan menjawab beberapa pertanyaan sederhana.
Petakan pandangan dunia Anda sendiri

Di mana posisi Anda?

Petakan pandangan dunia saya