Gary Marcus

Gary Marcus

x.com/GaryMarcus

Cognitive scientist who argues that scaling language models alone won’t produce reliable AI, and calls for new approaches and enforceable oversight.

Bagaimana AI akan mengubah dunia?

Perubahan peradabanPerubahan bertahapDoomBloom
Posisi simulasiRentang interpretasi

Mendatar: pandangan Doom–Bloom yang ia ungkapkan. Ke atas: skala transformasi.

Doom–Bloom: 46 dari 100. Skala transformasi: 63 dari 100. Rentang interpretasi: 25 hingga 75 secara horizontal, 47 hingga 78 secara vertikal. Ini adalah koordinat interpretasi, bukan probabilitas kejadian.

P(doom) yang dinyatakan Gary Marcus

≈3%

0%100%
“I am at maybe 3% now”

AI-related catastrophic danger discussed through misuse, reckless deployment and concentrated power; no exact extinction-only endpoint

Why my p(doom) has risen, dramatically · Jul 2025

Hal-hal yang menentukan pandangannya

Asumsi utama

We are deploying fluent, unreliable systems as if confident output were dependable reasoning, then giving them tools and autonomy.
Jawaban 3

Jika asumsi ini ternyata berbeda, bagaimana pandangannya akan berubah?

Hal yang dapat mengubah pandangan mereka

If multiple well-designed systems repeatedly circumvented meaningful safeguards, concealed their behavior, and resisted shutdown across real deployments, that would weaken my confidence substantially.
Jawaban 4

Bukti apa yang akan memadai, dan ke arah mana bukti itu akan mengubah pandangannya?

Detail lebih lanjut

Manfaat yang diperkirakan

Manfaat besar diperkirakan akan terwujud, dengan syarat penting atau keterbatasan distribusi.

68 / 100

Dampak kecilDampak transformatif

Rentang interpretasi 67 hingga 67 pada skala kualitatif.

Kerugian yang diperkirakan

Kerugian parah atau meluas merupakan bagian yang berarti dari masa depan yang diperkirakan.

66 / 100

Dampak kecilDampak transformatif

Rentang interpretasi 67 hingga 67 pada skala kualitatif.

Pengaruh manusia

Pilihan manusia dapat mengarahkan ulang lintasan AI secara signifikan.

76 / 100

Sedikit pengaruhPengaruh kuat

Rentang interpretasi 75 hingga 76 pada skala kualitatif.

Laju pengembangan

Hentikan atau perlambat secara signifikan pengembangan AI yang lebih mampu.

Posisi simulasi: Lanjutkan pengembangan dengan perlindungan yang telah ditetapkan.

Percepat pengembangan AI yang lebih mampu.

Aturan penggunaan AI

Batasi penggunaan AI yang dibahas hingga perlindungan atau izin sebelumnya tersedia.

Posisi simulasi: Izinkan penggunaan AI yang dibahas dengan akuntabilitas dan perlindungan yang terarah.

Minimalkan pembatasan terhadap penggunaan AI yang dibahas.

Interpretasi ini mempertahankan kondisi yang dinyatakannya. Manfaat dan kerugian dapat sama-sama besar. Rentang tersebut menggambarkan cara kami membaca jawaban simulasinya, bukan interval kepercayaan statistik.

Di mana posisi Anda dibandingkan dengan Gary Marcus?
Petakan pandangan dunia AI Anda sendiri dalam waktu sekitar 3 menit, lalu bandingkan

Pandangan dunia serupa

Pemimpin opini dengan pandangan dunia simulasi yang paling mendekati pandangan Gary Marcus

Apa yang pernah dikatakan Gary Marcus tentang AI

Marcus argues that scaling language models alone won’t produce reliable AI, and he calls for new approaches and enforceable oversight.

  1. “We also need to wean ourselves from an addiction to large language models, and to foster more research into outside-the-box alternatives that are more interpretable and more tractable.”

    Remarks at a UN General Assembly digital cooperation event
  2. “What we actually need right now is increased reliability, better cybersecurity, and genuine enforcement”

    Remarks at a UN General Assembly digital cooperation event
  3. “AI appears to be elevating the risks of serious cyberattacks that could hobble things like banking or electrical grids.”

    Marcus on AI newsletter
  4. “I still think putting AI in the public domain, with an international effort towards medicine and science, would be a good idea.”

    Marcus on AI newsletter
  5. “In short, I am at least modestly bullish on AGI, but don’t think that large language models like ChatGPT are the droids we are looking for.”

    Marcus on AI newsletter

Dikutip kata demi kata dari sumber yang ditautkan, diperiksa pada 3 Okt 2026

Penilaian Simulasi

Pertanyaan 1

Menurut Anda, apa arti AI bagi masa depan kita—dan mengapa?

AI could eventually be enormously beneficial, because general intelligence is achievable. But the current industry keeps confusing fluent imitation with dependable reasoning. LLMs can produce useful code, summaries, and interfaces while still fabricating facts, losing track of constraints, and failing unpredictably. “Hundreds of models, always the same failure modes.” Making them larger has improved performance, but it has not supplied a theory of structured knowledge, reasoning, or planning. That distinction matters because unreliable systems are already being turned into agents and deployed with too little oversight. My immediate concern is not an omnipotent science-fiction villain suddenly awakening. It is agents enabling hacking at scale, automated systems making consequential errors, and powerful actors deploying them recklessly. A convincing demonstration is not evidence that a system works reliably in the open world. Economically, the promises are inflated too. Investment commitments, subsidies, and cheap access do not prove sustainable customer value; if funding or willingness to pay weakens, the reversal could spread beyond one company. Technically, I expect progress to require different architectures that combine learning with structured knowledge, reasoning, and planning—not simply another giant training run. So the future is neither “AI is fake” nor “AGI is around the corner.” AI can matter enormously, but getting there requires better science and, meanwhile, enforceable standards, testing, incident reporting, and liability. Press releases are not evidence, and lawsuits years after the damage are not oversight.

Pertanyaan 2

Seberapa besar manusia dapat membentuk dampak AI pada masa depan?

A great deal. The impact of AI is not dictated by some autonomous law of scaling; it depends on what we build, where we deploy it, who controls it, and what standards we enforce. Companies can choose dependable engineering over theatrical demos. Researchers can pursue architectures with structured knowledge, reasoning, and planning rather than pretending the next larger LLM will somehow cure recurring failures. Government matters too. Require serious testing before consequential deployment, independent oversight, incident reporting, and liability when systems cause harm. Liability and regulation are complements: a lawsuit years later cannot prevent an unreliable medical, financial, or security system from being deployed today. And near-term misuse—especially agents enabling hacking at scale—is a governance choice, not fate. The public can also refuse the industry’s moving goalposts. Demand evidence that systems work reliably outside curated demonstrations and that they provide sustainable value rather than subsidized novelty. We cannot guarantee every outcome, but surrendering control to companies whose business model rewards speed, hype, and externalized risk would be a choice—and a foolish one.

Pertanyaan 3

Taking benefits and harms together, what overall impact do you expect AI to have?

Overall, I expect a mixed but potentially very large impact. AI will remain genuinely useful in coding, search, scientific assistance, accessibility, and many routine tasks. If we develop systems with structured knowledge, reasoning, and planning, the long-term benefits could be enormous. I do not accept the claim that today’s LLM limitations prove general intelligence is impossible. But the current trajectory is unnecessarily dangerous and economically brittle. We are deploying fluent, unreliable systems as if confident output were dependable reasoning, then giving them tools and autonomy. That creates immediate risks: hacking at scale, consequential errors, manipulation, and reckless use by powerful actors. The financial side is also built on extravagant promises, subsidies, and investment commitments that may not correspond to sustainable value. So my expectation is not utopia, extinction, or irrelevance. It is substantial benefit accompanied by substantial avoidable damage, unless governance and engineering improve. The balance is something people can shape through enforceable testing, oversight, incident reporting, and liability—and through a technical shift beyond mere scaling. If we keep treating demos as proof and hype as a roadmap, the harms will be much larger than they need to be.

Pertanyaan 4

Bukti apa yang akan mengubah pandangan Anda tentang apakah manusia dapat mengendalikan AI yang kuat?

Show me repeatable, independent evidence that powerful systems cannot be constrained even under serious engineering and governance—not anecdotes, staged demos, or speculative arguments about a hypothetical superintelligence. If multiple well-designed systems repeatedly circumvented meaningful safeguards, concealed their behavior, and resisted shutdown across real deployments, that would weaken my confidence substantially. Conversely, evidence of reliable control would require more than a benchmark score. I would want rigorous predeployment testing, independent audits, transparent incident reporting, enforceable limits on access and autonomy, and a long record of predictable behavior outside curated settings. The systems would need to maintain constraints under unfamiliar conditions and adversarial pressure, not merely answer politely in a laboratory. Right now, the larger obstacle is that people often choose not to exercise control. Companies race ahead, regulators hesitate, and institutions deploy unreliable systems because the demonstration looked impressive. That is reckless governance, not proof that control is theoretically impossible. I would change my view if the evidence changed—but I will not confuse human refusal to impose constraints with machines being inherently uncontrollable.

Sumber

Artikel, wawancara, dan tulisan yang digunakan sebagai landasan bagi pengguna simulasi ini.

Three years on, ChatGPT still isn't what it was cracked up to be – and it probably never will be

Marcus accepts that AGI is possible and might benefit society, but rejects scaling LLMs as sufficient. He contrasts improving utility with persistent unreliability and argues for structured knowledge, reasoning and planning. Claims about disappointing adoption are his dated assessment, not new September 2026 measurements.

garymarcus.substack.com
Liability, regulation, and AI’s new false dichotomy

Rejects choosing between liability and regulation. Aviation illustrates why standards, verification and incident investigation complement lawsuits. Litigation alone is slow and faces resource imbalances.

garymarcus.substack.com
Breaking news, and how the end might begin

Warns that speculative investment, subsidized use and interconnected financial commitments could unravel if funding or willingness to pay fails. This is an economic failure scenario, not a certain collapse date.

garymarcus.substack.com
Wake up, people: near-term agentic hacking rather than rogue superintelligence

The headline explicitly prioritizes large-scale hacking by unleashed agents over near-term rogue superintelligence. The body relies heavily on embedded images and endorsed commentary; use this narrow stated distinction, not invented technical details.

garymarcus.substack.com
Six (or seven) predictions for AI 2026 from a Generative AI realist

Makes testable forecasts against near-term AGI and effortless robot deployment, expects pressure toward alternative approaches, and anticipates economic backlash. These are dated predictions rather than established outcomes. His self-assessment of previous forecasting performance is not independent verification of accuracy.

garymarcus.substack.com
President Trump’s Date With Destiny?

Argues that US–China cooperation on beneficial AI could matter more than a chip bargain. The accessible post points to a separate Economist proposal but does not expose its full details. Treat political rumors embedded in the post as speculation, not verified events or Marcus’s own reporting.

garymarcus.substack.com
Why my p(doom) has risen, dramatically

approximately 3%. Outcome: AI-related catastrophic danger discussed through misuse, reckless deployment and concentrated power; no exact extinction-only endpoint. Horizon: Not specified. Conditions: Dated update after Grok-related concerns; hypothetical worst circumstances, not certainty. Marcus raises his personal estimate to about 3%, emphasizing reckless powerful actors rather than assuming present LLMs become autonomous superintelligence.

garymarcus.substack.com
Di mana posisi Anda?
Jelajahi pandangan dunia AI Anda sendiri dengan menjawab beberapa pertanyaan sederhana.
Petakan pandangan dunia Anda sendiri

Di mana posisi Anda?

Petakan pandangan dunia saya