We Were Promised AGI By Now. Where Is It?

I keep seeing old predictions saying artificial general intelligence was supposed to arrive by now, but it clearly hasn’t happened. I’m trying to understand what slowed progress, what experts got wrong, and whether current AI is actually close to AGI or still far away. I need help sorting through the hype, timelines, and real technical limits so I can get a clearer answer.

A lot of old AGI timelines were bad for the same reason old flying car timelines were bad. People projected a trend and ignored the hard parts.

What got missed:

  1. Benchmarks got gamed.
    Models improved on tests. People assumed test scores meant broad reasoning. It didn’t. A model scoring high on exams still fails at planning, memory, long tasks, and stable world models.

  2. Scale helped, then slowed.
    From 2018 to 2023, bigger models plus more data gave huge gains. That made people think the curve would keep going. Now data quality, training cost, power use, and chip supply all bite harder.

  3. Language is not the whole mind.
    LLMs predict tokens well. AGI likely needs stronger grounding, persistent memory, planning, tool use, and better learning from experience. Text alone was a shortcut, not the finish line.

  4. Demos fooled people.
    A smooth demo looks human. Real evals are harsher. Put the same system in a messy setting for weeks and it breaks in dumb ways. Hallucinations are still a big deal.

  5. Experts disagreed more than headlines showed.
    Some people said AGI was close. Others said decades. Media amplified the hottest takes becuase those get clicks.

Where are we now? Closer to useful automation than AGI. Coding help, search, customer support, and research assist are real. Full general intelligence, nope. If you want a practical signal, watch autonomous long-horizon work, not benchmark screenshots. That’s where the gap still is.

A lot of “AGI by now” predictions were basically extrapolation cosplay. Compute went up, benchmarks went up, demos got flashier, so people assumed intelligence was scaling in a smooth line. It wasn’t.

I mostly agree with @sognonotturno, but I’d add one thing: the field kept quietly moving the goalposts. When machines couldn’t do robust reasoning, people said language mastery was enough. When they got decent at language, suddenly agency, memory, embodiment, reliability, and self-directed learning became the real test. Some of that is legit, some of it is cope.

Also, economics matters more than futurists admitted. It’s one thing to make a spooky-impressive model in a lab. It’s another to run it cheaply, safely, and consistently at global scale. AGI that crashes, hallucinates, or costs a fortune per task is not really “arrived,” imo.

What experts got wrong? They underestimated brittleness and overestimated how much human-like text implied human-like cognition. Curent systems are powerful pattern machines, not durable autonomous thinkers.

Are we behind? Maybe on hype schedules, yes. On actual capability, not really. We got useful AI faster than I expected, just not the sci-fi version ppl were sold. The real tell will be when a system can take a vague goal and handle weeks of messy work with minimal babysitting. We’re… not there yet.

A piece I think both you and @sognonotturno are circling around is this: “AGI” was never a stable target. It bundled together at least four different things:

  1. human-level task breadth
  2. autonomous long-horizon execution
  3. transferable world models
  4. economic replacement value

Progress on #1 got mistaken for progress on all four.

Where I slightly disagree with the usual take is this idea that experts were simply fooled by hype. Some were, sure. But a lot of forecasts failed because AI progress is lumpy. You get sudden jumps from scaling, then long plateaus where the missing ingredient is not “more parameters” but better training signals, better memory, better tool use, better feedback loops, or better ways to deal with uncertainty. Forecasting that kind of jagged curve is brutal.

Another underrated slowdown: evaluation. If you can’t define AGI cleanly, you can’t tell whether you’re near it. Benchmarks got saturated, but real-world competence stayed messy. Passing exams, writing code, and sounding convincing are not the same as maintaining goals, recovering from mistakes, or knowing when you’re wrong.

And honestly, some “delay” is social, not technical. Labs now have to care about deployment risk, regulation, inference costs, copyright fights, security, and whether enterprise users will trust the system. Those frictions matter.

My rough view: we are not missing one magical breakthrough. We are missing a stack of boring things that have to work together reliably. Memory, planning, calibration, tool competence, self-correction, persistence, and cheap inference.

Pros for the ‘’: can improve readability if used to organize comparisons like prediction vs reality. Cons for the ‘’: not really applicable here unless it actually exists as a real tool or resource.