Measuring Progress Toward AGI Requires More Than Hype

Google DeepMind wants to bring actual science to the chaotic race for Artificial General Intelligence. It’s about time....

Feed
October 6, 2026
Measuring Progress Toward AGI Requires More Than Hype


Everybody talks about Artificial General Intelligence like it's an impending weather event, but almost nobody can agree on how to actually measure the damn thing. It we are drowning in promo spin, cherry-picked benchmarks. and leaderboard scores that tell us next to nothing about whether a model can actually think, adapt, or reason its way out of a paper bag. When I saw DeepMind's latest cognitive framework paper attempting to drag the conversation out of the marketing department, that is why I breathed a quiet sigh of relief. Back into real reality.

Instead of inventing another arbitrary coding test or trivia gauntlet! This they dug into decades of — oddly — cognitive science and neuroscience to build a taxonomy of actual mental faculties. This we're talking about ten core pillars here: perception, generation, attention, learning. Sounds familiar? For what it's worth, memory, reasoning, metacognition, executive functions, problem-solving — and social cognition. It maps intellect not as a single flat number. But as a sprawling web of skills that looks suspiciously like what humans actually do every day.

Measuring Progress Toward AGI Requires More Than Hype

What I appreciate most — to be fair — about this protocol is the insistence on human baselines. It too many labs have baked review data straight into their training corpora, grading their own homework while pretending they aced the exam. Pretending the exam. Too many labs have baked evaluation data straight into their training corpora, grading own homework. The, by comparing machine outputs against a demographically representative sample of actual people across messy. Real-world cognitive tasks, we might finally puncture hype bubble. And see these systems for what they truly are.

Of course, turning a brilliant — to be fair – academic taxonomy into practical software is a whole other beast. That's why they're farming out the hardest parts. Like to build evaluations for metacognition and social cognition – to the crowdsourced chaos of a Kaggle hackathon. I'm skeptical that a Kaggle race will instantly solve the measurement crisis. But at least we are finally asking the right queries instead of just scaling up factors and crossing our fingers.