What is AGI? Nobody agrees, and it’s tearing Microsoft and OpenAI apart.

The reported $100 billion profit threshold we mentioned earlier conflates commercial success with cognitive capability, as if a system’s ability to generate revenue says anything meaningful about whether it can “think,” “reason,” or “understand” the world like a human.

Sam Altman speaks onstage during The New York Times Dealbook Summit 2024 at Jazz at Lincoln Center on December 04, 2024 in New York City. — Sam Altman speaks onstage during The New York Times Dealbook Summit 2024 at Jazz at Lincoln Center on December 4, 2024, in New York City.

Credit:

Eugene Gologursky via Getty Images

Depending on your definition, we may already have AGI, or it may be physically impossible to achieve. If you define AGI as “AI that performs better than most humans at most tasks,” then current language models potentially meet that bar for certain types of work (which tasks, which humans, what is “better”?), but agreement on whether that is true is far from universal. This says nothing of the even murkier concept of “superintelligence”—another nebulous term for a hypothetical, god-like intellect so far beyond human cognition that, like AGI, it defies any solid definition or benchmark.

Given this definitional chaos, researchers have tried to create objective benchmarks to measure progress toward AGI, but these attempts have revealed their own set of problems.

Why benchmarks keep failing us

The search for better AGI benchmarks has produced some interesting alternatives to the Turing Test. The Abstraction and Reasoning Corpus (ARC-AGI), introduced in 2019 by François Chollet, tests whether AI systems can solve novel visual puzzles that require deep and novel analytical reasoning.

“Almost all current AI benchmarks can be solved purely via memorization,” Chollet told Freethink in August 2024. A major problem with AI benchmarks currently stems from data contamination—when test questions end up in training data, models can appear to perform well without truly “understanding” the underlying concepts. Large language models serve as master imitators, mimicking patterns found in training data, but not always originating novel solutions to problems.

But even sophisticated benchmarks like ARC-AGI face a fundamental problem: They’re still trying to reduce intelligence to a score. And while improved benchmarks are essential for measuring empirical progress in a scientific framework, intelligence isn’t a single thing you can measure, like height or weight—it’s a complex constellation of abilities that manifest differently in different contexts. Indeed, we don’t even have a complete functional definition of human intelligence, so defining artificial intelligence by any single benchmark score is likely to capture only a small part of the complete picture.

Source link

What's Hot

Condé Nast, other news orgs say AI firm stole articles, spit out “hallucinations”

Moveworks Surpasses $100M ARR as Its AI Copilot Boosts Productivity for Millions of Employees Worldwide

Report: Mistral AI in talks with MGX on $1B equity financing round

What is AGI? Nobody agrees, and it’s tearing Microsoft and OpenAI apart.

Grok praises Hitler, gives credit to Musk for removing “woke filters”

Mike Lindell lost defamation case, and his lawyers were fined for AI hallucinations

IBM Power11 Raises the Bar for Enterprise IT

Pioneer Works Hosts a MSCHF Sculpture You Can Take Home by the Inch

Lorna Simpson’s Former Home and Art Studio Hits the Market for $6.5 M.

Prospect New Orleans Will Not Mount Next Edition in 2027

New London Art Fair for Women-Led Galleries, And More: Morning Links

Condé Nast, other news orgs say AI firm stole articles, spit out “hallucinations”

Moveworks Surpasses $100M ARR as Its AI Copilot Boosts Productivity for Millions of Employees Worldwide

Report: Mistral AI in talks with MGX on $1B equity financing round

What's Hot

What is AGI? Nobody agrees, and it’s tearing Microsoft and OpenAI apart.

Why benchmarks keep failing us

Related Posts

Subscribe to Updates