Nvidia CEO Declares AGI Achieved, Sparking Debate Over Definition and Benchmarks
Huang’s take on AGI was narrowly framed: he described it as the capacity of an AI system to run a billion‑dollar company autonomously. That view stands in contrast to the more common, expansive definition that an AGI can perform any intellectual task a human can. Because the term was left undefined, critics argue the claim muddles the conversation about what constitutes AGI and how progress should be measured.
In September 2026, OpenAI launched GPT‑6 Astra, its most sophisticated generative model to date. The offering comes in a standard version and a “Pro” variant, priced at $10 and $50 per million tokens, respectively. OpenAI markets Astra as a leap forward in computer use, coding, cybersecurity, and scientific reasoning. Benchmark tests that pit Astra against Claude Fable 5.1 show improved performance across a range of tasks, but the gains are not described as a jump toward human‑level general intelligence. According to OpenAI’s system card, Astra meets a “critical threshold” for certain security tasks, yet it still requires human oversight for many complex problems.
A pivotal metric for abstract reasoning and fluid intelligence is the ARC‑AGI test. The ARC‑AGI‑2 version, released in 2024, was designed to challenge models that had excelled on earlier iterations. The test evaluates a system’s ability to solve novel grid‑puzzle problems. Recent results indicate that GPT‑6 Astra scores roughly 4 % on the ARC‑AGI‑2 benchmark, whereas human participants achieve near 85 %. The disparity underscores that current models, including Astra, have not yet reached the level of reasoning that most definitions associate with AGI.
Huang’s declaration brings to the fore the field’s broader uncertainty. While Nvidia, OpenAI, Google, Meta, and others continue to push the envelope with increasingly capable models, there is no consensus on a single metric that can confirm AGI. Researchers stress the importance of rigorous, transparent benchmarks such as ARC‑AGI and call for clear definitions that distinguish narrow AI from general intelligence.
Regulators and policy makers are also paying close attention. The absence of a shared definition complicates efforts to assess safety risks, ethical implications, and potential regulatory requirements. Some experts argue that without a clear standard, it is difficult to determine whether a system poses an existential risk or merely represents incremental improvement.
As of September 2026, the AI community remains divided. Huang’s claim has not been corroborated by independent evidence, and the latest models, including GPT‑6 Astra, have not demonstrated the breadth of capabilities required for AGI. OpenAI continues to roll out Astra in phases, while other companies are releasing competing models around the same time. The coming months are likely to see additional benchmark releases, further model comparisons, and ongoing discussions about the definition and measurement of AGI.
In short, Nvidia’s CEO has asserted that AGI is achieved, but the claim lacks a clear definition and supporting data. Current leading models, such as OpenAI’s GPT‑6 Astra, show measurable improvements over previous generations yet fall short of the human‑level general intelligence most researchers associate with AGI. The field continues to evolve, with new benchmarks and model releases poised to shape the conversation in the near future.