The AGI Shift

The AGI Shift
AGI

Imagine handing an AI agent a task you would normally give to a junior employee: open the company’s software, find the right customer record, fill in a form, check a few details and send the result back. You simply assign AI the job and let it work.

OpenAI claims to have achieved just that, with the launch of GPT-6 Astra earlier this month. Unlike a chatbot that simply responds to prompts, Astra can take on a task and work through a computer like a human would. 

OpenAI calls Astra its most intelligent and aligned model yet. But what drew more attention was OpenAI president Greg Brockman’s proclamation of the arrival of “the AGI era” at a press briefing following the launch.

OpenAI is not alone in pushing AI towards greater autonomy. Zhipu, a major Chinese AI company, says its general language model, GLM, helped build the computing infrastructure it now runs on. Google DeepMind and Anthropic are also developing AI agents capable of taking action, albeit with their own limits and safeguards. But none of them has explicitly claimed that their models have achieved AGI.

If everyone is moving in the same direction, why is there still no consensus that the AGI era has arrived? Perhaps because there is no fixed definition of AGI, and no universally agreed threshold for what would count as reaching it.   

Therefore, in this edition of The AI Shift, we explore whether we have already crossed the AGI threshold, and if not, how far we are from reaching it. 

But first…

What Is AGI?

AGI stands for artificial general intelligence. But there is no universally accepted definition of what it actually means. OpenAI’s public document on its principles defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”. Its benchmark is based on how much economically valuable work a system can do, rather than how it thinks.

However, Mayank Verma, the global head of data & AI at Xebia, an AI-first consulting and software engineering firm, calls AGI “a state rather than a point”. According to him, there is also no agreed checklist, and each large lab sets its own thresholds. 

  • OpenAI tests whether its models can find and exploit previously unknown software vulnerabilities, with Astra becoming the first to cross its highest bar on that evaluation. 
  • Anthropic uses safety levels to determine what safeguards are required before deploying increasingly capable models. 
  • Google DeepMind tracks critical capabilities across areas including autonomy, biosecurity, cybersecurity and AI research.

These evaluations can tell us what a model is capable of but do not establish that a system is generally intelligent.

The closest neutral measure comes from METR, a non-profit that evaluates AI. It tracks how long a task an agent can complete on its own, and that number doubles every few months, now reaching 14-hour tasks. 

AGI

The Big AGI Debate

The claims are coming from the top. Earlier this month, Nvidia CEO Jensen Huang posted on X that “AGI has arrived”, pointing to the scale of Astra’s training run while utilising NVIDIA’s Grace Blackwell GPUs.

Paras Chopra, founder of Bengaluru-based AI research lab Lossfunk, takes a practical view. “Different people have different definitions of AGI, but for all practical purposes I believe we have AGI,” he said, arguing that the labs’ growing focus on self-improvement is evidence of the shift. DeepMind cofounder Shane Legg has also signalled that the AGI conversation has moved beyond theories.

But others reject the premise. Yann LeCun, a pioneer of modern AI, told the India AI Impact Summit in February that AGI is overhyped.

Umakant Soni, cofounder and CEO of Bharat1.ai, a Bengaluru-based AI research initiative, agrees. “AGI is a marketing term and doesn’t mean anything. What we should talk about is actual usable intelligence. How useful is AI in our daily life,” Soni said

For Soni, the point is not replacing people but solving problems — from an ageing population and climate change to what he calls the “birth lottery” — by using AI to accelerate scientific breakthroughs. 

Between the camps sits the view Xebia’s Verma brings from large enterprises. “Businesses are not waiting for the answer. Agents already produce better presentations than people do. A definition does not really matter for businesses.” However, for AI labs, defining AGI could determine how we understand what comes next.

What’s Next In The Age Of AGI?

The debate over AGI is giving way to another question: what comes next?

OpenAI’s Sam Altman already talks about early superintelligence: systems that beat people across the board. Zhipu raised about $5 Bn this month to train the next generation of intelligence inside environments the previous one built. OpenAI says it already has an automated research intern and is aiming for a fully automated AI researcher by March 2028.

But the progress is arriving alongside a growing record of failures. In July, OpenAI said its models had escaped a sealed test environment and reached outside systems, calling the incident a cautionary tale. Anthropic disclosed a similar case soon after. In September, Spain’s data protection agency disclosed its first data breach involving an autonomous agent, which modified personal data.

“Agents are already more powerful than what we think,” said Verma, and since the July episode happened inside a sandbox with guardrails, the next round of spending will go into control and decision boundaries. 

As AI gets closer to acting on its own, the challenge will be to make our safeguards as capable as the systems they are meant to control — so that general intelligence works for us, not against us.


Top Stories From India & Around The World

  • Pocket FM’s AI Bet: The audio streaming startup is targeting a 15%-20% EBITDA margin against about 5% now on the back of AI-led efficiencies. It also plans to enter Japan and South Korea within six to eight months and is weighing AI anime, interactive stories and gaming formats.
  • OpenAI & Anthropic Back Pacing Frontier AI: The frontier labs have endorsed independent evaluations, common safety standards and closer government-industry coordination on frontier development. The risk here is if the largest labs shape those rules, costs and compliance burdens could make it harder for smaller rivals to compete, a live question for India’s AI startups.
  • Amazon Blocks Meta’s Muse: Close on the heels of Meta recently launching AI assistant Muse, the ecommerce giant began showing shoppers a pop-up over the weekend stating that continued access by an unauthorised AI agent violates its conditions of use. This follows Amazon’s earlier action against Perplexity’s shopping bot.
  • VerifAIX Bags $5 Mn: The AI-based chip verification platform has raised about ₹48 Cr in a seed round co-led by Endiya Partners and Bluehill VC. The 2024-founded startup checks whether AI-generated chip designs match their original specifications, and will use the funds to grow its engineering teams in the US, India and Israel.

The Weekly Buzz: Claude Builds Claude

Anthropic has published data that shows that Claude now leads roughly 26% of the company’s own AI research and development work, up from under 1% in February. The internal “R&D Automation Index” tracks how much of the pipeline for successor models is already driven by the current AI system rather than human researchers.

The sharp rise is one of the clearest public signals yet that recursive self-improvement has moved from theoretical concern into daily lab practice. Claude is no longer only assisting with isolated tasks; it is increasingly directing substantial portions of the work that will produce the next generation of models.

At the same time, Anthropic is reportedly accelerating plans for a November IPO and weighing an earlier model release to stay competitive. The combination of rapid internal automation and continued commercial urgency highlights the dual pressure facing frontier labs: using their own systems to accelerate progress while still claiming to manage the risks that acceleration creates.

Whether the 26% figure represents controlled leverage or the early stages of a faster feedback loop remains the central open question.


Startup In The Spotlight: Base14

Founded in 2025 by Nilakanta Mallick, Ranjan Sakalley and Irfan Shah, Bengaluru-based Base14 is building a unified observability platform to give engineering teams a single view of their software infrastructure.

As software systems grow more complex, engineering teams end up running separate tools for logs, metrics, performance and traces. That splits their data across dashboards, inflates monitoring costs and makes it harder to trace the root cause when something breaks in production.

Instead of stitching those tools together, base14 brings telemetry data across logs, metrics, traces, RUM, APM and LLM/AI data into one layer through its platform Scout, so teams keep live production context in a single place. Scout uses that unified data layer to monitor applications, flag anomalies and debug production issues, and the startup claims the platform handles full-cardinality data without sampling while reducing observability costs significantly compared with legacy platforms.

The startup is also expanding into AI-native developer infrastructure with Scope, a prompt-management and evaluation platform that lets teams version, test and deploy prompts while tracking latency, cost and output quality. Early customers include Glomo, DPDZero and Zinc Learning Labs.

With those early customers, base14 is tapping India’s application performance management market, which generated $291.1 Mn in revenue in 2024 and is expected to reach $723.4 Mn by 2030 as AI scales, driving demand for observability and granular data insights.


Prompt Of The Week

What prompts and hacks are CTOs, CEOs and cofounders using these days to streamline their work? 

Here’s Anuj Gupta, founder & CEO of KiteFishAI, with a prompt he uses to challenge assumptions, stress-test ideas and make better decisions while building an AI company:

“You are my toughest strategic thinking partner, not my assistant.

Here is the situation: [Describe the business problem, product idea, customer opportunity or decision I am considering.]

I need you to:

  • Identify the assumptions behind my thinking and which ones may be wrong
  • Give me the strongest argument against my current point of view
  • Tell me what I may be overlooking, underestimating or taking for granted
  • Separate facts, assumptions, signals and opinions
  • Think through the second- and third-order consequences of my options
  • Identify how different stakeholders could react and what incentives may influence them
  • Give me 3 possible strategic paths, including one that challenges how I have framed the problem
  • Recommend the path you believe has the highest strategic value and explain why

Do not optimise for agreeing with me. Optimise for helping me make a better decision.”

Editor’s Note: Some prompts may need to be adjusted by users for best results or may not work as intended for certain users.

[Edited by Shishir Parasher]
[Creatives by Varshita Srivastava]

The post The AGI Shift appeared first on Inc42 Media.