Enterprise AI Spend Gets A Reality Check

The cost of using AI is falling rapidly, with the price of processing a million tokens dropping by roughly 10X every year. The new price benchmark is being set across the market: DeepSeek V4 Flash charges $0.14 per Mn input tokens, OpenAI’s GPT-5.6 Luna costs $0.20, while Meta’s Muse Spark 1.2 is priced at $1.25 for the same volume.
On the surface, this looks like a simple equation: cheaper intelligence should mean more AI. But inside enterprises, the conversation is moving in a different direction.
As the cost of inference falls, the bigger question is no longer how much AI companies can afford to use, but where they should use it, and whether the returns justify the spend.
Tech leaders are increasingly looking beyond the price of a token to its productivity gains, revenue impact and business outcome. They are putting budgets around AI features, tracking which teams are consuming the most tokens, deciding which tasks deserve expensive frontier models and which can be handled by cheaper alternatives.
While the AI cost curve is collapsing, the bar for proving value is rising. So, as intelligence gets cheaper, what makes an AI token worth spending? That’s the question we’re digging into in today’s edition of The AI Shift.
New Token Economics
As humans, the first instinct is always to treat cheaper tech as a licence to use more of it, and AI is no exception. However, the leaders we spoke with are approaching the shift differently: the goal is not to use fewer tokens at any cost, but to prevent tokens from being spent on work that does not matter.
According to Ed Huang, CTO and cofounder of TiDB, an open-source database company, token cost is no longer a problem for enterprises. They are asking more important questions: “Do we know where our tokens are going, and are they producing enough value to justify their usage?”
That value-first approach is now translating into basic financial controls inside AI teams. At Eightfold AI, an AI talent platform, no feature can use an AI model until it is approved.
Thiyagaraj T, director of engineering at the company, said a team that wants to build something with AI must explain what it is making, which model it will use and what it will cost per month. Platform owners can then challenge the estimate before approval. In one case, a projected cost was reduced from $2,000 to $400, an 80% cut, before the experiment began.
What we understood from our conversations with experts is that organisations are now considering AI unit economics before production. At runtime, each team’s usage is being capped with its own monthly budget, and every call is logged and compared with the original estimate. AI consumption is therefore being treated less like an open-ended experiment and more like a budget request that has to earn its place.
According to Anuraag Kochhar, CTO of AI services company ShepHertz, model efficiency and falling token prices have already cut software development costs by up to 30% on some projects. But the bigger shift is that organisations are no longer just looking for cheaper ways to do the same work. They are rethinking how AI projects are approved, measured and tied to business outcomes.
AI Budgets Are Getting Guardrails
As organisations gain control over AI usage, they are changing how they budget for it. The question is no longer simply how much to spend on AI, but which features deserve access to that budget.
While many enterprises are concerned about rising token consumption, their financial controls have not kept pace. This is because fewer than two in three organisations currently actively monitor usage with clear budget limits.
Therefore, there is a growing enterprise interest in Chinese and other open-weight models. While open-weight models are not replacing premium systems, companies are using them selectively for high-volume, less sensitive workloads.
According to Kochhar, 25-40% of Indian enterprises are piloting or partially deploying such models for customer support, content generation and internal tools. Depending on the use case, they can reduce inference costs by 30-70%, although compliance, data governance and reliability remain barriers to wider adoption.
The approach is simple: use cheaper models for routine tasks and reserve expensive frontier models for complex reasoning or sensitive work.
Operational Cost Or Capital Bet?
How companies budget AI today will determine where it appears on the balance sheet tomorrow. Unlike traditional software, which usually comes with a predictable per-seat subscription, AI spending is consumption-based and spread across product, research, support and cloud infrastructure.
For most companies, token costs will initially remain part of operating expenses, much like cloud spending. But AI’s wider cost base, including model access, data preparation, engineering, evaluation, monitoring, security and human review, may push organisations to track it as a separate internal category, even if accounting systems do not.
As companies build more software in-house, engineering costs may move towards capitalised development, while GPUs and long-term cloud commitments could be treated as capital expenditure. The immediate shift, however, will happen within operating budgets, where teams will have to justify AI usage alongside headcount, cloud infrastructure and software licences.
AI may not appear as a single line on the balance sheet, but companies will increasingly need to treat it as one internally. For Indian enterprises, the real benchmark is whether every token can be linked to a rupee of value. Those that build this traceability will manage AI as an investment; the rest will continue treating it as a bill.
Top Stories From India & Around The World
- Cars24’s Enterprise AI Bet: The used-car marketplace has launched Deployment Inc, a company that helps big businesses put AI to work in their daily operations. It comes with $5 Mn from Cars24 and plans to hire 50 engineers.
- Eternal Spins Out Nugget: The foodtech giant has spun off its AI platform into a separate business. Nugget, which handles customer support and sales calls for banks, manufacturers and other enterprises, earned ₹7.2 Cr last year and is cash-flow positive.
- OpenAI Restricts Astra Over Cyber Concerns: OpenAI has restricted work on Astra, its next major AI model, over concerns it could carry out cyberattacks on its own. The company is holding the model back until it is confident it is safe.
- Sarvam To Raise $74 Mn: The newly minted AI unicorn is planning to raise ₹74 Mn from NVIDIA, Glade Brook, Gaja Capital and Indigo Ventures as part of its ongoing $300 Mn+ Series B round. Earlier this year, Sarvam had bagged $234 Mn in a funding round led by HCL Tech.
- Centre Infuses ₹276 Cr Into AI COEs: The funds were disbursed to set up four centres of excellence in AI, which are being set up at IIT Ropar, Kanpur and Madras, and IISc Bangalore. The CoEs have already published research papers and developed AI models in partnership with industry players.
The Weekly Buzz: AI Just Designed Working Viruses From Scratch
Researchers at Stanford and the Arc Institute used genome language models to design complete bacteriophage genomes, the first time AI has generated fully functional viruses end-to-end. Of hundreds of AI-written designs that were chemically synthesised and tested in the lab, 16 successfully infected and killed a group of bacteria, with several outperforming the natural strains they were modelled on in both replication speed and killing efficiency.
The models, Evo 1 and Evo 2, trained on millions of viral genomes, produced sequences carrying dozens to hundreds of novel mutations never seen in nature. Some of the AI-designed phages even incorporated functional genetic modules from distant evolutionary relatives, something human designers would struggle to achieve rationally.
The work, published in the Science journal, demonstrates that generative AI can now move beyond proteins and individual genes to design entire viable genomes. The breakthrough opens a promising path for phage therapies against antibiotic-resistant bacteria, while simultaneously raising sharp biosecurity questions. Critics note that the same techniques could theoretically be directed at more dangerous pathogens.
Startup In The Spotlight: eyecandy robotics
Founded in 2025 by Alqama Shaikh, Raghuvamsi Velagala and Mankaran Singh, Bengaluru-based eyecandy robotics builds physical AI characters. As consumer robotics fills shelves with utilitarian assistants and vacuums, the startup is addressing the fact that very little in the category actually gives people something to form an emotional bond with.
Instead of another functional gadget, eyecandy is chasing a rare fusion of robotics, AI and entertainment, an offshoot of the ongoing increase in robotics deployments. The founders’ stated ambition is to make their characters 10X more engaging than anything on shelves today, with lifelike, responsive personalities at the heart of the product.
The startup is targeting the $16 Bn global consumer robotics market, a category expanding quickly as generative AI makes lifelike, responsive characters commercially viable.
As generative AI brings down the cost of realistic, responsive interaction, the startup believes that consumers will increasingly favour robots built for emotional engagement over purely functional devices, opening the door to a new category of entertainment-first hardware.
Prompt Of The Week
What prompts and hacks are CTOs, CEOs and cofounders using these days to streamline their work?
Here’s the prompt Manish Choudhary, cofounder and CEO of Flexprice, uses to turn raw sales call notes into a structured coaching breakdown instead of generic feedback.
“You are a founder-first sales coach for B2B startups selling to mid-market and large enterprises. I’ll share an anonymous call breakdown (call frame, talk-track skeleton, outcome so far). Treat me as an operator, not a beginner.
Produce three sections, no preamble:
- Critical review: Name the biggest failure mode and why it risks the deal. List 5-7 material problems, each with the exact moment, why it’s risky, and what better looks like. Flag high-risk vs low-risk issues. Give 3-5 prioritised fixes.
- Ideal call reconstruction: Rebuild as a 45-60 min time-blocked flow: discovery, positioning, demo, pricing, mutual action plan. Show how to handle pushback in the moment.
- Follow-up playbook: Pre-call prep, an opening that surfaces objections upfront, two scoped proposal options, pilot framing with success criteria, and a close with written recap plus a 1-10 commitment question.
Be specific and practical, no jargon, ground everything in the call notes, optimise for real deal outcomes.”
Editor’s Note: Some prompts may need to be adjusted by users for best results or may not work as intended for certain users.
Edited by: Shishir Parasher
Creatives by: Varshita Srivastava
The post Enterprise AI Spend Gets A Reality Check appeared first on Inc42 Media.


Superadmin 










