The Case For Evals In A Multi-Model World

The Case For Evals In A Multi-Model World
The Case For Evals In A Multi-Model World

Another day and another new model grabs the limelight. Product teams scramble to see how it can be used for their own operations and systems. 

But a new model can look great in a demo and still make a product worse. 

It may cost less but respond too slowly, pass offline tests but fail with real customers, or behave differently once traffic and context change. That is why the engineering question is shifting from which model to use to how teams prove a model change is safe and worth making.

At Inc42’s The CTO Summit 2026 last week, leaders from Swiggy, Rapido, Razorpay, ShareChat, Shadowfax and Meesho described building evaluations or evals and monitoring into the model lifecycle. And as platforms and products grow, evals are quickly becoming a key differentiation in the AI strategy. 

The accounts of these unicorns and massive listed tech companies point to a practical shift in the product and software development lifecycle: model choice is becoming a continuing production decision, and evaluation is the evidence needed to make it. 

Evals, short for evaluations, are systematic tests used to measure how well an AI model or system performs a specific task.

But what should teams measure, and how can they tell whether the measurement itself is reliable? Let’s unravel it in today’s edition of The AI Shift.

Why Evals Are Becoming Critical

Swiggy’s experience shows the cost of not building for change early. Madhusudhan Rao, CTO at Swiggy, said an early contact-centre agent was tied to one model. When the provider faced capacity constraints, migrating took nearly a month. Swiggy now uses evaluations and experimentation to make model changes easier. “Treat the model as replaceable, invest in the workflow, the evals,” Rao said.

The point is not to test models in isolation. A switch has to be assessed against the task and the product: does it improve quality, latency or cost without disrupting the customer experience? Nitin Jain, CTO, ShareChat & Moj, described evaluations as an internal check and A/B testing as a way to observe customer impact. “Eval and A/B testing are just one or like different ways of doing the similar thing,” he said.

 A switch has to be assessed against the task and the product: does it improve quality, latency or cost without disrupting the customer experience?

Shadowfax offers a production example. Vaibhav Khandelwal, cofounder and CTO, Shadowfax, said the company was on its third model iteration for a delivery-partner support bot. Evaluations and observability let the team compare iterations on cost, speed and accuracy. The useful question for an AI team is not whether a model scored higher once, but whether the new version performs better on the product’s actual operating measures.

Those measures cannot be chosen from a generic checklist alone. A support bot may fail by repeating itself, going silent or misunderstanding a delivery partner, while a commerce assistant may fail to understand a request. Teams first need to examine real outputs and identify the specific errors that matter to their users.

Anand Jain, head of engineering, Meesho, said evaluations give the company confidence when comparing one model with another. The transcript does not specify Meesho’s evaluation metrics or thresholds, but his point highlights a gap CTOs should address: what does “better” mean for each workflow, and what level of regression is acceptable before a model change is rejected?

An evaluation system also needs validation. Eternal’s engineering team, in a post about its in-house Gavel evaluation infrastructure, says it compares automated judges with human review. 

At the time of writing in July 2026, Gavel had already evaluated more than a million items and some of its AI judges were close to human quality checks on roughly 90% of critical flaws. That example illustrates both the potential and the caveat: automated scoring can extend review, but teams still need to check whether the judge catches the failures they care about.

Monitoring Has To Meet Action

Even a well-designed pre-release evaluation cannot guarantee that a model will keep performing once users and conditions change. Monitoring needs to show what is happening in production, identify a meaningful regression and connect it to a decision: keep the model, adjust it or roll back. Khandelwal’s emphasis on observability and Rao’s account of a slow provider migration show that evaluation and operational readiness are linked.

For companies, the loop should run from real user outputs to failure categories, from those failures to targeted tests, and from test results to a release or rollback decision. Cost belongs in that loop too. Shadowfax compares model iterations on cost as well as accuracy and speed; Swiggy’s model portability work is intended to make changing providers less disruptive.

For companies, the loop should run from real user outputs to failure categories, from those failures to targeted tests, and from test results to a release or rollback decision.

Eternal’s Gavel post adds a practical illustration of scale: it says automated review of a voicebot call costs around ₹10, compared with roughly ₹450 for human review. That is a company-reported example, not evidence that human oversight can be removed. The stronger takeaway is that automation can widen coverage while human reviewers define and validate the quality bar.

As Indian companies move AI systems into customer-facing and operational workflows, evals will matter most when they change what teams do: block a weak release, surface a failure early, or show that a cheaper model is good enough. The advantage will go to teams that can make model changes quickly without asking users to absorb the risk.


Top Stories From India & Around The World

  • The Rise Of Bank-Built AI: Indian banks are increasingly co-building and investing in AI startups instead of simply buying off-the-shelf tools. Inc42 notes that HDFC Bank, Canara Bank, IDFC FIRST Bank and Bajaj Finance are among institutions taking a more active role in shaping financial AI.
  • Google Unveils Gemini 4 Argon: Google has launched Gemini 4 Argon for complex software engineering, enterprise knowledge work and cybersecurity, with a 1 Mn-token context window. The model is initially being made available to trusted cyber defenders before a wider rollout. 
  • Indian Startups Put AI Computing In Orbit: TakeMe2Space, SatLeo Labs and EON Space Labs have sent AI computing and Earth-observation payloads into orbit aboard SpaceX’s Falcon 9. TakeMe2Space’s satellite is designed to process imagery in orbit, reducing the need to send large volumes of raw data back to Earth. 
  • OpenAI Disrupts Model-Distillation Campaign: OpenAI says it disrupted a coordinated effort involving more than 15,000 users to extract protected reasoning from its models and use it for adversarial distillation. The company has since strengthened monitoring, account controls and protections around hidden reasoning.
  • NVIDIA GPT-6 Astra Ultrafast Touts Scale: NVIDIA says OpenAI’s GPT-6 Astra Ultrafast, running on Blackwell GPUs, can deliver up to eight times faster token generation than Astra Standard. The acceleration is aimed particularly at agentic coding, tool use and other latency-sensitive workloads.

The Weekly Buzz: Anthropic Expands Into Healthcare Ahead Of Massive IPO

Anthropic pushed Claude deeper into regulated industries this week while market chatter intensified around a potential pre-Thanksgiving IPO that some reports peg near a $2 trillion valuation. The company made Claude generally available for government use, expanded capabilities into healthcare billing and prior authorization, and enabled U.S. users to connect Apple Health and Android Health Connect records.

At the same time, Anthropic released Claude Sonnet 5.5—priced at the same level as its predecessor yet delivering faster output and lower cost per task—and continued building out Claude Code adoption with major banks and consulting firms. Separate reporting revealed Broadcom is preparing to lend Anthropic up to $42 billion in chip-lease financing, underscoring the scale of the company’s infrastructure commitments.

The dual track of product expansion into sensitive domains and preparation for a public listing comes as Anthropic also faces the same industry-wide questions about agent safety and self-improving systems that have hit other labs. Company executives have publicly called for stronger evaluation standards and oversight of recursive self-improvement.If the IPO materializes on the rumored timeline, Anthropic would become one of the most highly valued pure-play AI companies to go public, giving it fresh capital to fund both frontier model development and the specialized vertical deployments it is now prioritising.


Startup In The Spotlight: ByteAsk

Founded in 2026 by IIT-Delhi alumni, ByteAsk is building AI coding agents specifically for industries where C/C++ power mission-critical systems, including defence, automobiles, aerospace, medical technology and finance.

As AI coding tools become more capable, the bigger challenge for these industries is not simply generating code. C/C++ systems often sit inside environments where reliability, performance and safety matter as much as speed, making conventional AI-assisted coding less suitable for sensitive development workflows.

ByteAsk is building its platform to work with both open- and closed-source AI models, while allowing organisations to deploy it entirely on-premise. This enables companies to keep sensitive source code and intellectual property within their own infrastructure while using AI to assist with C/C++ development.

The startup is currently developing two offerings, ByteAsk Harness and a B2B enterprise product. It has raised $1 Mn from Y Combinator and follows a business model combining subscriptions with enterprise services.

ByteAsk operates in the AI coding market, estimated at ~11 Bn in annualised value as of April 2026. Its differentiation lies in focusing on C/C++, particularly for performance-critical and safety-sensitive systems where reliability can matter as much as code generation speed.


Prompt Of The Week

What prompts and hacks are CTOs, CEOs and cofounders using these days to streamline their work? 

Here’s Yash Varyani, Chief Technology Officer (CTO) and co-founder of Drizz, with a prompt he uses to review pull requests.

“You are a CTO reviewing a pull request. I’ll paste the diff. Skip line-by-line review; use five gates.

Gate 1: Problem

Right reviewers? Meaningful or housekeeping? How do other industries solve it? Product-agnostic version? Two-way or one-way door? Impact on retention, acquisition, expansion?

Gate 2: Production

Blast radius? Deployment and change risks? Monitoring needed before merge? Metrics affected, and how are they improved?

Gate 3: Architecture

Before vs after: better, or just more complex? Improved architecture or added surface area? Patterns used or missed (Port Adapter, DDD, CQRS)? Control plane separate from data plane? Two-way door? Better alternatives? Can Customer Success run this without engineering?

Gate 4: Code Health

Mental model, focus areas, best file review order? Cross-repo consistency? Flag coupling, cohesion, circular dependencies, SOLID violations, broken abstractions, type safety gaps, redundancy, blocking async code. Questions answered for engineering, PM, sales, CS? Old code cleaned up? Matches repo structure and UX conventions?

Gate 5: Ship Readiness

Works across Android, iOS, watch, car, device sizes? Client CI/CD impact? Multi-tenancy respected? Memory or resource risks? Unit tests enough, or integration needed? Customer reporting impact? More auditable? Central docs updated? More agentic approach? Can a machine learn from the output and improve the process?

End with the top risk, the biggest missed opportunity, and a verdict: SHIP, REVISE, or REDESIGN.“

Editor’s Note: Some prompts may need to be adjusted by users for best results or may not work as intended for certain users.

[Edited by: Nikhil Subramaniam]
[Creatives by: Varshita Srivastava]

The post The Case For Evals In A Multi-Model World appeared first on Inc42 Media.