BharatGen CEO On Why Models Alone Can’t Help India Gain In The AI Race

With many homegrown foundation model initiatives competing for talent, compute and enterprise mindshare, the question of who will define the country’s AI future is becoming increasingly consequential.
Enter BharatGen. Launched as a non-profit consortium anchored at IIT Bombay and backed by the Department of Science and Technology, BharatGen is not your typical AI startup.
With ₹988.6 Cr in funding from the IndiaAI Mission — the largest single allocation under the government’s ₹1,500 Cr AI push — it is building multilingual and multimodal AI models across all 22 scheduled Indian languages. Its model families — Param, Patram, Sooktam and Shrutam — span text, vision and speech, and its open-source releases have already crossed 1 lakh downloads.
But unlike for-profit challengers such as Sarvam AI or Tech Mahindra’s Project Indus, BharatGen operates as a non-profit with a mission that extends beyond model releases.
In the last 18 months, BharatGen has published over 30 research papers and placed 100-plus interns across academia and industry. Today, the company is building what it calls the “Bharat Data Sagar”, a dataset repository aimed at capturing India’s linguistic and cultural diversity at scale.
In an interaction with Inc42, BharatGen CEO Rishi Bal laid out the organisation’s roadmap for the next 12 to 18 months. He spoke about the challenges of building AI for India’s underserved languages, and explained why the non-profit model is a feature, not a limitation, for India’s sovereign AI ambitions.
Edited excerpts…
Inc42: What does the roadmap for BharatGen models look like over the next 12 to 18 months?
Rishi Bal: Let me start with context. BharatGen is not a typical startup. We chose not to set it up as a for-profit company. We set it up as a non-profit. Our goal is to contribute to growing the whole AI ecosystem in India.
We also build models. But if you think of an iceberg, the models and applications are the visible part. The invisible part is significantly larger. If you don’t have a core group doing deep research across India, you’re not going to have the people who can create this next generation of AI. Our team has published over 30 papers in the past 18 months across top global journals. That’s an important piece because without a robust research ecosystem, you can’t be a global player as a country.
The model journey is very much on. Our 17 billion parameter model has gone through continuous progression in capabilities. The model we have today is significantly better than what we launched earlier. We have over 1 lakh downloads across our models. We have larger models in the pipeline. You should expect to see us launch new models every quarter — text, vision, and speech releases across our four model families: Param, Patram, Sooktam, and Shrutam.
On the application side, we are also building purpose-driven AI tools for India’s specific needs. Krishi Sathi, for instance, is a farm bot equipped with text-to-speech and data-driven insights to guide farmers. e-VikrAI is an AI assistant for Indian sellers to support and elevate their business operations. These are not just proofs of concept, they are live products reflecting the kind of real-world impact we are aiming for.
Inc42: How are you approaching data collection, especially for languages where digital text is scarce?
Rishi Bal: There is a massive data unlock that has to happen. When you go from English to even a European language like Spanish or French, there’s a big drop in data availability. From there to Hindi, there’s another order of magnitude drop. From Hindi to Marathi, Telugu, Tamil, Punjabi, Bengali, one more order of magnitude. The rest are even further behind.
So we have teams on the ground approaching publishers that own out-of-print books, old newspapers, and digitisation organisations working on heritage preservation. Just scanning documents doesn’t make them usable.
You have to convert them into native digital text, and that is hard and expensive. We have set up a world-class pipeline that converts scanned documents to digital text. We offer this pipeline to heritage preservation groups, we convert their material for free, and if they’re okay with it, we keep one copy to train our models. We don’t republish or sell it.
We also offer content-as-a-service and quid-pro-quo arrangements where enterprises get access to our models in return for data.
Inc42: What does the cost structure look like? How much of your spending goes into training versus data collection?
Rishi Bal: Data collection is labour-intensive, and luckily in India that is relatively cheap. It takes a lot of time and effort — we have people based in different parts of the country going out and meeting publishers and digitisation partners, but it is not capital intensive. The capital intensive part is training the models, and that is the bulk of the expenses for anybody building foundation models.
Roughly 80% goes into training and 20% into data collection is a good way of looking at it.
Inc42: What are the top two challenges BharatGen is facing in developing these models?
Rishi Bal: One is the talent ecosystem. LLM building is fundamentally a people business. It is quite challenging to hire, train, and retain the kind of talent needed for this work, especially as a non-profit.
The second is a more systemic challenge. Because we are not a typical AI startup with a product-launch-and-revenue trajectory, what we are trying to do is grow an ecosystem. And growing an ecosystem requires patience.
When we started, the media and many people were saying don’t even bother trying, India should not attempt this. We went from “it’s not possible” to proving it is possible. Now the impatience has shifted to “we can never be as good as Chinese models.” It’s a journey. We have to be persistent as a nation.
Inc42: How does the non-profit structure affect your revenue model and commercial strategy?
Rishi Bal: As a non-profit, I don’t have to return 10x capital to investors. But over time, I do want BharatGen to reach self-sustenance. Foundational model building is very expensive, so I think it will take about five years to get to that point.
The first 18 months were focused on model building. Once we proved these models are excellent, demand started emerging organically. Companies routinely approach us saying they’ve tried open-source and commercial models but are running into challenges. That is where we create custom models, custom solutions, and applications. We are working with a nationalised bank on a range of solutions, with a couple of state governments on paid commercial projects, not just proofs of concept, and with educational institutes. These are all committed, paid engagements.
Inc42: How do you structure the commercial engagement? Is it inference-based or something else?
Rishi Bal: It depends on the sensitivity of the data. In some cases, it involves customising models that we deploy into the client’s private infrastructure, which is more capital-intensive upfront. In other cases, we run a service or solution for them, which follows a different model, inference-based plus a fixed-cost component. I am seeing both models play out.
Data sensitivity is a big deal. It’s not just about having local data; it’s about local infrastructure and local AI components. That is emerging as a key factor in our conversations.
Inc42: What technical characteristics can we expect from the models launching in 2027?
Rishi Bal: First, you will see a range of model sizes. We don’t believe one size fits all. Different users need different cost-performance operating points.
Second, all our models will have heavier Indian focus, more Indian data, more Indian specifics. Think about a word like “gas.” Gas means petrol in America, but it means something else here. We need that local context baked in.
Third, heavy emphasis on inference cost optimisation. We are working with hardware partners and data centre partners to drive down costs. And while context windows will continue to grow, our approach is to focus on smarter context, how do you be smart about what chunks of the context window you’re looking at to drive down memory and compute?
In terms of scale, our partnership with the IndiaAI Mission commitments point toward models that are roughly one order of magnitude larger. We have already built on a corpus of 22 trillion tokens that is multilingual and multi-code, and we hope to come close to doubling that.
Inc42: How are you tackling the compute challenge, given global demand for GPUs?
Rishi Bal: Compute is in intense demand globally. Water and power requirements of data centres have led to clampdowns on investments in the US, Europe, and Singapore. We are lucky that India still sees data centre growth. But global demand is chasing supply globally, and hyperscalers are trying to lock Indian suppliers into multi-year contracts.
Our solution is to stay flexible and creative. We are not locked into one hardware infrastructure or one hosting provider. We work with three to four data centre companies, including local providers and hyperscalers. The strategy is to go where the supply is and be willing to adapt. Beyond renting GPUs, we are also building strategic compute partnerships. Earlier this year, BharatGen signed a landmark MoU with Larsen & Toubro to co-develop India’s sovereign AI compute platform covering AI chips, data centres, and foundational AI models, a long-term play to strengthen the infrastructure backbone for India’s AI ecosystem.
Inc42: Is BharatGen planning to build world models, similar to what some global labs and Tech Mahindra are attempting?
Rishi Bal: I think we need to build those models in India as well. We should not be left behind. The aspiration is to build frontier models in India. Why settle for something else? There will be plenty of people who will tell you it won’t happen, just like they said about LLMs two or three years ago. But if you don’t start building now, you won’t be able to do it when the need comes. The best time to dig a well was 10 years ago, but the next best time is today.
Inc42: What’s the approach to internal AI adoption at BharatGen? Do you use your own models exclusively?
Rishi Bal: Wherever we can use our models, we do. Our 17B parameter model handles a lot of our internal needs. But we don’t yet have a trillion-parameter model, so when a use case requires that scale, we use other options as well. We are looking forward to the day when we can largely run our organisation on our own models.
Edited by Nikhil Subramaniam
Creatives by Abhyam Gusai
The post BharatGen CEO On Why Models Alone Can’t Help India Gain In The AI Race appeared first on Inc42 Media.


Superadmin 










