Colonised by the corpus: Sovereign AI must be authored in Indian tongues

By: Brijesh Singh
Last Updated: August 2, 2026 03:34:11 IST

Bias is routinely magnified in Indian languages when measured against English inside the same model family, exposing a frailty at the heart of the multilingual AI dream; translation is not authorship.

Ask an artificial intelligence model to speak of an Indian village, and more often than not it will open the same little box of fossils; monsoon clouds, mud walls, a bullock cart caught in sepia, preserved as if rural life were an insect in amber. Ask it then to think through a Tamil Siddha medical text, a Manipuri martial lineage, or the fiscal logic of a gram panchayat, and it either falls into a pious silence or begins to hallucinate. This is not a bug, waiting for some neat patch. It is the lawful child of a global AI infrastructure fed on data which was never collected with most Indians in mind, and now it sits, quietly but heavily, at the centre of the question of whether India will grow its own intelligence or lease another’s. The wound is deeper than absence. Researchers looking at representational bias in large language models have found a winner-takesall tendency; when training data leans toward one picture of a group, the model does not merely inherit the lean, it sharpens it, like a knife on stone. The dominant image pushes out the smaller ones, and the skew multiplies with every training cycle. The unsettling part is this; diversification of training data, by itself, does not stop the drift. The bias lives not only in the corpus but in objective functions, in sampling mechanics, in the hidden plumbing of learning, so the familiar prescription touches the surface and misses the root. For India this distortion grows a linguistic edge. Bias is routinely magnified in Indian languages when measured against English inside the same model family, exposing a frailty at the heart of the multilingual AI dream; translation is not authorship. A model that has learnt the world first in English and then pours it into Hindi or Bengali carries the first distortions through the membrane of translation, often making them denser. To treat Indian languages as coloured surfaces on which English knowledge is painted, and not as epistemologies in themselves, is to guarantee that fairness will remain cosmetic, a layer of varnish on an alien grain. Below misrepresentation lies something colder, invisibility. Large portions of India’s population are missing from the datasets that train global models not because they are wrongly represented, but because no one captured them at all. These data voids are structural, born of absent digital infrastructure rather than merely an algorithm’s failure to be inclusive. A farmer in a Scheduled Area without a smartphone, a craftsperson whose livelihood never leaves the offline world, a tribal community whose oral traditions leave no written footprint; they are not misread by AI. They are unread. And unreadness is a harder abyss than error.

Even where Indian-language life exists online, much of it remains opaque to the machines that harvest the web. Indic activity on the internet is app-first, video-first, voice-first; it flowers inside WhatsApp groups, YouTube channels, voice notes, and other halfclosed courtyards, not on crawlable pages waiting politely for a scraper. The dominant data-gathering habit of global AI, designed around the open web, passes by this abundance like a blind traveller. To close the gap is not merely to scrape better; it requires deliberate, statebacked collection of data, the kind of public capacity that can endure beyond startup appetite. The economics of contemporary models add yet another sieve. Tokenizers built primarily around English place a 1.4 to 5 times token cost on Indic languages, so a Hindi sentence eats far more tokens than its English twin. The consequences are simple and not small; inference becomes costlier for Indian users and developers, and the effective context length of the model shrinks when it handles Indian-language text. A system which can hold a long English document in its working memory may lose the thread halfway through a Marathi one. This is not a decorative inefficiency at the margin; it decides which Indian-language applications can even be imagined. There is also a loop tightening, which few beyond research rooms have reckoned with. As AI-generated content floods the web, models increasingly train on the excretions of other models, a circular referencing of mirrors facing mirrors, risking the hardening and amplification of old distortions. For low-frequency Indic traditions, texts, dialects, knowledge systems that appear only rarely in any corpus, this is particularly dangerous. One hallucinated version may become the canonical version for future models, and absence, through repetition, becomes permanent distortion. India sits on a resource which could rebalance much of this game, and yet it lies mostly untouched. The country’s manuscript heritage, estimated at roughly ten million items, remains largely undigitized and inaccessible, scattered across temples, homes, regional archives, like seeds locked away before the monsoon. These are not merely relics for antiquarian tenderness; they are the foundationstones of sovereign AI, corpora authored in Indian languages, from Indian soil, bearing knowledge systems never encountered by the global pipeline. To let them decay is, in effect, to let the training data of the future rot quietly.

The narrowness is also in the hands that build the machines. The developers shaping global models are perhaps ten thousand in number, overwhelmingly male and largely Westerneducated, and such concentration colours everything; what becomes a benchmark, what is optimized, what is ignored as noise. This is not an accusation of bad faith. It is the plain recognition that epistemic perspective follows demography. When those defining general intelligence occupy a thin slice of human experience, the resulting models carry that thinness as their default, however dazzling the engineering may look. Benchmarks are beginning to show the hollow. Even the strongest multilingual models perform poorly on India-specific knowledge, as MILU and similar evaluations have shown, stumbling over questions an informed Indian schoolchild might answer. Without locally designed, dedicated benchmarks, these deficiencies remain invisible to the very teams that build the systems. One cannot improve what has not been measured in one’s own grammar of reality. The exit, therefore, is not finer translation bolted on to foreign foundations. Sovereign AI requires native-authored corpora, culture-first training objectives, and evaluation frameworks co-designed with the communities for whom the technology is supposedly built. English dominance in AI is not some technical destiny; it is also the continuation of colonial power structures older than the machine. To answer it needs governance interventions, data trusts, public investment, regulatory framing, not only more vernacular pages thrown into the furnace. Some scaffolding is now appearing. AI Kosh and Bharat Data Sagar are assembling a national data repository, with 367 datasets catalogued as of May 2025, meant to provide the sovereign data foundation on which indigenous AI development depends. The number is modest beside the scale of hunger, but it signals a shift in posture; from consuming global models to becoming custodian of national knowledge. What remains suspended is whether India’s intelligence future will be written at home or imported under licence. The manuscripts, the languages, the living traditions of a billion people form an epistemic inheritance no other country possesses in such abundance. An AI that actually sees India can be built, but only if India treats its own knowledge as infrastructure, not ornament, and builds institutions worthy of that inheritance.

* Brijesh Singh is a senior IPS officer and an author (@brijeshbsingh on X). His latest book on ancient India, “The Cloud Chariot” (Penguin) is out on stands. Views are personal.

Most Popular

The Sunday Guardian is India’s fastest
growing News channel and enjoy highest
viewership and highest time spent amongst
educated urban Indians.

The Sunday Guardian is India’s fastest growing News channel and enjoy highest viewership and highest time spent amongst educated urban Indians.

© Copyright ITV Network Ltd 2025. All right reserved.