Inclusion, Rights & Development · 2026-04-28

AI’s English Problem—and Why We Should Care

AI governanceinfrastructure governancedigital public infrastructureglobal southIndiaLLMspublic infrastructurelegitimacy
Key Insight

The article’s central contribution is to frame language as AI infrastructure rather than interface localization, but its governance model remains insufficiently operationalized because it does not define who controls linguistic datasets, who can authorize reuse, and how communities can contest downstream model behavior.

Review

This TechPolicy.Press perspective by Sushant Kumar and Ananya Mukherjee argues that English dominance in AI is not a marginal usability problem but a structural barrier to digital inclusion. Its central move is to shift language from the periphery of product design into the core infrastructure of AI access. The article uses India’s Bhashini initiative as the primary example, alongside African multilingual AI efforts such as Lelapa AI, Masakhane, and African Language Lab, to show how language datasets, model development, speech systems, and translation pipelines are becoming prerequisites for meaningful participation in AI-mediated services.

The governance significance of the article lies in this infrastructure framing. If AI systems are primarily trained, evaluated, and deployed through high-resource languages, then English becomes more than a language of convenience. It becomes a gatekeeping layer for public service access, administrative legibility, economic participation, and institutional voice. In that sense, the article is not merely describing a model performance gap. It is identifying a control-plane problem: the communities whose languages are poorly represented in training data are also the communities least able to shape how AI systems classify, translate, advise, and mediate their interactions with the state and the market.

The article directly advances the analysis when it connects language exclusion to power concentration. It correctly observes that most AI systems are built within institutional and commercial environments far removed from low-resource language communities. The resulting asymmetry is not just that models perform poorly in Odia, Gujarati, Assamese, Hausa, or other underrepresented languages. The deeper problem is that linguistic communities may become dependent on systems whose categories, context assumptions, and representational defaults were not designed with their authority or consent. The example of models defaulting to four seasons when local ecological reality recognizes two is not a trivial cultural error. It illustrates how model outputs can normalize the worldview embedded in dominant training corpora.

Bhashini’s community-driven data collection model is presented as an alternative governance pathway. The article treats BhashaDaan and similar initiatives as mechanisms through which citizens can contribute speech, text, and local linguistic knowledge to AI development. This is valuable because it resists the purely extractive model in which communities appear only as raw data sources for platform expansion. It also recognizes that public AI infrastructure in multilingual societies cannot be produced solely through private-sector scaling or frontier model adaptation. There is a legitimate role for public institutions, academic partners, linguistic experts, and community contributors in building shared linguistic assets.

Yet the article does not fully resolve the governance contradiction inside its own proposal. Community contribution is not the same as community control. A language dataset can be crowdsourced, open, or locally curated while still lacking durable rules for consent, provenance, licensing, revocation, benefit sharing, misuse reporting, and downstream accountability. The article argues for rights-respecting dataset construction and structured partnerships instead of unpermitted scraping, but it does not specify what institutional machinery would make those rights enforceable after data enters model development pipelines. That omission matters because language datasets are not inert public goods. They are strategic inputs into commercial systems, state service delivery, identity formation, and cultural representation.

The methodology is illustrative rather than empirical. The article draws on prominent initiatives, public statistics, and examples from India and Africa to make a policy argument. This is appropriate for a perspective essay, but it limits the falsifiability of the claims. There is no comparative evaluation of model performance across languages, no governance audit of Bhashini’s dataset lifecycle, no evidence framework for measuring community authority, and no operational test for whether a multilingual AI system is culturally grounded rather than merely translated. The piece is therefore best read as a strong agenda-setting intervention, not as a complete governance model.

The largest gap is the absence of institutional design detail. Treating language as infrastructure should lead to hard questions: who is the steward of a language corpus, who has standing to challenge its use, how are dialect choices made, how are minority variants protected from standardization pressure, how are consent conditions carried into downstream models, and what happens when public-sector AI systems produce harmful or culturally invalid outputs in a low-resource language. These are not implementation details. They are the governance layer that determines whether multilingual AI becomes participatory infrastructure or a more inclusive front end for centralized extraction.

The article also underplays the risk that public multilingual AI infrastructure can consolidate state power as much as it expands access. In a country such as India, language infrastructure connected to welfare, health, agriculture, grievance redress, and public administration can improve reach, but it can also reshape dependency between citizens and the state. Voice-first and translation-mediated systems may make services more accessible while also making automated intermediation harder to contest for users with low literacy, low institutional power, or limited ability to escalate errors. Inclusion without redress can become a softer form of administrative control.

The paper’s novelty is not in claiming that low-resource languages are underserved. That claim is now widely understood. Its value lies in naming language as essential AI infrastructure and connecting multilingual AI capacity to public service delivery, Global South sovereignty, and institutional legitimacy. The next step is to make that claim operational. Multilingual AI systems need governance artifacts: dataset provenance records, consent and licensing terms, community review boards, model cards at the language and dialect level, error reporting channels, redress workflows, public procurement conditions, and audits that evaluate cultural validity alongside accuracy.

The article should therefore be used as a framing document for language infrastructure governance. It establishes why English-centric AI is structurally exclusionary, but it should not be treated as sufficient guidance for building legitimate multilingual AI systems. The hard governance work begins where the article stops: converting community participation into enforceable authority, converting linguistic inclusion into accountable service delivery, and converting language datasets into governed public infrastructure rather than unconstrained model fuel.

Key Insight

The article’s central contribution is to frame language as AI infrastructure rather than interface localization, but its governance model remains insufficiently operationalized because it does not define who controls linguistic datasets, who can authorize reuse, and how communities can contest downstream model behavior.

Continue exploring

Related reviews

More in Inclusion, Rights & Development
Inclusion, Rights & Development · 2026-06-24

How Can AI Support Language Digitization and Digital Inclusion?

Stanford Institute for Human-Centered Artificial Intelligence and Stanford SILICON

The paper correctly identifies language digitization as the precondition for AI inclusion, but its proposed social contract remains incomplete until communities have enforceable authority over data reuse, standards decisions, model deployment, and the right to withdraw or contest harmful representations.

Digital Public Infrastructure · 2026-05-04

DPI@2047 for Viksit Bharat: A Strategic Roadmap to Enable Non-linear Inclusive Socio-economic Growth

NITI Aayog / NITI Frontier Tech Hub

DPI@2047 treats digital public infrastructure as market-making state capacity, but it does not operationalize the governance layer that must decide who controls data flows, AI-mediated decisions, ecosystem access, revocation, redress, and accountability across decentralized implementation.

Economic & Market Infrastructure · 2026-04-16

Mapping India’s Data Centres: Aspirations, Realities and Futures

Digital Futures Lab

The report’s core contribution is to show that India’s data-centre buildout is not a neutral scaling exercise but a governance choice that reallocates water, energy, land, subsidy, and political priority toward compute infrastructure without yet building the disclosure, accountability, and redress mechanisms needed to legitimate that shift.

AI Governance · 2026-04-14

AI Governance, Safety and Infrastructure

Global Network Initiative and Centre for Communication Governance, National Law University Delhi

The briefing’s central contribution is showing that standards, safety institutions, and infrastructure concentration are converging into one governance problem, but it stops short of specifying the enforceable control points that would actually redistribute power.