Inclusion, Rights & Development · 2026-06-24

How Can AI Support Language Digitization and Digital Inclusion?

Stanford Institute for Human-Centered Artificial Intelligence and Stanford SILICONOriginal paperMarkdown source
AI governanceinfrastructure governanceinclusion, rights & developmentdigital public infrastructuresovereigntylegitimacyLLMsenvironmental impact
Key Insight

The paper correctly identifies language digitization as the precondition for AI inclusion, but its proposed social contract remains incomplete until communities have enforceable authority over data reuse, standards decisions, model deployment, and the right to withdraw or contest harmful representations.

Review

Stanford HAI and SILICON's white paper makes a necessary intervention in the AI inclusion debate: the decisive constraint is not merely that many languages lack enough data for large models. More basic infrastructure is absent. A language without stable script encoding, usable fonts, keyboard input, locale data, transcription capacity, and safety or accessibility tooling cannot reliably function across the digital environment. It will consequently struggle to generate data, gain representation in models, or support participation in digital services. The paper therefore reframes linguistic exclusion as an infrastructural dependency problem rather than a narrow natural language processing deficit.

This is the paper's central contribution. Its nine-tool account of digital inclusion distinguishes foundational tools from supporting ones and makes clear that connectivity is insufficient when people cannot read, write, search, learn, or obtain protection in their own language. The distinction between a digitally disadvantaged language and a low-resource language is equally important. Corpus size is not a proxy for participation. A language can have usable data but lack Unicode, rendering, or input support; it can also have limited data while retaining meaningful digital capacity through public investment and community stewardship. This prevents AI discourse from treating data volume as the sole measure of linguistic value or technical readiness.

The paper's staged mapping of AI assistance is practical. It identifies plausible roles for grapheme-to-phoneme systems, morphological analysis, OCR, speech recognition, language identification, translation, text-to-speech, forced alignment, and large language models across script development, transcription, and higher-order tools. Its caution is also sound: these systems can reduce particular bottlenecks, but they cannot decide whether a community should adopt a script, which orthography is legitimate, what should be documented, or whether linguistic material should be made reusable. Those are collective decisions with consequences for identity, authority, education, access to public services, and cultural continuity.

The white paper is at its best when it treats language digitization as a social contract. It recognizes that communities have often encountered extraction through missionary archives, colonial classification, corporate data collection, and contemporary model development. It also recognizes Indigenous data sovereignty and the need for community control. This matters because digitization converts linguistic practice into an asset that can be indexed, trained on, monetized, standardized, and governed at a distance. Once a script, corpus, tokenizer, or speech model becomes embedded in platforms, it can determine which dialect is legible, whose usage is treated as correct, which content becomes discoverable, and which users receive moderation, accessibility, or service support. Technical defaults redistribute decision rights.

Yet this is where the paper stops before the hardest institutional work. Community participation, co-design, benchmarks, convenings, and capacity building are necessary but not sufficient governance mechanisms. They describe an inclusive process without specifying durable authority. Who owns a corpus after it is used to train a commercial or open-weight model? Who can authorize a new downstream use? Can consent be limited by purpose, time, community membership, or geography? What provenance record follows a dataset and its derivatives across model training, fine-tuning, and deployment? Who can challenge a model's representation of a language, suspend a harmful integration, or secure correction when a public service, school, or platform privileges an imposed standard? The paper names data governance, but it does not provide a rights architecture for these decisions.

That omission is not a minor implementation detail. The proposed AI stack is largely built on standards bodies, cloud providers, frontier-model developers, app platforms, and public institutions that retain far greater technical and financial capacity than language communities. A community may be invited to contribute recordings, annotations, cultural knowledge, or evaluation labour while another actor controls storage, model weights, pricing, deployment terms, and the future value created from those contributions. In that arrangement, community-centered design can become consultation around an extractive pipeline. The paper should distinguish clearly between participation, stewardship, ownership, and governance authority. They are not interchangeable.

The discussion of standards requires a similarly sharper power analysis. Unicode encoding, CLDR contributions, keyboard designs, tokenization, and orthographic conventions are indispensable infrastructure, but none are neutral technical chores. They create canonical forms. A community may be internally diverse, use several scripts, or reject an externally preferred writing system. A standards decision can therefore settle political disputes by making one form interoperable and rendering others inconvenient, invisible, or expensive. The paper acknowledges that script development is political, but it offers no institutional protocol for disagreement: representation rules, deliberation, minority-dialect protection, appeal, revision, or periodic review. Without these, "community choice" risks becoming a claim made after intermediaries have already defined the available options.

The paper also does not operationalize its theory of successful inclusion. Its proposed global dashboard would be useful, but counting keyboards, OCR models, or language support can produce a misleading readiness score. A tool may exist but be inaccurate, unaffordable, unavailable on dominant devices, inaccessible to older users, governed by an untrusted provider, or rejected by speakers. Tool availability is not legitimacy, adoption, safety, or sustained capacity. The dashboard should therefore report provenance, maintenance ownership, funding horizon, interoperability, device and platform coverage, model error by dialect and context, community approval, data-use terms, environmental cost, complaints, and redress outcomes. It should also show where communities have decided not to digitize or not to make data reusable. Absence is not always a deficit.

The paper rightly notes that advanced AI can impose unequal token, cost, and emissions burdens on speakers of languages that are inefficiently represented in model architectures. This deserves more than a brief equity observation. Tokenization and API pricing convert linguistic design choices into recurring economic penalties. When a Bengali, Amharic, or Santali speaker pays more for comparable access, the system effectively taxes linguistic difference. Providers should disclose tokenization inequities, publish comparable cost and energy metrics, and remediate them through pricing, model design, or service obligations. Otherwise, AI inclusion will reproduce language hierarchy through the commercial infrastructure of access.

Methodologically, this is a policy synthesis and technical landscape scan, not an empirical evaluation of tools or governance outcomes. It draws together standards, practitioner experience, community initiatives, and published research to establish a coherent account of the digitization pipeline. That scope is appropriate for a white paper, but it limits the force of some recommendations. The report does not compare implementation models, measure adoption, evaluate long-term maintenance, test community governance arrangements, or establish that a particular AI tool improves capability without creating new dependency or harm. Its claims about acceleration should therefore be treated as plausible hypotheses conditioned by local capacity, funding, trust, and control, rather than universal effects.

The next version of this agenda should become an assurance framework rather than another inventory of promising tools. Every language digitization initiative should establish a community mandate, decision-rights map, data governance charter, provenance and licensing scheme, benefit-sharing terms, maintenance owner, platform-interoperability plan, independent evaluation protocol, complaint channel, correction and withdrawal process, and remedy for downstream harms. Funders should finance the continuing labour of stewardship, documentation, review, and repair, not simply model development or initial data collection. Technology providers should accept enforceable duties around attribution, purpose limitation, disclosure, audit access, incident response, and suspension of harmful uses.

The paper's most enduring insight is that language becomes a condition of citizenship in AI-mediated society. The languages that can be entered, rendered, searched, moderated, translated, and understood by systems will have a claim on digital public life that others do not. This is why language digitization cannot be governed as benevolent preservation or market expansion. It is an allocation of institutional visibility and agency. The paper establishes that diagnosis well. Its unresolved task is to build the control architecture that prevents inclusion from becoming a more culturally fluent form of extraction.

Key Insight

The paper correctly identifies language digitization as the precondition for AI inclusion, but its proposed social contract remains incomplete until communities have enforceable authority over data reuse, standards decisions, model deployment, and the right to withdraw or contest harmful representations.

Continue exploring

Related reviews

More in Inclusion, Rights & Development
Inclusion, Rights & Development · 2026-04-28

AI’s English Problem—and Why We Should Care

TechPolicy.Press

The article’s central contribution is to frame language as AI infrastructure rather than interface localization, but its governance model remains insufficiently operationalized because it does not define who controls linguistic datasets, who can authorize reuse, and how communities can contest downstream model behavior.

Digital Public Infrastructure · 2026-06-26

Digital Public Infrastructure in Africa: A Leapfrog Catalyst for Inclusive Growth and Prosperity

United Nations Development Programme, Regional Bureau for Africa and Digital, AI and Innovation Hub

The paper establishes DPI as state capacity, fiscal infrastructure and continental bargaining power rather than a technology stack. Its unresolved governance problem is that it calls for safeguards, sovereignty and inclusion without specifying the enforceable controls, failure metrics and redress rails that would make those claims operational at population scale.

Digital Public Infrastructure · 2026-05-04

DPI@2047 for Viksit Bharat: A Strategic Roadmap to Enable Non-linear Inclusive Socio-economic Growth

NITI Aayog / NITI Frontier Tech Hub

DPI@2047 treats digital public infrastructure as market-making state capacity, but it does not operationalize the governance layer that must decide who controls data flows, AI-mediated decisions, ecosystem access, revocation, redress, and accountability across decentralized implementation.

Economic & Market Infrastructure · 2026-06-24

Municipal Tokens as Urban Policy Tools: The Case of LVGA and the MyLugano App

P2P Financial Systems International Workshop

The paper establishes LVGA and MyLugano as municipal market infrastructure rather than a blockchain novelty. Its unresolved governance problem is that programmability increases the city's power to steer eligibility, spending, visibility, identity, and merchant participation, but the paper does not yet specify the enforceable controls, appeal rights, exclusion metrics, or fiscal accountability needed for legitimate urban infrastructure.