AI Governance · 2026-05-09

Governing Artificial Intelligence in India: Data Sourcing, Synthetic Content, and Technological Sovereignty

Kautilya School of Public Policy Working Paper #3Original paperMarkdown source
AI governanceAI regulationdigital sovereigntyIndiaprovenancetransparency and accountabilitystate capacityinfrastructure governancerights-based frameworksmodel governancesovereignty
Key Insight

The paper connects data sourcing, synthetic content, and technological sovereignty as one governance loop rather than three separate policy problems. Its central gap is that it proposes institutional remedies without fully specifying the enforcement architecture, evidence duties, revocation mechanics, and redress pathways needed to make those remedies operational.

Review

*Governing Artificial Intelligence in India: Data Sourcing, Synthetic Content, and Technological Sovereignty* is valuable because it starts from a systems premise: India’s AI governance problem is not only about model behavior, but about the political economy of inputs, outputs, and infrastructure. The paper links three concerns that are often treated separately: training-data extraction, synthetic content, and dependence on non-indigenous technology. That framing is directionally correct. Data sourcing determines whose knowledge becomes machine-legible. Synthetic content determines whether provenance and authenticity survive downstream reuse. Infrastructure dependence determines whether the state has enough jurisdictional and operational leverage to govern either layer.

The central contribution is the paper’s insistence that AI governance should not be anchored primarily in definitional battles over AI. Its use of an “AI+Systems” frame is more useful for public policy because institutional harm emerges when AI is embedded into creative, educational, commercial, and administrative systems. This shifts the analysis from abstract capability to situated power. Artists, students, teachers, cultural communities, consumers, developers, cloud providers, data marketplaces, regulators, and ministries do not occupy symmetrical positions. Some actors operate the system, some use it, and some are absorbed into it without meaningful knowledge or consent. That asymmetry is the real governance object.

The paper is especially strong on the extraction problem. It correctly treats unlicensed scraping as a failure of authority, not merely a copyright inconvenience. When creative works, personal images, and community-held cultural expressions are converted into training inputs without consent, the system reallocates value and expressive control away from creators and communities toward model developers. The recommendations on copyright amendments, TDM rules, mandatory licensing, AI societies, collective remuneration, machine-readable noAI signals, dataset disclosures, provenance documentation, watermarking, and a MeitY scraping registry all point toward a more enforceable control plane for training data.

The synthetic content section is less developed but institutionally important. The paper recognizes that deepfakes and imitative outputs degrade public trust, academic integrity, cultural authority, and consumer understanding. Its proposed labelling approach for AI-generated content in indigenous artistic styles is useful because it distinguishes provenance from legitimacy. A label should not merely say that an output is AI-generated. It should prevent the output from being mistaken for culturally authoritative work. That distinction matters in India, where many cultural forms are not just aesthetic categories but community-governed expressions with moral, economic, and representational stakes.

The sovereignty argument is also right in principle. Dependence on external compute, cloud infrastructure, model supply chains, and pricing regimes creates a regulatory weakness that cannot be solved by domestic AI ethics guidelines alone. If the critical infrastructure of AI remains outside effective domestic leverage, Indian policy may be reduced to negotiating after the architecture has already allocated control. The paper’s reliance on the IndiaAI Mission, public-private compute infrastructure, and collective bargaining with large AI firms is pragmatic, but it also reveals a deeper tension: sovereignty cannot be purchased only as subsidized access. It has to be built as enforceable optionality across compute, datasets, standards, audit capacity, procurement, and public-interest deployment.

The main weakness is operational specificity. The paper proposes tribunals, transparency reports, DPIAs, scraping registries, licensing bodies, public awareness programs, metadata mandates, and an independent board for bargaining with technology firms. These are plausible instruments, but the paper does not fully specify the institutional mechanics that would make them enforceable. Who validates dataset disclosures? What is the penalty for false or incomplete registration? How are opt-outs propagated into already trained models? What happens when a community revokes permission? What evidence must a developer produce during a dispute? How will tribunals access technical logs, training-data records, and model update histories? Without these details, governance remains mostly declaratory.

The methodology is suitable for an exploratory working paper, but not strong enough to support some of the stronger policy claims. The paper uses design thinking, logical framework analysis, and theory of change, and it openly acknowledges that its analysis is abstract rather than empirically driven. That honesty is welcome. The limitation is that the paper sometimes presents institutional recommendations without testing administrative capacity, market incentives, regulatory costs, litigation load, startup impact, or compliance feasibility. A scraping registry, for example, is a strong idea only if paired with audit standards, disclosure schemas, confidential filing rules, public summaries, complaint triggers, and enforcement thresholds.

The novelty of the paper lies in its integrated Indian framing. It does not treat data governance, synthetic media, and technological sovereignty as isolated silos. It sees them as a feedback loop: opaque data sourcing weakens model accountability; weak accountability enables synthetic content harms; foreign infrastructure dependence limits domestic monitoring and enforcement. That is the right strategic frame for AI governance in India. The paper should now be converted from a policy map into an institutional architecture. Each proposed mechanism needs a mandate, accountable authority, evidence requirement, trigger condition, review cycle, appeal route, and measurable success indicator.

Read as a governance artifact, the paper is best understood as a strong agenda-setting intervention rather than a complete regulatory design. Its core insight is that India’s AI governance challenge is not simply to regulate model outputs. It is to govern the full chain through which cultural material, personal data, compute access, synthetic content, and institutional authority are transformed into infrastructure. The next step is to make the proposed controls testable, auditable, and contestable.

Key Insight

The paper connects data sourcing, synthetic content, and technological sovereignty as one governance loop rather than three separate policy problems. Its central gap is that it proposes institutional remedies without fully specifying the enforcement architecture, evidence duties, revocation mechanics, and redress pathways needed to make those remedies operational.

Appears in these collections

Continue exploring

Related reviews

More in AI Governance
AI Governance · 2026-03-14

Advancing Indigenous Foundation Models

White paper

The paper’s decisive analytical move is treating indigenous foundation models as public-interest infrastructure, but it stops short of specifying the assurance, procurement, and lifecycle governance machinery needed to make that ambition operational.

Read review · uploaded white paper: Advancing Indigenous Foundation Models
AI Governance · 2026-05-11

AI Governance at the Frontier: Unpacking Foundational Assumptions

Center for Security and Emerging Technology

The report's decisive analytical move is to treat governance proposals as bundles of assumptions rather than competing slogans. Its unresolved weakness is that assumption-mapping becomes policy-relevant only when each assumption is translated into testable institutional capacity, enforceable authority, and observable failure conditions.

AI Governance · 2026-04-16

AI Index Report 2026

Stanford Institute for Human-Centered Artificial Intelligence (HAI)

The report’s most important contribution is showing that AI capability, compute, capital, and measurement power are concentrating faster than governance systems can adapt, leaving a small set of actors with growing influence over both AI’s trajectory and the terms on which it is evaluated.

AI Governance · 2026-04-14

AI Governance, Safety and Infrastructure

Global Network Initiative and Centre for Communication Governance, National Law University Delhi

The briefing’s central contribution is showing that standards, safety institutions, and infrastructure concentration are converging into one governance problem, but it stops short of specifying the enforceable control points that would actually redistribute power.