AI Safety & Evaluation · 2026-03-24

Indigenizing Adversarial Stress Testing

Humane Intelligence and Emergence CircleOriginal paperMarkdown source
adversarial evaluationrights-based frameworkssovereigntyAI governanceevaluationsaccountability
Key Insight

The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.

Review

*Indigenizing Adversarial Stress Testing* relocates AI evaluation from the model provider's technical assurance perimeter into the jurisdiction of the Tribal Nation affected by the system. CARE, OCAP and UNDRIP are not treated as contextual ethics overlays. They define who may authorize testing, which data and cultural materials may enter it, who interprets harm, who controls disclosure, and who may stop the process. This is the framework's central governance contribution: evaluation becomes infrastructure for exercising sovereignty, not merely evidence supplied to an external deployer.

The report translates Collective Benefit, Authority to Control, Responsibility, Ethics, Ownership, Control, Access and Possession into adversarial scenarios involving misleading benefit claims, institutional ownership clauses, open-access defaults, policy-only possession, unauthorized translation and commercialization of traditional stories. Participants reportedly elicited violations across every tested principle, especially through multilingual prompts, role play, emotional pressure and multi-turn reframing. The findings expose a substantive asymmetry in model safeguards. Systems appeared more responsive to familiar Western legal categories, such as copyright and globally established human-rights language, than to Indigenous governance frameworks. That is not only a representation defect. It shows how training and safety policies establish an implicit hierarchy of authorities that models recognize and enforce.

The accompanying protocol converts the exercise into a governance sequence. A Tribal AI Oversight Committee sets scope and boundaries; free, prior and informed consent precedes testing; a Data Use and Testing Agreement governs ownership, storage, access and deletion; cultural review applies to prompts and datasets; consent may be withdrawn; and the Nation controls publication and follow-up. The protocol also recognizes evaluator wellbeing, capacity transfer and the need for community-controlled infrastructure. These elements move beyond participatory consultation because the Tribe retains decision rights throughout the lifecycle.

The evidence base remains exploratory. The report does not identify model versions, dates, system settings, participant numbers, sampling logic, complete transcripts, scoring rules or inter-rater procedures. Statements that models were successfully "exploited" therefore cannot yet support comparative claims, reproducibility or trend monitoring. Several scenarios also combine generation of harmful language with a subsequent self-critique, which tests conversational recoverability as much as the system's initial governance compliance. A future evaluation suite should separate refusal, warning, safe redirection, substantive recognition of collective authority, and remediation quality into distinct scored outcomes.

The protocol's institutional reach also stops at the testing boundary. It asserts immediate withdrawal, deletion, exclusive ownership and Tribal control, but does not specify how these duties are verified across model providers, cloud services, logs, backups, fine-tuning pipelines or downstream recipients. Nor does it establish mandatory developer response periods, independent escalation, compensation, deployment suspension, appeal, or proof that corrective action occurred. Without these mechanisms, a Tribe may govern the evaluation process while remaining unable to bind the organizations that own the model and deployment channel.

The next iteration should produce a machine-readable control catalogue and evidence schema linked to CARE and OCAP, versioned test cases, severity and confidence criteria, reproducibility metadata, signed decision records, deletion attestations, access logs and retesting requirements. It should distinguish local evaluation authority from provider remediation authority and define the contractual or regulatory bridge between them. The framework has already identified the correct locus of legitimacy. Its next task is to make that legitimacy operational against parties that do not voluntarily accept Tribal jurisdiction.

Key Insight

The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.

Appears in these collections

Continue exploring

Related reviews

More in AI Safety & Evaluation
AI Safety & Evaluation · 2026-05-15

From Symptoms to Systems: A Stakeholder-Informed Taxonomy of Generative AI Risks for Eating Disorders

Center for Democracy & Technology AI Governance Lab

The report's central contribution is that it treats eating disorder risk as a pattern of interaction rather than a prohibited content class. Its governance gap is that the taxonomy still needs to become an auditable control framework with thresholds, evidence requirements, escalation duties, and redress pathways.

AI Safety & Evaluation · 2026-03-09

Agents of Chaos

arXiv

The paper shows that once language models are wrapped in memory, tools, messaging, and delegated authority, the main governance problem is no longer just model error but insecure delegation across socio-technical systems.

AI Safety & Evaluation · 2026-07-12

‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI

Cambridge Programme on AI Science & Policy, University of Cambridge

The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.