AI Safety & Evaluation · 2026-07-12

‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI

Cambridge Programme on AI Science & Policy, University of CambridgeOriginal paperMarkdown source
generative AIAI safetyAI risk managementadversarial evaluationdeployment controlsrisk assessmentcyber riskgeopolitical riskinstitutional readinesstransparency and accountability
Key Insight

The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.

Review

“God has helped us, and so will AI” reports 57 semi-structured interviews with 27 former members of the two principal Boko Haram factions, ISWAP and JAS, conducted in northeast Nigeria during 2025 and 2026. Fifteen participants claimed knowledge of organizational AI use. Their accounts describe access to several leading chatbots, training delivered through transnational militant networks, dedicated internal AI units, rank-based access controls, paid accounts, and routine use of model outputs for technical problem-solving, logistics, operational security, planning, and organizational learning. The report’s central claim is not merely that terrorists have queried chatbots. It is that an armed organization has begun converting general-purpose AI into institutional capability.

That distinction materially changes the governance problem. Most deployed safeguards treat misuse as an interaction between a provider and an individual account. The organization described here operates across providers, accounts, intermediaries, languages, devices, and command structures. Refusal by one model becomes routing friction rather than denial. Expertise is pooled, prompts are refined with external assistance, outputs are interpreted by specialists, and guidance is passed to personnel who may never touch the model. The effective user is therefore a networked institution. Safety controls focused on prompt classification and account suspension govern the visible endpoint while leaving the capability-production system largely intact.

The fieldwork provides rare evidence about the concealed organizational layer of AI adoption. Earlier analysis often relied on propaganda and online supporter activity, which privileges what militant groups choose to publish. This study instead asks how access, expertise, and decision rights are arranged internally. The reported creation of specialist units is especially consequential because it converts model access from individual knowledge into a repeatable organizational service. It also concentrates interpretive authority. Commanders decide which problems are submitted, specialists determine how questions are framed, and leadership controls which outputs become operational instructions. AI is not autonomously directing violence, but it may change who within the organization can produce credible advice and how quickly that advice can circulate.

The research design is careful about access, translation, respondent protection, and the known weaknesses of testimony from former armed-group members. Multiple entry points, repeat interviews, role-based questioning, cross-respondent triangulation, and the inclusion of participants who denied knowledge of AI use all improve credibility. The report also distinguishes institutional adoption from proven capability uplift and explicitly states that it cannot establish whether AI enabled attacks that were otherwise impossible.

The evidentiary ceiling remains lower than the report’s most expansive language sometimes suggests. The core claims depend on retrospective self-report without chat logs, account records, device forensics, model outputs, subscription evidence, or independent observation of the reported units. Logo recognition establishes familiarity more readily than sustained use, and several participants with knowledge of organizational practices had not directly operated the systems. Some detailed operational stories may be accurate, partially reconstructed, exaggerated, or retrospectively attributed to AI after an outcome occurred. Triangulation across respondents reduces but does not eliminate correlated narratives, especially where knowledge travels through the same command network.

The report also combines four analytically distinct propositions: access, use, institutionalization, and uplift. It provides substantial qualitative evidence for the first three, but uplift remains largely perceived. Claims that AI improved precision, reduced casualties, accelerated innovation, or enabled tactical adaptation lack a counterfactual baseline. The relevant comparison is not AI versus no information. It is AI versus search engines, manuals, video, human trainers, captured expertise, and existing transnational support. Without task-level reconstruction and outcome comparison, the study cannot determine whether models supplied novel capability, compressed search and synthesis costs, increased confidence in existing plans, or simply became the most visible component of a broader learning system.

The safeguard findings are similarly time-bound. Most reported use concerns 2023 and 2024, with limited visibility into mid-2025. Models, moderation systems, account enforcement, and access policies have changed since then. The report can reasonably conclude that organized users perceived restrictions as manageable during the observed period. It cannot establish current bypass rates across the named providers or determine which outputs came from which model version. Provider-specific accountability therefore remains unresolved, and naming multiple companies risks implying equivalent failure without equivalent evidence.

The policy section identifies the need for cooperation among AI companies, governments, intelligence agencies, law enforcement, and researchers, but does not specify the governance of that cooperation. This omission matters. Information-sharing systems redistribute surveillance and enforcement power. They require defined legal authority, evidentiary thresholds, purpose limitation, retention rules, auditability, contestability, and redress for wrongly flagged users. In conflict-affected regions, expanded account monitoring and identity linkage can expose journalists, researchers, aid workers, dissidents, minority communities, and people seeking legitimate dual-use information. A safety response that treats secrecy as proof of maliciousness or relies on opaque watchlisting can create a second infrastructure of harm while claiming to mitigate the first.

The report’s most operationally important implication is that evaluation must move from harmful-answer testing toward adversarial workflow testing. Models should be assessed against coordinated users who distribute tasks, rephrase requests, combine benign outputs, compare systems, translate across languages, and insert human expertise between model output and action. Evaluation should separately measure information novelty, time saved, error correction, confidence effects, and downstream task performance. Providers also need cross-account and cross-session risk controls, but these controls should be independently audited and bounded by due-process requirements rather than treated as a discretionary extension of private intelligence power.

Governments should avoid interpreting the findings as a mandate for broad restrictions on general-purpose access. The reported organization benefited not only from explicitly harmful answers but from ordinary functions such as troubleshooting, translation, planning, and synthesis. Restricting those capabilities for entire regions would externalize security costs onto legitimate users while sophisticated actors route around them. A more defensible approach combines targeted disruption of known networks, controlled access for exceptionally high-risk capabilities, stronger provenance and audit mechanisms, multilingual safety evaluation, and institutional channels through which researchers can disclose evidence without converting academic fieldwork into an intelligence collection proxy.

The report changes the baseline for AI-terrorism analysis. The immediate risk is not necessarily a model inventing an unprecedented weapon. It is AI becoming an always-available advisory layer inside an organization already capable of violence, experimentation, recruitment, procurement, and transnational learning. That layer can lower coordination costs, preserve expertise, and make external knowledge easier to operationalize. The corresponding governance test is whether safety institutions can observe and constrain organizational misuse while remaining lawful, reviewable, and proportionate. The report establishes the need for that test, but leaves its institutional design largely open.

Key Insight

The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.

Appears in these collections

Continue exploring

Related reviews

More in AI Safety & Evaluation
AI Safety & Evaluation · 2026-05-15

From Symptoms to Systems: A Stakeholder-Informed Taxonomy of Generative AI Risks for Eating Disorders

Center for Democracy & Technology AI Governance Lab

The report's central contribution is that it treats eating disorder risk as a pattern of interaction rather than a prohibited content class. Its governance gap is that the taxonomy still needs to become an auditable control framework with thresholds, evidence requirements, escalation duties, and redress pathways.

AI Safety & Evaluation · 2026-03-09

Agents of Chaos

arXiv

The paper shows that once language models are wrapped in memory, tools, messaging, and delegated authority, the main governance problem is no longer just model error but insecure delegation across socio-technical systems.

AI Safety & Evaluation · 2026-03-24

Indigenizing Adversarial Stress Testing

Humane Intelligence and Emergence Circle

The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.