AI Safety & Evaluation · 2026-08-06

Position: LLMs Can't Jump

ICML 2026

The inability to generate new premises is not only a model-capability gap; it is a governance boundary for institutions that delegate scientific agenda-setting to AI. World models may expand the space of machine-generated hypotheses, but without rules for evidentiary status, validation, attribution, and contestability they also concentrate authority over what counts as a plausible explanation.

AI Safety & Evaluation · 2026-08-06

Not All LLM Reasoning is Visible in the Chain-of-Thought

arXiv

Chain-of-thought monitoring is not an accountability mechanism when consequential computation can occur without an interpretable token trace. The governance implication is not simply that monitors need better detection, but that institutions must stop treating model-generated explanations as sufficient evidence of intent, compliance, or safe internal process.

AI Safety & Evaluation · 2026-07-12

‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI

Cambridge Programme on AI Science & Policy, University of Cambridge

The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.

AI Safety & Evaluation · 2026-05-15

From Symptoms to Systems: A Stakeholder-Informed Taxonomy of Generative AI Risks for Eating Disorders

Center for Democracy & Technology AI Governance Lab

The report's central contribution is that it treats eating disorder risk as a pattern of interaction rather than a prohibited content class. Its governance gap is that the taxonomy still needs to become an auditable control framework with thresholds, evidence requirements, escalation duties, and redress pathways.

AI Safety & Evaluation · 2026-03-24

Indigenizing Adversarial Stress Testing

Humane Intelligence and Emergence Circle

The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.

AI Safety & Evaluation · 2026-03-09

Agents of Chaos

arXiv

The paper shows that once language models are wrapped in memory, tools, messaging, and delegated authority, the main governance problem is no longer just model error but insecure delegation across socio-technical systems.