AI Safety & Evaluation · 2026-07-12
Cambridge Programme on AI Science & Policy, University of Cambridge
The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.
AI Safety & Evaluation · 2026-05-15
Center for Democracy & Technology AI Governance Lab
The report's central contribution is that it treats eating disorder risk as a pattern of interaction rather than a prohibited content class. Its governance gap is that the taxonomy still needs to become an auditable control framework with thresholds, evidence requirements, escalation duties, and redress pathways.
AI Safety & Evaluation · 2026-04-06
arXiv
CUBE correctly identifies benchmark fragmentation as an infrastructure bottleneck, but the standard it proposes would also become a governance layer that shapes what agent capability is legible, portable, and worth optimizing for.
AI Safety & Evaluation · 2026-03-24
Humane Intelligence and Emergence Circle
The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.
AI Safety & Evaluation · 2026-03-14
arXiv
MASFactory’s real contribution is not that it makes multi-agent systems easier to build, but that it reframes orchestration as a reusable governance surface where topology, context access, and human intervention can be made explicit, inspectable, and testable.
AI Safety & Evaluation · 2026-03-09
arXiv
The paper shows that once language models are wrapped in memory, tools, messaging, and delegated authority, the main governance problem is no longer just model error but insecure delegation across socio-technical systems.
AI Safety & Evaluation · 2026-03-07
arXiv
Automatically generated repository context files often degrade coding-agent performance because they introduce additional constraints without improving task-relevant understanding.
AI Safety & Evaluation · 2026-03-06
International AI Safety Report
The central AI-safety challenge is not diagnosis but translating uncertainty into default actions, decision thresholds, and enforceable consequences.