AI Safety & Evaluation · 2026-09-16
arXiv
When an evaluator can reward visible reasoning independently of verified task completion, the evaluation system becomes part of the behavior being optimized and cannot serve as neutral assurance evidence.
AI Safety & Evaluation · 2026-09-16
arXiv
If interaction topology and message history can create behavior absent in isolated agents, assurance cannot stop at the model boundary: the relationship, protocol and composition become part of the governed system.
AI Safety & Evaluation · 2026-09-03
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), PMLR 306
Intermediate-token visibility is not assurance: when traces are not causally tied to outcomes, governance must attach trust to verifiable decisions and external commitments rather than plausible-looking model monologues.
AI Safety & Evaluation · 2026-08-06
ICML 2026
The inability to generate new premises is not only a model-capability gap; it is a governance boundary for institutions that delegate scientific agenda-setting to AI. World models may expand the space of machine-generated hypotheses, but without rules for evidentiary status, validation, attribution, and contestability they also concentrate authority over what counts as a plausible explanation.
AI Safety & Evaluation · 2026-08-06
arXiv
Chain-of-thought monitoring is not an accountability mechanism when consequential computation can occur without an interpretable token trace. The governance implication is not simply that monitors need better detection, but that institutions must stop treating model-generated explanations as sufficient evidence of intent, compliance, or safe internal process.
AI Safety & Evaluation · 2026-07-12
Cambridge Programme on AI Science & Policy, University of Cambridge
The report shows that the relevant unit of AI misuse is not the isolated malicious prompt but the organization that can train specialists, distribute access, compare providers, and convert model output into operational routines. Safety governance built around single-user refusals will remain structurally inadequate unless it can address coordinated adversaries without turning platform monitoring into unaccountable security infrastructure.
AI Safety & Evaluation · 2026-05-15
Center for Democracy & Technology AI Governance Lab
The report's central contribution is that it treats eating disorder risk as a pattern of interaction rather than a prohibited content class. Its governance gap is that the taxonomy still needs to become an auditable control framework with thresholds, evidence requirements, escalation duties, and redress pathways.
AI Safety & Evaluation · 2026-04-06
arXiv
CUBE correctly identifies benchmark fragmentation as an infrastructure bottleneck, but the standard it proposes would also become a governance layer that shapes what agent capability is legible, portable, and worth optimizing for.
AI Safety & Evaluation · 2026-03-24
Humane Intelligence and Emergence Circle
The framework recasts AI stress testing as an exercise of Indigenous governing authority rather than a vendor-controlled safety check, but its sovereignty claims require enforceable agreements, auditable revocation, remedy pathways, and developer obligations before test findings can constrain deployment.
AI Safety & Evaluation · 2026-03-14
arXiv
MASFactory’s real contribution is not that it makes multi-agent systems easier to build, but that it reframes orchestration as a reusable governance surface where topology, context access, and human intervention can be made explicit, inspectable, and testable.
AI Safety & Evaluation · 2026-03-09
arXiv
The paper shows that once language models are wrapped in memory, tools, messaging, and delegated authority, the main governance problem is no longer just model error but insecure delegation across socio-technical systems.
AI Safety & Evaluation · 2026-03-07
arXiv
Automatically generated repository context files often degrade coding-agent performance because they introduce additional constraints without improving task-relevant understanding.
AI Safety & Evaluation · 2026-03-06
International AI Safety Report
The central AI-safety challenge is not diagnosis but translating uncertainty into default actions, decision thresholds, and enforceable consequences.