AI Governance · 2026-09-16

Designing Loyalty: AI Agents and Conflicts of Interest

Stanford Institute for Human-Centered Artificial Intelligence (HAI)Original paperMarkdown source
AI agentsdelegationaccountabilityauthoritygovernance-by-design
Key Insight

A duty of loyalty becomes governable only when delegated authority, conflicts of interest, execution boundaries, revocation, and evidence of action can be made observable and enforceable at the point an agent acts.

Review

Smith, Wu, and King identify a structural problem that becomes unavoidable once AI systems move from recommending actions to executing them: the entity presented to a user as an agent may operate inside commercial incentives controlled by developers and deployers whose interests diverge from the user's. The brief proposes treating developers and deployers as fiduciaries in high-stakes consumer settings, with a duty of loyalty bounded by the delegated task. It translates that duty into disclosure of material commercial relationships, restrictions on self-preferencing, purpose limitation for data acquired through delegation, and affirmative consent before specified high-stakes actions.

The governance contribution goes beyond fiduciary language. The authors connect legal obligation to infrastructure: task-scoped and revocable agent credentials, verification of user identity, controlling entity and authorization scope, data minimization, incident reporting, and regulatory registration. This matters because interoperability protocols can otherwise become de facto governance by defining what systems permit while leaving authorization, consent and commercial-use constraints to implementers.

The unresolved issue is execution. A legal duty can establish who owes an obligation and create liability after breach, but it does not by itself determine whether a particular action is admissible before execution. "Best interests" also remains context-dependent, especially where price, convenience, privacy and user preference conflict. The proposal therefore depends on an institutional bridge from fiduciary obligation to machine-verifiable delegation, conflict disclosure, policy enforcement, evidence retention and revocation. The brief recognizes much of this infrastructure but does not specify a complete enforcement architecture or redress lifecycle.

Its durable implication is that agent governance cannot stop at identity or provenance. Knowing which agent acted and which company controls it supports attribution, not legitimacy. A governed agent relationship must also establish whose authority the agent exercises, the permitted task and duration, which conflicting incentives are disallowed, how authority can be withdrawn, and what evidence allows a user or regulator to challenge what happened.

Key Insight

A duty of loyalty becomes governable only when delegated authority, conflicts of interest, execution boundaries, revocation, and evidence of action can be made observable and enforceable at the point an agent acts.

Appears in these collections

Continue exploring

Related reviews

More in AI Governance
AI Governance · 2026-09-08

AI Agents Push Humans Out of the Loop

arXiv

Human oversight is not a governance control merely because a person remains in the loop; it is effective only while the system preserves the attention, expertise, independence, and decision capacity required to exercise authority over it.

AI Governance · 2026-08-03

Critique of Agent Model

arXiv

The paper correctly identifies that advanced agents redistribute control by internalising goals, identity, deliberation, and learning, but it mistakes architectural visibility for governability: an inspectable module is not an accountable institution unless authority, constraint, revocation, evidence, and redress are executable around it.