Proving Loyalty

When an agent is not your agent

3 min read
Proving Loyalty

Fiduciary law has an unusual feature that most evidence problems do not. When a fiduciary is accused of disloyalty or carelessness, the burden often runs the other way. The fiduciary must prove it acted properly, rather than the challenger proving it did not. That allocation exists because the fiduciary controls the information and acts in another's interest, so the law makes it answer for its own conduct. This project asks what happens to that bargain when the fiduciary is an AI agent, and the only contemporaneous account of what it did and why is the record the agent generated about itself.

The question is narrow and, I think, new: can a self-generated record discharge a fiduciary's burden of proof, and what would make it reliable enough to do so? This is not ordinary admissibility. The burden is reversed and heightened, which means the record is the centerpiece of the fiduciary's defense. A record that fails is worse than a weak exhibit: it leaves a fiduciary with no way to show it was loyal.

To carry that burden, the record has to answer a connected chain of questions, and a gap at any link breaks the proof. It begins with what the agent did, and the authority it claimed to act under. From there, the questions are whether anyone observed the action as it happened or whether it ran unwitnessed, and whether the action can be traced through the system to the inputs, decisions, and effects that produced it rather than surfacing as an unexplained output. The chain ends at durability: whether the record is indelible, sealed so it cannot be quietly edited after the fact to fit the account the fiduciary now needs to give.

I propose evaluating that record against a few testable factors: whether the action was observable and traceable in the first place, whether the record is tamper-proof, whether the party vouching for it is independent of the agent that acted, and whether the verification reproduces. Independence is the factor I expect to do the most work, because a fiduciary proving its own loyalty with a record it generated is exactly the conflict the burden-shifting rule was built to police. A self-record may need an independent verifier in the loop before it can carry the burden at all.

The deeper problem is a paradox of its own. The proof is generated by the party that must persuade, and that party owes a duty of candor. A fiduciary cannot satisfy a duty of disclosure with a record that performs thoroughness while hiding a defect. The reliability question and the candor question turn out to be the same question asked twice: a record that looks complete but obscures a skipped check fails as evidence and breaches the duty it claims to document in the same stroke.

The empirical core tests whether evaluators treat these records as burden-carrying in practice, and whether they are fooled by form. Using simulated agent episodes and the records they produce, I plan to vary the substantive defensibility of the underlying action while holding the record's polish and tamper-proofing constant, then measure whether legal evaluators separate a genuinely loyal action from a well-documented disloyal one. The overtrust result, if it appears, would matter beyond this setting: it would show that the very mechanism meant to enable accountability can instead launder a breach.

The chain above reflects how agents are built today, and the technology will change. The fiduciary claim underneath does not. Whatever tooling produces the record in five years, a fiduciary will still have to show what its agent did, that the action was observed and traced, and that the record of it cannot be altered after the fact. Those are first principles of the fiduciary role. They hold whatever the architecture, and they are what keep the inquiry durable.

The questions I am still working through: whether the burden-shifting framing holds across the fiduciary contexts I am drawing on or whether I am overgeneralizing from a few; whether verifier independence should be a precondition for admissibility or one factor among several; how to design the overtrust experiment so it measures evaluator calibration rather than familiarity with a polished-looking record; and who, in agency-law terms, the record's statements should be attributed to. Where I want to land is a defensible account of when a fiduciary AI agent can prove its own loyalty, and when the law should refuse to let it try alone. If you work in fiduciary law or agent systems and see a hole in it, I want to hear it.