When the better lawyer wins, the LLM isn't deciding cases
it's grading rhetoric
A recent research paper feed legal arguments to frontier LLMs and varied the quality of the advocate making each argument. Then they looked at whether the model agrees with a legal point of view based on the merits, or based on how well the argument was written.
A judge who is "unduly persuadable" — who decides cases based on the skill of the advocate rather than the strength of the case — is, in any common-law tradition, a failed judge. That is the standard. The paper asks whether the systems being proposed as first-instance decision-makers in administrative and judicial contexts meet it.
For product counsel and AI governance teams, this connects to something I have been writing about for a while: persuasive wrongness. LLMs do not separate the form of an argument from its substance the way a trained adjudicator is supposed to. They reward fluency. It is a problem when the output is a benefits eligibility decision, a tenant dispute, an asylum claim, or a school discipline appeal.
The interesting design question is whether you can build adjudicative systems that explicitly discount rhetorical sophistication and weight evidence and applicable rules instead. That is a hard architecture problem, not a prompt-engineering one.