Evaluating Agents
For agents, the trajectory is the test
When the precedent hasn’t been set yet, we get to write it
For agents, the trajectory is the test
The governance implications are uncomfortable
Do you understand what you produced, and can you defend it when it's wrong?
Accountability is more than I think it is
The open question is who holds the record once the agent acts.
The scientists and academics here are starting to work on judgment.
The duty stays with a human or a company, not the AI.
As agents start doing real legal work, we have no clean way to prove what they actually did