Ken Priore The writing · kenpriore.com
Reflections · 2026-09-02 · 3 min read

Messy Data

Data is the new oil is the "new" claim for Big Law

Messy Data

Debevoise's Data Blog has a metaphor for a firm's accumulated work product: crude oil. Real value, sitting in the ground, useless in its current form. It becomes jet fuel only after refining.

Their refinery has five stages. Collection decides which work product is worth preserving at all. Separation pulls durable legal knowledge apart from duplicative drafts, client-specific facts, privileged material that cannot be reused, and advice that has gone stale. Conversion turns bespoke output into standardized, non-overlapping content organized so that both lawyers and models can navigate it. Quality control settles who approves what enters the platform, how changes get tracked, how often material is reviewed, and what happens when the law moves. Distribution puts the result where someone does the work. STAAR, their client-facing AI risk platform, is the proof of concept.

The claim I would double back on: as frontier models commoditize, the refinery becomes more valuable rather than less. That is the data-moat argument applied to professional services, and it is directionally right. If everyone has access to roughly equivalent reasoning, the differentiator is what you feed it and how tightly you control the feed. What the argument skips is that each release does more of the refining on its own. Longer context and better retrieval let models work messier corpora every year, so the conversion stage, all that careful standardizing and deduplicating and organizing, is partly a patch for model weakness. Patches for model weakness have a short shelf life.

The separation stage is doing governance work under a data-engineering label. Deciding, before ingestion, what is privileged, what is client-confidential, what carries reuse restrictions, what is jurisdictionally scoped, and what has expired: those are the determinations a privacy or AI governance function makes at design time. No model release makes them for you. Comprehension is the model's problem; permission is yours. Do the work at the front and it is metadata. Do it after retrieval and it is an incident.

I have watched that sequencing decision go both ways. Building the global privacy program at Atlassian, the classification work that felt slow up front is what let us scale requirements to thousands of marketplace partners without re-litigating every case. Teams that skip it do not save time. They defer the work to the moment of highest pressure.

In-house legal has the same crude oil and different geology. There is no cross-client privilege problem, which removes the hardest separation case. But in-house teams hold something firms do not: the negotiation history, the escalation record, the decisions made under deadline with incomplete facts. That is the corpus that would make an internal AI useful, and it is the corpus nobody has curated, because it lives in email threads and comment histories rather than in executed documents. The final contract tells you what was agreed. It does not tell you what you would never agree to again.

The practical move is to start with separation rather than collection. The instinct is to gather everything and sort later, which produces a large pile with unknown provenance and no reliable way to establish which parts are safe to reuse. Pick a narrow, high-frequency domain instead: DPAs, security exhibits, standard AI clauses. Classify before you ingest. Assign owners. Set expiration triggers. Then measure whether anyone uses the output, because a refinery producing fuel nobody burns is just an expensive archive.

Debevoise ends in the right place: none of this happens organically. Every refinery is built on purpose, staffed on purpose, and maintained on purpose. The firms and legal functions that treat this as an IT project will produce a searchable archive. The ones that treat it as an accountability design problem will produce something a lawyer will rely on.

Refining Law Firms’ Crude Oil into Jet Fuel – Leveraging Legal Data for AI – Debevoise Data Blog
← PreviousShow Me the Receipts
On the record
Follow
Check your inbox to confirm.