Hi there, I’m Martin, Flank’s CTO.
Ask a GC why they won’t let AI near their live contracts and the answer is always some version of trust.
I think that objection rests on an assumption that’s about to stop being true. Last week I wrote about why the model was never the thing that mattered; this goes one level down, into what changes when legal judgement gets cheap.
A cheap scan is a different instrument
I spend a lot of time picturing the in-house legal team of five years from now. Almost none of that picture is about the models getting cleverer. The interesting part is almost always somewhere else. It is what becomes possible once something that used to be expensive becomes something you can do constantly, without thinking about the cost. The clearest version of that I have seen lately did not come from legal technology at all. It came from a company that makes pictures.
When Midjourney, until now a maker of images and video, said it was building a full-body medical scanner you step into like a spa pool, the obvious reaction was that someone had lost the plot. Bear with the detour, because it lands on legal in the end.
The objection most doctors raised was the standard, correct one about screening healthy people. Cancer is rare in any one person at any one moment, so run an imperfect scanner across a low-prevalence population and most of what you flag will be false. You frighten healthy people and send them for biopsies they never needed, and for whole-body screening of people with no symptoms the trials have generally failed to show they live any longer for it. A single imprecise snapshot, taken once, can do more harm than good.
That arithmetic is sound. It also quietly assumes the scan is expensive, so you take it once. Make the scan cost almost nothing and you stop taking it once. You take it every week, and a weekly scan is a different instrument. You stop asking “is there a tumour here”, which an imprecise tool answers badly, and start asking “is this changing, in this person, in a way it was not three months ago”, which the same tool answers far better. The screener never got more accurate. The economics changed what it was for, and the old objection quietly stopped applying. I am borrowing the logic here, not the clinical claim; serial screening has hazards of its own.
The same inversion is sitting in front of legal work, and almost nobody is pricing it in.

🔒 The objection we have all accepted
Ask a general counsel why they will not let an AI system look after their contractual obligations and the answer is some version of trust. The model gets things wrong. It is confident when it should not be. You cannot put it in the critical path of something that matters, because if it is wrong once, in the wrong place, you are exposed. I find this completely reasonable, and I also think it rests on an assumption that is about to stop being true.
The assumption is that a legal judgement is a single, expensive, one-shot act. A lawyer reads the contract, forms a view, and the view goes out of the door. Because forming the view is costly, in time and in salary, each one has to count. It is load-bearing the moment it is made. In that world trust has to be a gate. You either trust the judgement enough to act on it or you do not, and an imprecise judgement fails, because there is no second look to catch the error before it leaves the building.
That is the legal equivalent of the once-a-year scan, and it carries the same hidden dependency on cost.
The economics changed what it was for, and the old objection quietly stopped applying.
Reading the derivative
Collapse the cost of forming a legal view to almost nothing and the trust-as-gate stops making sense. Not because the contract changes: once you sign, the words are fixed, and you comply or you pay for not complying. What moves is everything around the words. The contract is a stationary line, and the business it governs never stops moving, so the thing worth watching is not the obligation but the widening gap between what you committed to and what the business is actually doing. A snapshot tells you whether you are onside today; the trajectory tells you that you are two quarters from breaching a covenant nobody has read since signing. The obvious objection is the right one: re-reading the same clause every morning buys you nothing, because a model that misreads an indemnity today misreads it the same way tomorrow. That error is correlated, not random, so it does not average out. The signal comes from triangulation. The obligation is read out of the contract once and mostly holds still, while what you track over time is the externally measured fact it constrains: the leverage ratio against its ceiling, the volume against a cap, the days left before a notice window closes. Those numbers live in systems that are not the contract, and they are real measurements rather than the model’s reading of itself. The model earns its keep on the hard part, extracting the obligation from messy contract language and knowing which fact it is tied to.
All of this concerns contracts already in force, the ones in the drawer nobody re-reads, not the deal still on the table. Today those obligations are read once, at signing, by someone expensive, then filed while the business moves underneath them. They surface again only when something has already gone wrong: the renewal has auto-extended, or the counterparty’s lawyer is the one who noticed the breach first. A system that re-reads the drawer every day does not need to be a better oracle than the lawyer who signed it. It needs to notice the gap starting to widen and to say so while there is still time to act. The single judgement was never the valuable part. The movement is.
🔍 Legal observability
There is a word for this in software, and as an engineer I keep reaching for it. We call it observability. A running system emits a continuous stream of signals about its own state, so you watch it live rather than reconstructing what happened after it has fallen over. Before observability, you debugged by reproducing yesterday’s crash.
Legal work, almost everywhere, is still pre-observability. A company’s contractual position is something you reconstruct after the fact, under pressure, usually because a dispute has forced the question. The review, the opinion, the periodic audit are all snapshots of a thing that changes continuously between them. I think the most consequential shift of the next couple of years is that the snapshot stops being the unit of legal work, and a continuously current picture of your obligations takes its place. Not “what did we agree two years ago”, reconstructed in a hurry, but “what are we obligated to right now, and are we meeting it”, answered before anyone has had to ask.

A picture on its own is not the point, though, and this is where I expect most of the category to stop. A dashboard that shows you the drift still leaves the work of closing it on the same expensive desk it always sat on. Seeing the obligation move and meeting it are different jobs. The first is a tool. The second is the thing a legal department actually needs, and it is the harder half, because it is where someone has to do the work rather than display it.
What this looks like sitting on the inbox
I find this easiest to think about not in the abstract but as the question we actually argue over when we decide what to build next. So let me put it where the work already is.
Almost every legal request a business generates ends up, in the end, as an email. It might begin as a message in a chat tool or a question in a corridor, but it lands in legal’s inbox, because that is the one surface the business already uses without being told to. So that is where the work belongs. A layer sits on the inbox, receives each request, works out what it is, does the routine work end to end, and routes the genuine judgement calls to a person with the context already assembled. None of that is novel on its own. The part that matters for trust is that the layer sits on the flow continuously, and in the open.
You are not handing it a contract and waiting to see whether you can trust the answer it hands back. You are watching it work, on everything, all the time, and everything it does is logged, attributable and replayable. That last part is not a feature footnote. Continuous monitoring creates its own duty: an obligation you can now see drifting is one you are now on the hook to have acted on, and the silent miss, the thing that should have surfaced and did not, becomes the failure you most have to engineer against. So the system records every decision, and every non-decision with it.
And once it is recording, the by-product is a kind of intelligence the department has never really had. Not just the obvious counts, how many NDAs came in this week, who handled what, how much went out without a lawyer ever touching it, but the pattern underneath them. Which objections your team keeps conceding even though the playbook says hold the line. Which clauses on third-party paper keep getting waved through against your own standard. Where the quiet gap between the policy you wrote and the practice you actually run has opened up. Nobody commissions a study to find that out. It is the kind of pattern that only shows up once something is watching the whole flow, day after day, rather than a sample of it. When a request reaches the kind of call that should never be made by reading a trend, the system stops and escalates to a person.

That changes what trust is. It stops being a gate you stand at, deciding once whether to let the machine through. It becomes a standard you can audit as it runs: work that is shown, retraceable, and stood behind, rather than a black box you cannot see into. And stood behind is meant fairly literally. The commercial model is built on outcomes rather than licences, which means the work is priced against an agreed standard rather than a seat, and when it falls short the rework, and the obligation to put it right, sit with the people who built the system, not only with the lawyer left supervising it. Visibility you cannot act on, and cannot hold anyone to, is not trust. It is a nicer view of the same exposure.
⚖️ What to do now is change what you measure
If any of this is right, the move for a legal team now is not to launch an eighteen-month programme to clean and tag every contract it owns before anything can begin. That is the old reflex, and it is the wrong one. It puts your most expensive people on your most basic work, which is the exact mismatch most departments are already drowning in, and it treats the current picture as something you build up front rather than something you get as a by-product of the work being done.
What I would change instead is the test you apply when you buy. Most legal AI is still bought the way the once-a-year scan was justified, by how impressive it looks in a single demo, on one good day. That test rewards the snapshot. It tells you almost nothing about whether a system can sit on your real flow of work for six months and hold a current picture of it. So change the question. Do not ask a vendor to dazzle you once. Ask what it does on an ordinary Tuesday when nobody is watching: what share of last week’s intake it closed without a lawyer touching it, what it escalated and why, and which obligation it surfaced that nobody had looked at since signing. Measure the work that got done, not the activity around it.
Visibility you cannot act on, and cannot hold anyone to, is not trust. It is a nicer view of the same exposure.
What the cheap thing makes possible
Here is the part the demo never shows, and the part Midjourney understood. The interesting question about a new capability is rarely what it can do today, on its own, once. It is what becomes possible when the thing that used to be expensive is suddenly something you can do ten thousand times a day. The scanner did not win by being a better scanner. It won by being cheap enough to turn from a test you take once into a thing that watches you, and that shift did more than improve screening. It moved the whole question onto different economics, and it changed what happens on an ordinary day.
The interesting question about a new capability is rarely what it can do today, on its own, once. It is what becomes possible when the thing that used to be expensive is suddenly something you can do ten thousand times a day.
Legal is at exactly that point, and most of the market is still measuring the old thing. The old question is whether an AI system can read a contract as well as a senior lawyer on a good day. The question that matters is what a legal department becomes when reading the whole drawer, continuously, costs almost nothing. The answer is not a faster lawyer. A faster tool hands the judgement back to you sooner and leaves it yours to sweat over, which is the ceiling of every instrument that only speeds the human up, and the reason a department whose real problem is expensive people doing inexpensive work never gets unstuck by buying one. What changes at scale is different in kind. The routine work that used to consume expensive people stops reaching them. The drift nobody had the hours to watch becomes visible. The lawyer’s day starts with the irreversible calls that were always theirs to make, rather than the inbox that never should have been. The work gets done by the system. The judgement stays with the person.
The work gets done by the system. The judgement stays with the person.
That is the two-year question, and it is not whether the capability arrives. I am fairly confident it does. What I am less sure of is how much of it survives contact with the mess: how much of a real contract reads cleanly enough to track, and how many obligations turn out to have no clean external number to read them against. That part is genuinely hard, and I will not pretend otherwise. The open question is whether legal teams change what they reward fast enough to use what does work, or keep buying the impressive one-day demo while the thing that would actually move their position runs quietly on the inbox, day after day, waiting to be measured by the right test. So the question worth asking of anything you put in front of your team is not how well it performs once. It is what your department becomes when it runs on every contract, every day, and the looking finally costs nothing.
✳️


