Hi there, I’m Lorna. I lead Legal Alignment at Flank.
Every legal AI conversation arrives at the same question eventually: where’s the return?
I think it’s sitting in the part of the plan nobody wrote. The stretch after go-live, with no owner and no budget line.
The roadmap that ends at go-live
I’ve lost count of the AI delivery plans I’ve seen that end at rollout.
They’re often impressively detailed up to that point. Discovery, scoping, playbook building, testing, pilot, go-live, adoption workshops. Then the document simply stops, as if adoption were the finish line.
On more days than not, my work is the part that comes after that line. So I can tell you with some confidence that the go-live milestone is actually the starting line. Everything that determines whether an agent earns its keep happens after it.
Think about it. No business stops operating on launch day. Launch is when the real work begins - the customers arrive, the assumptions get tested, the operating model proves itself or doesn’t. We all know this instinctively about businesses. Somehow we forget it about AI agents.
Personally, I think this delivery-heavy planning habit explains a lot about why legal AI keeps disappointing people. Too much energy goes into testing and deploying. Not nearly enough goes into what happens next.
Everything that determines whether an agent earns its keep happens after it.
🏗️ You’re not buying legal tech, you’re hiring a new way of working
That’s because deploying an AI agent is a different kind of transaction to buying tech.
Traditional software is finished when it ships. You configure it, you train the users, and the vendor’s job is largely done. The tool does the same thing on day 400 as it did on day 4, whether you love it or not.
An agent doesn’t work like that. An agent is hired to do a job - review the NDA, turn around the supplier contract, tell sales whether they can accept 60-day payment terms on the deal that’s mid-negotiation. And how well it does that job depends on things that only exist after deployment: the playbook it follows, the feedback it receives, the escalations it raises and how those get resolved.

I always like to think of it as onboarding rather than installation. A new joiner doesn’t arrive fully formed. They arrive with strong general skills and close to zero knowledge of how your organisation actually likes things done. The first months are corrections and context. Nobody considers that a flaw in the hire. That’s simply how capability becomes useful.
Agentic AI is a whole new way of working. It’s a new system inside your organisation, and like any system, it needs a solid build followed by regular maintenance. Skip the second part and the first part slowly stops mattering.
In fairness, the delivery focus makes sense
I want to be fair to the people writing those roadmaps, because the delivery focus didn’t come from nowhere.
A launch has everything an organisation knows how to fund. It has a date, an owner, a budget line and a ribbon to cut. If you’re a general counsel, your board wants to hear about AI progress, and “we’ve gone live” is a clean answer to give them. If you run legal ops, you’re likely measured on delivery milestones, and the roadmap reflects exactly what you’re being asked to prove.
Maintenance has none of those things. It has no end date, no obvious owner and no moment of glory. Nobody gets promoted for tending something that’s already live, and no budget template I’ve ever seen has a line for “keep teaching the agent”.
Skip the second part and the first part slowly stops mattering.
And procurement is built around the old model. You evaluated the vendor, negotiated the contract, ran the implementation. When you bought your CLM or your e-signature tool, nobody stood up a permanent team to keep tuning it. Why would agents be any different?
However, it’s crystal clear to me that agents are different, for the reason above. Their performance is shaped by what happens after go-live, not before it. Applying the buy-tech mental model to a hire-a-new-way-of-working problem is how organisations end up with agents that launch well and then drift.
Nobody gets promoted for tending something that’s already live, and no budget template I’ve ever seen has a line for “keep teaching the agent”.
🔍 What AI agent maintenance actually looks like
Day-to-day, maintenance is less glamorous and more interesting than it sounds.
It means scrutinising performance at scale, not just spot-checking the demo cases. It means collecting feedback from the people actually using the agent and iterating the playbook based on what they say. And it means measuring impact honestly, including the uncomfortable questions. Is human involvement genuinely tapering off? If not, why not?

An example from our own operations at Flank. We run agents on our own standard contracts - events, sponsorship deals, NDAs, supplier agreements. No matter how comfortable I am with our internal playbooks on those documents, real end users interacting with the agent directly will always surface friction points that would otherwise have been absorbed by a human lawyer without anyone noticing.
Questions phrased in ways no lawyer would phrase them. Requests the playbook never anticipated, because a human lawyer would have handled them on autopilot and never thought to write them down. Before the agent, that knowledge lived in someone’s head and stayed there.
That friction is key information. Every friction point is an opportunity to adapt the playbook, and the agent, to what the organisation actually needs.
That’s the continuous learning loop the roadmaps leave out. And it’s the mechanism by which the agent becomes worth what you paid for it.
So what does success look like? I would describe it as a taper curve. In the early months, the humans in the loop make corrections often. Correction here rarely means the agent got something wrong, by the way. More often it means adjusting base positions, adapting the agent to your actual preferred negotiation posture on particular types of deals, proper tailoring to your organisation. Say your playbook holds firm on a 12-month non-solicit as standard, but on low-value sponsorship deals the business is relaxed about conceding it. The agent has to learn that posture, and the only way it learns is through the humans in the loop teaching it. In a word: training.
Over time, those corrections should genuinely reduce. Success is when human involvement sits at escalation points only - the moments where human judgment is needed. Even then, you keep a system for assuring quality throughout, like regular spot checks on samples of the agent’s output. The taper is earned, then verified on a schedule.
That's the continuous learning loop the roadmaps leave out.
As for who owns all of this: the legal function does, and that means legal ops just as much as the lawyers, wherever a legal ops team exists. I want to be clear that this is not another burden to load onto an already stretched legal team. Ultimately, agentic systems need a cockpit to be operated from, and someone at the controls. At Flank we have the benefit of a dedicated team for it - a legal AI alignment team, focused on aligning what the AI does to the organisation’s actual needs. Personally, I think every organisation deploying agents into legal work will eventually have one, whether it’s built by the lawyers, run by legal ops, or borrowed from the vendor. The name matters less than the cockpit existing at all.
And if I’m honest, I think this is the opportunity sitting in plain sight. For a GC, the alignment function is how legal AI becomes something you can stand behind in front of the board, with evidence rather than anecdotes. For a legal ops lead, owning the learning loop means owning the value story - you become the person who can actually show what the AI investment is returning, in a market where almost nobody can. That’s a far more interesting position to be in than being the person who ran the implementation.
This is where the legal AI ROI went
Now connect that back to the question everyone in legal AI keeps asking. Where’s the return?
MIT’s Project NANDA report, The GenAI Divide: State of AI in Business (August 2025), made headlines with the finding that 95% of enterprise GenAI pilots delivered no measurable P&L impact, despite $30 to 40 billion of investment. The methodology got picked apart, and fairly, but the pattern it described matched what many of us see on the ground: high adoption, low transformation. Tools that demo well and then sit brittle inside real workflows.

Closer to home, Thomson Reuters reported in 2026 that 82% of legal departments either don’t measure AI return on investment or don’t know whether they do. In other words, most legal teams have no learning loop running on their own investment either.
Imo those two findings are the same finding.
A pilot that ends at deployment has no mechanism for compounding.
Usage spikes at launch, then drifts as the un-ironed friction points pile up. Workarounds creep back in. Eighteen months later the tool gets blamed for underdelivering, when what actually underdelivered was the plan around it. Meanwhile, the deployments that do generate returns are the ones where somebody kept scrutinising performance, kept collecting feedback, kept iterating. The MIT authors found the winners were workflow-integrated systems built with partners who stuck around after launch. That’s maintenance by another name.
The timing matters here too. Thomson Reuters also reports that only 15% of organisations currently use agentic AI, while another 53% are planning or evaluating it. That’s a very large number of delivery plans being written right now. On current form, most of them will end at go-live.
For a legal ops lead, owning the learning loop means owning the value story - you become the person who can actually show what the AI investment is returning, in a market where almost nobody can.
A test you can run this week
If any of this stings, there’s a check that takes five minutes.
Whether you’re the GC who signs the budget or the legal ops lead who writes the plan, pull up your AI roadmap, or the vendor’s delivery plan, or the business case that got the budget approved. Find the go-live line. Read what’s written after it.
If the answer is a renewal date and not much else, you now know where your ROI is going to go.
What should be written there instead? Ownership of the learning loop, for a start. Who reviews performance, who collects end-user feedback, who iterates the playbook, and how often. A definition of what tapering human involvement should look like for your use cases. A spot-check schedule that survives after the novelty wears off.
I won’t pretend the steady state is a solved problem. I can describe the taper curve, but I can’t yet tell you how long the taper should take for a given organisation, or at what point a taper that isn’t happening means the product is wrong rather than the playbook is simply young. We’re learning that in real deployments, one feedback loop at a time (which is rather the point).
The direction, though, I’m sure of. Most agent deployments that disappoint were never going to be rescued at launch. The gap sits in the long, unbudgeted, unowned stretch after go-live - and almost nobody is planning for that part yet.
Lorna Khemraz is Lead Counsel at Flank and leads its legal AI alignment team, which exists to align what agents do to what the organisation actually needs.
✳️


