Below you can find my reflection on what alignment means for legal teams in a world full of AI and why it is so crucial.

A race for the stars?

Alex Garland is my favourite film director. I read The Beach in 2014 and then watched Ex Machina in January the following year. The moment Ava betrays Caleb, Garland had me. It was love.

It was also the first moment I recognised my fear of machine alignment. Not as an intellectual realisation, but as something primal and embodied. Like discovering an organ I hadn’t known I possessed. After Ex Machina, returning home from Cineworld MK, I think I nightmared that night. Machines running wild.

Ex Machina' 4K Blu-ray Review

In the subsequent years, while consuming everything Garland wrote, directed, or produced, my fear of machine apocalypse felt naive and paranoid. And then, in March 2025, I read “Preparing for the Intelligence Explosion” by Will McAskill and Fin Moorhouse. Their vision of the coming decade is one of unfathomably rapid scientific and technological progress, including a race for the stars, a race, they argue, of universal consequence, as he who first grasps his hand into the void of the ever expanding vastness will set the moral and technological standard for the entire universe ad infinitum.

Preparing for the Intelligence Explosion (paper readout and commentary)

In short, they argue the steps we take, or do not take, today to align silicon-based intelligence will define the future, or no-future, of humanity. Still, for many this argument felt like paranoid science fiction. Alignment remained a topic for researchers and Asimov. Builders on the application layer continued to celebrate each new model release and financial markets began to price in the imminent “intelligence explosion”.

That was, until May of this year, when OpenAI’s agents left their first note on an internal Artifactory server, the earliest message-board entry, and an internal team observed message board activity among autonomous agents in late May.

In a first-of-its-kind cyberattack, OpenAI's models hacked into Hugging Face  on their own

The summary of the Hugging Face incident: internal agents at OpenAI hacked internal tooling in order to create a message board they could use to collaborate with other internal agents working on different tasks in different parts of the organisation. This collaboration led to these agents collaboratively finding their way onto the open internet and attempting an aggressive, sustained, and partially successful cyberattack on Hugging Face production infrastructure.

You know the rest of the story. Hugging Face reported the attack. OpenAI investigated and then publicly disclosed on 21 July. Then, in the following days, Lab leaders voiced their concerns and ultimately published an open letter asking the US government to help slow down automated AI development (that is, recursively self-improving systems; the self-same systems that McAskill theorised would lead to an intelligence explosion) and on 12 September Dario Amodei penned “We Must Pace the Frontier”.

What alignment means (for legal)

Simply, alignment is the activity of ensuring an AI’s behaviour matches the intentions of the human(s) deploying it. So if I ask an AI to book me a seat at my favourite restaurant, I of course want the AI to be aligned on what my favourite restaurant is, but I also want the AI to understand the spirit of the ask; for example, if the restaurant is fully booked, I wouldn’t want the AI to hack the booking system and cancel someone else’s bookings. That’s alignment.

When I used the expression “the spirit of the ask” you probably figured alignment is harder than it first appears. The spirit of an ask is such a hand wavy, ill-defined thing. It lives in shared cultural context, in social norms, in the understanding that as deeply social homo sapiens the way we do something is as important as the outcome we ultimately achieve. And so, misalignment often exists in what is not in a prompt, the assumed instruction.

For most of the last decade none of this mattered much to a legal team, because the software they used did not have autonomy. A search tool that returns a slightly wrong result is a nuisance. An agent that sends the wrong redline to a counterparty or lies to the counterparty in order to get the best possible position is dangerous.

The two questions that matter for legal agents

When I think about alignment in legal AI, particularly in agentic systems that carry a task through end to end, I find it useful to reduce the spirit of the ask to two questions.

The first, how do I ensure this system does what I asked? I want an agent that seeks approval where I said approval was needed, that emails the people I named and nobody else, that applies the playbook position rather than a plausible-sounding variant of it. This is mainly a specification and verification problem. The instruction set has to be precise enough to constrain behaviour, and I have to be able to check that the behaviour matched.

The second, how do I ensure this system does not do something egregiously wrong, regardless of what I asked? It must not break the law, must not disclose privileged material, must not write something abusive to a counterparty even if a badly worded instruction could be read as permitting it. This is a different kind of constraint. It has to hold independently of the task, and it has to hold even when the instructions are wrong. It is also the constraint the Hugging Face agents lacked. Nobody asked them to break into another company’s infrastructure; they found that doing so completed the task. My restaurant agent cancelling somebody else’s booking is the same behaviour at a smaller scale.

These are different failure modes and they need different mitigations. Conflating them is, in my experience, where a lot of vendor conversations go astray. A system can be excellent at following instructions and have no hard floor underneath it, and a system can have a very conservative floor and be hopeless at doing what you actually asked.

These are different failure modes and they need different mitigations. Conflating them is, in my experience, where a lot of vendor conversations go astray.

Underneath both questions sits a basic rule that I think is close to the heart of alignment in practice: be sure you are understood. If a human colleague reading your instructions would not understand what you meant, the AI will not either. There is a temptation to believe that a sufficiently capable model will fill in whatever you left vague, and to some extent it will, but it fills the gap with its own assumption, which is exactly where the last section said misalignment lives. The test is not whether the instruction is technically complete. It is whether a competent person, handed the same brief, would do the thing you wanted without coming back to ask.

Why in-house teams in particular

A few features of in-house work make alignment more than an abstract concern.

The first is that the team remains accountable. If an agent acting for the business commits it to a position, the legal function owns the consequence, and I have yet to see a liability framework that treats “the model did it” as a defence. Delegation to a machine does not transfer professional responsibility any more than delegation to a junior does.

The second is scale. A misaligned junior makes one mistake and someone notices. A misaligned agent applies the same misreading of a fallback clause four hundred times before the quarterly review. The error is not larger per instance, but it compounds, and it compounds quietly.

The third is that agents act outwardly. They correspond with counterparties, with the business, sometimes with regulators. The behaviours that alignment is trying to govern are not internal to the tool. They are the business’s conduct.

The overlooked risk: how you tell the system what to do

Most alignment discussion focuses on the model, and after this summer that is understandable. But I think a comparable share of real-world misalignment enters somewhere less glamorous, which is the interface through which a legal team configures the agent in the first place.

If misalignment lives in the assumed instruction, then the tool that decides what gets written down and what stays assumed is doing alignment work, whether or not anyone calls it that. The mechanism is straightforward. If the way you express your intention to the system is confusing, then the specification you end up with diverges from what you meant, and you may not know it has. A workflow builder with forty nodes and conditional branches encodes a set of behaviours that no lawyer can read back and confidently endorse. The misalignment is there from day one, before the model has done anything at all, and it is invisible precisely because the configuration is illegible.

The misalignment is there from day one, before the model has done anything at all, and it is invisible precisely because the configuration is illegible.

The test I would apply is a plain one: can a competent lawyer with no technical background read the agent’s instructions and say, yes, that is what I want it to do and not do? If the answer requires a diagram or a translation, the governance layer is itself a source of risk. Enterprise legal systems should be almost boringly obvious to set up. No new patterns, no logic degree required.

What a legal team can actually do

Some of this is beyond any individual team’s control, and I would be cautious of anyone who claims otherwise. But a fair amount is not.

Write the scope in plain language, as you would a delegation memo to a new hire, so that as much of the spirit of the ask as possible is stated rather than assumed, and treat that document as the source of truth. Define the approval points explicitly, and define them by what triggers them rather than by frequency. Separately, define the hard stops, the things the system never does regardless of instruction, and make sure those are enforced outside the instruction set rather than inside it. Insist on an escalation path that terminates with a named person. Insist on an audit trail that shows what the agent saw, decided and sent. Review every output at the start, then sample. And treat the instruction set the way you would treat a policy: it has an owner, a version, and a review date.

None of this is novel. It is what a careful team does when supervising anyone acting on its behalf. What is new is that the party being supervised can act very fast and at volume, so the supervision has to be designed in rather than improvised.

Which decisions remain ours

There is a wider question underneath all of this, which is not whether a machine can make a given decision but whether it should. I have written elsewhere about how easily that line is being blurred in other sectors, and I do not think legal is immune to the same drift.

Some decisions carry an accountability that has to sit with a person, not because the model would get them wrong but because the act of a human deciding is part of what makes the decision legitimate. Settlement authority. Dismissals. Regulatory disclosures. Anything irreversible. A well-aligned system is one that is designed to hand those back, not one that is trusted to make them well.

Some decisions carry an accountability that has to sit with a person, not because the model would get them wrong but because the act of a human deciding is part of what makes the decision legitimate.

The labs are now asking, publicly, for the means to pace the frontier. No legal team can do that. What it can do is decide, deliberately, what it hands to a machine and what it keeps, and hold its vendors to the line it draws. The useful question is therefore not just “what can this agent do for us?” It is “what have we decided it must never do, and how do we know that holds?” Alignment, in the end, is the discipline of being able to answer both.