Hi there, I’m Taariq. I lead Forward Deployment at Flank.
Today I’m going to share some insights from these enterprise deployments.
Deploying an agent is a hiring problem
There’s a well-trodden path in enterprise deployments, which starts with gathering requirements and ends in user acceptance testing (UAT). The agile methodology built upon this and brought iteration cycles down considerably.
Agentic deployments can’t follow the same pattern.
Firstly, model capabilities change so quickly that designing a system based on today’s capabilities doesn’t quite work.
Secondly, agentic is a buzzword that’s been thrown around, but you need to take the word ‘agentic’ literally to fully appreciate it. You’re giving a system agency to act on your behalf. That’s a very different operating model compared with using a tool to achieve a goal.
In truth, there’s no playbook with agentic AI. The closest we have is comparing agents to employees and I’d like to explore that thread in this post.

Writing a job description
A customer comes to you with a specific problem to solve. Unless it’s a top-down mandate to ‘use AI’, the problem is generally articulated well.
In my opinion, this is the most important part to get right because there’s a tension between scoping this description ‘properly’ (i.e. exhaustively) and building whatever comes to the customer’s mind.
If you define this job description in a rigid way, the customer will only get a fraction of the value they wanted from the agent, and ultimately the deployment will fail.
This is because the nature of agentic deployment is very different to others. There will be things the customer doesn’t mention up front – not because they’re negligent, but because we’re looking to automate human processes that often contain some variability and judgement in the moment.
Trying to document everything up front is akin to writing an onboarding manual for a new employee – but only writing it once, without giving them any opportunity to change what or how they do the work once they figure things out.
The other end of the spectrum, where you very loosely define what needs to be built, is equally dangerous. This way of working is essentially a gun for hire – you build exactly what the customer wants at the time. For example, when configuring a drafting agent, we were asked to include a tracked-changes version of the agent’s drafted agreement vs the template. This was a check introduced by the business at some point to ensure no typos got through. Agents don’t make typos – so this really isn’t needed, but what the team was actually saying is that the tracked changes document is how they built confidence in everything they send out.
What the first version of the job description leaves out
Stress-testing the design is easier said than done, so I’m sharing what’s worked for me. The key to a solid design is to unearth the unwritten process. The written process is useful up to a point, but it doesn’t cover the complexity of the particular task.
The first port of call is to shadow someone whilst they do the job. They’ll usually show you a relatively simple example of what they do. Ask for one that doesn’t follow that particular blueprint – this is harder for someone to show you because it requires more effort for them to pull together all the sources, explain why it’s an exception, and then show you what they did (whilst being wary that they may have broken some rules by going off the happy path)
Which leads me to the next point: design for exceptions as well as the happy path. It’s often comforting to go down the happy path because that’s where the majority of the volume often is (everyone likes to tell themselves this, but I’ve found it isn’t generally true).
My favourite question when probing this is to ask ‘what’s the most complex one you’ve done this week – one that an agent wouldn’t know what to do with because of the nuances and intricacies.’ It usually yields things that an agent can do perfectly well – if the deployment strategist can re-articulate the task in general terms.
There’s a balance to this too. You don’t want to design something that’s exhaustive – that would go against the golden rule of not overfitting.

Adoption is the most important test
You evaluate the performance of an employee soon after they join, and you do the same with an agent. You look at the outputs and you compare them to what’s expected. This is similar to the UAT phase in traditional deployments.
However, even when the agent’s evaluation is deemed ‘good’ that doesn’t mean it will be useful. That only happens when people trust the agent – and this gets built over time. At first, its outputs will be checked very closely, but this will gradually reduce as trust in the agent increases. Based on all of the deployments I’ve overseen, this trust-building phase is natural and no shortcut exists. Trust takes time to build (proportional to the outcome, in general) so this phase needs a lot of thought. The key question to ask before deploying an agent is “who’s doing the onboarding?”
This must be someone who does the job (and not someone from procurement or IT) and who has time to dedicate to the onboarding.
✳️


