Why most AI pilots fail
Pilots don't fail on the model. They fail because nobody mapped where it was supposed to go. Five recurring causes, and what to do about each of them.
Forward Deployed Engineering
Your competitors can buy the same model you can. The difference will be who can get it into their own processes. That's what we do: we come in, map how the work actually runs, and hand over a system running in production, built on the ERP and CRM you already have.
AS DOCUMENTED
AS IT ACTUALLY RUNS
Email arrives
Copied to a spreadsheet
Keyed into the ERP
Approval
The problem
The MIT Media Lab's NANDA study (The GenAI Divide: State of AI in Business, 2025) surveyed 153 leaders and reviewed more than 300 public deployments. It found that 95% of generative AI pilots produced no measurable business return. It puts the cause down to organisational adoption rather than model capability.
Ask someone how their process starts and they'll say "an email arrives". The reality is it arrives from forty senders, no two formatted alike, half the time the number that matters is inside a PDF, and the rule for where each case gets routed lives in one colleague's head. They never wrote it down, because nobody ever sat next to them for eight hours and asked.
So the first step here is always the same: someone goes in, sits with the team, and watches how the work happens.
The FDE's judgement
The most expensive mistake is handing everything to the model. Most of the work is solved better, cheaper and more reliably by deterministic software. An FDE's job is to decide, step by step, which is which. Same process as above, redesigned:
Unpack attachments, drop duplicates, land on one format. This can be done without a model, so it is.
From PDFs, screenshots, forwarded threads. Judgement belongs here, because no two senders are alike.
Required fields, totals, vendor master. If it fails, it doesn't move on.
Tie to the purchase order and delivery note, propose a cost centre, flag uncertainty.
The finance clerk sees the proposal and the source on one screen. They press Enter, not the model.
Through the existing API. We don't migrate systems, we build on them.
Whatever the system won't finish queues up for a named person, with a reason attached.
The process
Each phase is a prerequisite for the next. Without evals you don't know whether it works; without discovery you don't know whether you're building the right thing at all.
We sit with the team and record how the work actually happens, not what the process document claims. You get an operating map, an ROI matrix of what is worth automating and what isn't, and a concrete build plan.
Nothing goes live before it is measured. We assemble a golden dataset from your own historical cases and run the system against it. You get a number for where the system stands, and the exact place the errors come from.
We build on what you already run. The system starts in shadow mode alongside your team, then earns autonomy one category at a time. Every step is logged and reviewable. If you can't see what it did, you won't trust it, and you'd be right.
What you can hire
One senior engineer with your team two to five days a week. They map the process, they build it, and they own it in production.
DetailsFixed-price, two to four week assessment of one department. You get an operating map and an ROI matrix. No obligation to continue.
DetailsAn FDE, a data engineer and an eval owner on one programme. Worth it when you're taking several processes live at once.
DetailsWriting
Pilots don't fail on the model. They fail because nobody mapped where it was supposed to go. Five recurring causes, and what to do about each of them.
Every company has two processes: the one written down, and the one that runs. AI deployments build on the first. Here is how to capture the second.
What an FDE does during the two to four weeks of discovery, who they talk to, what they pull from your systems, and exactly what you get at the end.
A consultancy hands over a deck, we hand over a working system. Discovery is our first phase, and the same engineer who mapped the process then builds it. That leaves no handover gap between the people who understood the problem and the people who wrote the code.
No, we do the opposite. If someone spent two years and a large budget rolling out a system, the last thing they need is another migration. We build on the APIs of what you already run, and connect those systems to each other.
Two to four weeks after discovery starts we can give you a concrete number for how much time and money sits in the process. The first workflow running in production is typically live in month three, and visible in shadow mode well before that.
Whichever gives the best accuracy for the money on that specific step, and we measure it rather than guess. We build so the model stays swappable: your own eval set remains the basis for the decision even when something cheaper or better ships six months from now.
That's what the whole design is for. A process can go right one way and wrong several hundred ways. We design the exception branches as carefully as the happy path: anything below the threshold queues up for a human with the reason attached.
Yes, and it is the most common arrangement. One senior engineer embeds with your team two to five days a week and drives deployments from the inside. Details on the pricing page.
Tell us which department burns the most manual hours. We'll come back with a concrete proposal for what discovery would look like at your company. We don't send templates.