Why most AI pilots fail
Pilots don't fail on the model. They fail because nobody mapped where it was supposed to go. Five recurring causes, and what to do about each of them.
Forward Deployed Engineering · Budapest
Your competitors can buy the same model you can. The difference won't be who has access — it will be who can get it into their own processes. That's what we do: we come in, map how the work actually runs, and hand over a system running in production, built on the ERP and CRM you already have.
The problem
The widely cited MIT survey found that the overwhelming majority of generative AI pilots never reach production. Not because the model is bad — because nobody mapped where it was supposed to go.
Ask someone how their process starts and they'll say "an email arrives". The reality is it arrives from forty senders, no two formatted alike, half the time the number that matters is inside a PDF, and the rule for where each case gets routed lives in one colleague's head. They never wrote it down, because nobody ever sat next to them for eight hours and asked.
So we don't start by building. We start by going and looking.
The FDE's judgement
The most expensive mistake is handing everything to the model. Most of the work is solved better, cheaper and more reliably by deterministic software. An FDE's job is to decide, step by step, which is which. Same process as above — redesigned:
Unpack attachments, drop duplicates, land on one format. No model involved — none is needed.
From PDFs, screenshots, forwarded threads. Judgement belongs here, because no two senders are alike.
Required fields, totals, vendor master. If it fails, it doesn't move on.
Tie to the purchase order and delivery note, propose a cost centre, flag uncertainty.
The finance clerk sees the proposal and the source on one screen. They press Enter, not the model.
Through the existing API. We don't migrate systems, we build on them.
Whatever the system won't finish doesn't vanish — it queues up, with a reason.
The process
Each phase is a prerequisite for the next. Without evals you don't know whether it works; without discovery you don't know whether you're building the right thing at all.
We sit with the team and record how the work actually happens — not what the process document claims. You get an operating map, an ROI matrix of what is worth automating and what isn't, and a concrete build plan.
Nothing goes live before it is measured. We assemble a golden dataset from your own historical cases and run the system against it. You don't get an opinion that it 'looks good' — you get a number, and the exact place the errors come from.
We build on what you already run. The system starts in shadow mode alongside your team, then earns autonomy one category at a time. Every step is logged and reviewable — if you can't see what it did, you won't trust it, and you'd be right.
What you can hire
One senior engineer with your team two to five days a week. They map the process, they build it, and they own it in production.
Details →Fixed-price, two to four week assessment of one department. You get an operating map and an ROI matrix. No obligation to continue.
Details →An FDE, a data engineer and an eval owner on one programme. Worth it when you're taking several processes live at once.
Details →Writing
Pilots don't fail on the model. They fail because nobody mapped where it was supposed to go. Five recurring causes, and what to do about each of them.
Every company has two processes: the one written down, and the one that runs. AI deployments build on the first — here's how to capture the second.
What an FDE does during the two to four weeks of discovery, who they talk to, what they pull from your systems, and exactly what you get at the end.
A consultancy hands over a deck. We hand over a working system. Discovery isn't our deliverable, it's our first phase — and the same engineer who mapped the process then builds it. There is no handover gap between the people who understood the problem and the people who wrote the code.
No — we do the opposite. If someone spent two years and a large budget rolling out a system, the last thing they need is another migration. We build on the APIs of what you already run, and connect those systems to each other.
Two to four weeks after discovery starts we can give you a concrete number for how much time and money sits in the process. The first workflow running in production is typically live in month three, and visible in shadow mode well before that.
Whichever gives the best accuracy for the money on that specific step — and we measure it rather than guess. We build so the model stays swappable: your own eval set remains the basis for the decision even when something cheaper or better ships six months from now.
That's what the whole design is for. A process can go right one way and wrong several hundred ways. We design the exception branches as carefully as the happy path: anything below the threshold doesn't disappear, it queues up for a human with the reason attached.
Yes, and it is the most common arrangement. One senior engineer embeds with your team two to five days a week and drives deployments from the inside. Details on the pricing page.
Tell us which department burns the most manual hours. We'll come back with a concrete proposal for what discovery would look like at your company — not a template.