The pilot worked. That is the part everyone forgets when they describe what happened next. Weeks one through six were genuinely good — a task that used to eat an afternoon stopped existing, the team was delighted, someone put a slide together. Then the line went flat and stayed flat, and six months later the only thing still moving was the invoice.
This is the most common shape we see at intake, and it is almost never a technology failure. The model did what it was asked. The problem is what it was asked to do.
The shape of a plateau
Plot value against time on an automation project and you get a step, not a curve. There is a jump when the manual work disappears, and then a flat line forever, because the work only disappears once. You cannot remove the same afternoon twice.
The flatness is not a sign something broke. It is the correct behaviour of the thing you built. A cost you cut is cut. It does not keep cutting.
What makes it feel like failure is that the cost structure underneath it is not flat. Model spend recurs. Integration maintenance recurs. Vendor seats recur, and tend to grow as more people get added "just to try it." So you have a one-time benefit sitting on top of a recurring bill, and the ratio gets worse every month by design.
A plateau isn't the project stalling. It's the project doing exactly what it was scoped to do, for the second month running.
Why the next workflow is harder, not easier
The instinct at this point is to go find another task to automate. That instinct is right in direction and badly wrong in expected value.
You picked first the thing that was most obviously worth picking: high volume, clean inputs, unambiguous output, a willing owner. That is the easy one. It is easy precisely because all four of those things were true at once, which is rare.
The second candidate has maybe two of the four. It is worth less and costs more, and the third is worse again. Our own intake numbers put it at roughly 73% of AI pilots never reaching a second workflow — not because the programme was cancelled, but because someone did the arithmetic on candidate two and quietly stopped.
So the programme stalls at exactly the point where it was supposed to start compounding. And the reason it cannot compound is structural: nothing the first system did was ever written down in a form the second system could use.
What the loop actually is
The difference between a pilot that plateaus and a system that appreciates is not model quality, prompt sophistication, or budget. It is three pieces of plumbing that almost never make it into scope.
None of that is exotic. It is a day or two of engineering on most builds. It gets cut because the demo does not need it, and the demo is what gets approved.
The honest test
Ask what your system knows now that it did not know at launch.
If the answer is "nothing — it does the same thing, just faster than a person," you have an efficiency play. That is a legitimate thing to own. Price it as one, fund it as one, and stop expecting a curve out of a step.
If the answer names something specific — which of two approaches wins more often, which inputs predict a bad outcome, which cases should never have been automated — then you have a loop, and the value of the thing you own is going up while you read this.
What we do differently
We scope the outcome event before we scope the model. It sounds backwards and it is the single highest-leverage change we made to our own process, because it forces the conversation nobody wants to have early: what, precisely, will we count as this having worked?
Teams that can answer that in one sentence build systems that compound. Teams that cannot are usually about to build a very good pilot.
That question is the first thing a CRFT Scan puts on the table, and the reason half the fee credits into the build — if the answer is "there isn't an outcome event here," you should know that before anyone writes code, not after month six.



