.crft
← All learningsScope your system
Data7 min readJan 2026

The five levels of data readiness

What each level actually unlocks, and the one you can't build a learning system below.

Fritz Desir
Founder, AWSM LABS
Golden bokeh lights

"Our data is a mess" is the single most common thing we hear in an intake, and it is almost never useful information. Everyone's data is a mess. The question that matters is which kind of mess, because some kinds cost you a week and some kinds mean the system you want cannot be built yet at any budget.

So we stopped asking whether data is ready and started scoring it, one to five, per system in scope. The score goes in the CRFT Scan report next to each opportunity, and it does more work than any other number in there — it is usually what decides the order of the build.

The five levels

01
It exists, somewhere
The information is in the business — in inboxes, in someone's head, in a PDF a supplier sends monthly. There is no system of record you could query. At this level you cannot build anything that learns; you can build something that helps a person do the work while quietly creating the record.
02
It's captured, but unstructured
Tickets, call notes, freeform fields. A model can read it, which feels like readiness and isn't — you can extract from it, but you cannot reliably join it to anything else or count it. Good enough for assistance. Not good enough for a decision you'd defend.
03
It's structured and queryable
Events land in a warehouse with consistent shape. You can answer questions about the past. Most companies who describe themselves as data-driven are here, and this is the floor for a production system that makes decisions.
04
Outcomes are attached to inputs
Not just what happened, but what happened next — the ticket closed or reopened, the lead converted or went dark, the shipment held or slipped. This is the level where a learning loop becomes possible, because supervision exists without anyone hand-labelling anything.
05
The loop is closed and running
Outcomes flow back into retrieval, scoring or routing on a schedule, and the system's behaviour this month differs from last month because of what it observed. Very few organisations are here on any workflow. The ones that are have a moat.

Level four is the line

Everything below four can be made useful with AI. Only four and above can compound.

That is the whole reason we score it. A team at level three asking for a system that "gets smarter over time" is not asking for something impossible — they are asking for something that requires one more piece of work first, and it is much cheaper to say so in week one than to discover it in month five.

Level 3
You can automate a decision
The system can act consistently on what it sees. What it cannot do is find out whether it was right, so it will act exactly as well in year two as on day one.
Level 4
You can learn from the decision
Every action creates a labelled example at no marginal cost. This is the only cheap supervision most companies will ever have access to.

The move that unlocks the most

When a Scan comes back with a high-value opportunity sitting at level two or three, the recommendation is frequently not the opportunity itself. It is a smaller, duller build that moves the readiness level — structuring one data source, or emitting one outcome event that nobody was recording.

That is a hard thing to sell and we say it anyway, because the alternative is worse. An agent built on level-two data does not fail loudly. It produces plausible output on bad inputs, quietly, for months, and by the time anyone notices, trust in the whole programme is gone.

Sometimes the honest first build is the one that makes the second build possible. We would rather say that in week one than defend a model that has been guessing since March.

Reading your own score

You can approximate this without us. Pick the workflow you most want to put a system inside and ask, in order:

  • Is there a system of record, or is it in people's heads? (1 vs 2+)
  • Could I count the cases, or only read them? (2 vs 3+)
  • Can I tell, from data alone, how each case turned out? (3 vs 4+)
  • Does anything currently read those outcomes and change behaviour? (4 vs 5)

The first question you answer "no" to is your level, and the gap between that and four is your real project — whether or not anyone has scoped it as one.

We score readiness per system, before anything gets built.
Scored opportunities, quantified value, data-readiness levels, and the workflow as your team actually runs it. Half the fee credits into the build.
Scope your system →
getcrft.ai

Build the intelligence your company compounds on.

Five inputs, a published price, a report in 48 hours. Decide with numbers — then own what gets built.

Build your scope + price