Check the data before you fund the project

Data and AI projects are often approved on a slide and discover the data problems weeks into the build. A short readiness check on the real data behind one decision avoids that.

Datagist4 min read

A use case gets approved on a slide. A team is staffed, a budget is set, and the build starts. A few weeks in, the team finds that customer IDs are duplicated in the CRM, one source cannot be accessed without a security review, and the field the model depends on is mostly empty. By then much of the budget is spent, and the decision the project was meant to improve is no better.

None of these problems were hidden. Nobody looked at the data before the money was committed. A readiness check does that looking first, on one decision, with the real data.

Start from one decision

A readiness check that tries to assess “our data” in general produces a long report and no verdict. Start instead from one business decision you want to improve, for example which customers get a retention offer each month.

Sit down with the person who owns that decision and agree what the decision is and who acts on it, how you will know it has improved, and roughly what an improvement would be worth. Then list the systems that hold the data it depends on. That list is the scope of the check, and everything else can wait.

List the sources, then get real extracts

For each system on the list, record who owns it and how the data can be accessed. Note the business key that identifies a record, the fields the decision needs, how fresh the data has to be, and which other sources it must join to.

Then get the data itself. A recent full extract or a large sample is enough, or you can query the database directly where that is allowed. Documentation describes how a system is supposed to work. The extract shows how people actually use it, and the two are rarely the same.

Agree the thresholds before you run anything

Before profiling, agree with the decision owner what “good enough” means for this decision: how complete a field must be, how old the data can get, and what share of orders must match a known customer.

Setting these levels first matters more than it seems. If you set them after seeing the results, the discussion turns into an argument about whether the numbers are acceptable. If you set them before, the results speak for themselves.

What to check

Most readiness problems fall into a handful of types, and it helps to work through them in order. Start with access and ownership, since there is little point profiling data the team cannot get, or data with nobody behind it who can answer questions and fix problems at the source.

Then look at the data itself. Completeness asks whether the fields the decision needs are actually filled in. Validity asks whether values fall within the allowed codes and ranges. Uniqueness asks whether each business key identifies exactly one record, and timeliness asks whether the data is fresh enough for how often the decision is made. Where one source must join to another, measure how many keys find a match, because a join that silently drops a third of the records will quietly distort everything built on it.

Finally, identify which fields hold personal data and what rules apply to them. For those columns, keep counts and patterns only, never sample values, so the check itself does not become a privacy problem.

Rank the gaps and review them with the owners

The check will find gaps. Rank each one by the effort needed to fix it and by its effect on the decision. A gap that is cheap to fix and blocks the decision goes to the top. A gap that is expensive to fix and barely affects the decision may never need fixing.

Then walk the ranked list with the owners of each source. They often know things a profile cannot show, such as a field that is empty because its data moved to another system last year. Adjust the ranking where they know better.

End with a verdict

The check should end with a clear answer, backed by evidence. If the data can support the decision, proceed and plan the first build increment. If it can support the decision once specific gaps are closed, fund those fixes first and build afterwards. And if it cannot support the decision at a cost that makes sense, stop, and record why.

Stopping is a useful result. It costs a few weeks instead of a failed project, and it tells you which data investment would make the use case possible later.

Keep the configuration and the checks so they can be run again. As fixes land, re-running the same checks shows whether the gaps have actually closed.


Our Data Readiness Assessment runs this process with profiling and readiness checks built in, and ends with a draft Decision Brief for the decision owner to sign. If you have a data or AI project waiting for approval, tell us about it and we will tell you what the first step would be.

Working on something like this?

Tell us the business decision you want to improve. We will tell you what the first step would be and whether we are the right fit.

Start a conversation