Prove your match rules before you configure an MDM platform

Most master data programs find out whether their matching works months in, after the platform is set up. Here is how to test match rules on a sample of your own records first.

Datagist5 min read

A master data program usually starts with a platform decision. The team picks an MDM product, configures it, loads the sources, and only then sees how well it matches records. If the matches are wrong, the fix is more configuration, another load, and another wait.

There is a quicker way to find out. Before anyone configures a platform, take a sample of your own records, decide by hand which of them are the same customer, product, or supplier, and test your match rules against that answer key. You learn in weeks whether the rules work, and you carry rules you have already proved into whichever platform you choose.

Why matching is the hard part

The data models in an MDM program are fairly standard. Customer, product, and supplier look much the same from one organization to the next. Matching is where the work is, because it depends on how your data is actually entered.

The same customer can appear in five systems five different ways. The CRM record was typed in capitals and has no email address. The ERP stores the name as “Last, First” and uses a nickname. The online store has a login with a +shop tag in the email and dots in a Gmail address. Across all of them, the phone number turns up in four formats, and in one of them a digit is mistyped.

Meanwhile, two different people can look almost identical. Two customers called James Smith in the same ZIP code are a common case, and a rule loose enough to join the five records above can merge that pair by mistake. No platform setting settles this for you. Someone has to decide how much evidence is enough to call two records the same, and that decision depends on your data.

Build an answer key first

Start with a labeled sample: a few hundred records that the people who know the data have marked as the same entity or different entities. Include the hard cases on purpose. Nicknames, typos, shared addresses, family members, and companies with several branches are exactly the records that expose weak rules, so they belong in the sample.

This takes a few working sessions, not months. It is also the most useful thing the business can contribute, because it records what “the same customer” means in your organization. With the answer key in place, every change to a rule can be measured. You are no longer arguing about whether a rule looks right; you can see which pairs it gets right and which it gets wrong.

How the rules work

Most matching approaches follow the same sequence. The first step is standardizing the data: folding names to one case, expanding common nicknames, normalizing email addresses, putting phone numbers into one international format, and cleaning addresses and identifiers. Much of what looks like a difficult match becomes an easy one once both records are written the same way.

Next comes blocking, which limits comparisons to records that share a key, such as the same email, the same phone number, or the same last name and ZIP code. Without it, every record would be compared with every other, and the work would grow far faster than the data.

Each pair that survives blocking is then scored. Every field is compared with a method that suits it, exact comparison for identifiers, fuzzy comparison for names, numeric comparison for amounts, and the field scores are weighted and combined into one. Two thresholds turn the score into a decision. Above the upper threshold the pair is a match, below the lower one it is not, and anything in between goes to a person for review. Hard rules sit on top of the score, so that, for example, two records with different tax IDs are never merged however similar they look.

The pairs in the review band go to a data steward, and the steward’s decisions should hold on every later run, so nobody reviews the same pair twice. Finally, the matched records are grouped into a golden record. For each field, a survivorship rule picks the value to keep, such as the most recent, the one from the most trusted source, or the most complete, and the record notes where every value came from.

Measure the two kinds of error

Test the rules against the answer key and count two kinds of error. A false merge joins two different entities into one. It is the costly mistake: invoices go to the wrong account, privacy requests reach the wrong person, and once other systems have used the merged record it is hard to undo. A missed match leaves the same entity as two records. That costs less, and stewardship catches many missed matches over time.

So tune for false merges first. Change a weight, a threshold, or a blocking key, run the sample again, and confirm that no false merges have appeared before you look at anything else. Only then work on missed matches.

Set thresholds with the data owner

Where the review band sits is a business decision as much as a technical one. A wide band catches more uncertain pairs, but somebody has to review them. A narrow band saves review time, but lets more mistakes through. Agree the thresholds with the data owner based on how many pairs the stewards can review in a week and what a false merge costs in that domain. Supplier records that drive payments may deserve a stricter threshold than marketing contacts.

Keep the check running

Once the rules are proved, keep the answer key and the measurement. Run them in your build pipeline, so that any rule change that causes a false merge fails before it reaches production. Keep adding the pairs your stewards decide to the sample, and it will grow with your data.

Then choose and configure the platform

Now the platform work starts from rules you have already tested. Translate them into the product’s configuration, run it on the same labeled sample, and compare the results with what you measured before. If they differ, you know exactly which pairs to look at. This also makes the platform decision easier, because you can test a shortlist of products against the same answer key rather than comparing demos.


Our Master Data Starter packages this approach: data models for the common domains, match and survivorship rules to tune against your records, a stewardship workflow, and a precision check that runs in your pipeline. It works with the major MDM platforms and with custom builds. If you are planning a master data program, tell us about it and we will tell you what the first step would be.

Working on something like this?

Tell us the business decision you want to improve. We will tell you what the first step would be and whether we are the right fit.

Start a conversation