Modern Data

Master Data Starter

Starting models and matching rules for customer, product, and supplier data.

Data models, match and merge rules, and stewardship workflows for the master data domains most organizations need. We tune the rules against a sample of your records in the first cycles, so you see real match results early. It works with the major MDM platforms and with custom builds.

GIST stepGroundIterate
Download the PDF
What you get

One trusted record for each customer, product, and supplier.

Free download · PDF · 10 pages

Get the Master Data Starter overview

The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.

The problem

The same customer lives in five systems, five different ways.

Names are spelled differently, emails carry tags, phones come in four formats, and two different people share a name and a ZIP code. MDM programs spend months configuring match rules before anyone sees whether they work.

What this accelerator does

The Master Data Starter tunes match rules against a labeled sample of your own records in the first cycles, so you see real results early and prove the rules before configuring a platform.

What's included

Match rules you can prove before you commit.

Four parts, each adapted to your data, platforms, and controls. What we adapt for you is yours to keep.

01

Domain data models

Customer, product, supplier, and reference data: golden records, cross-references, lineage, stewardship, and change log.

02

Match, merge, and survivorship rules

Rules in plain configuration files, tuned against a labeled sample of your records.

03

Stewardship workflow

A review queue for uncertain pairs. Steward decisions override the score on every later run.

04

Scorecards and data contracts

Data quality scorecards, measured match precision and recall, and a contract template per source.

How it works

How matching works.

  1. Map and standardize

    Names folded and nicknames expanded, emails normalized, phones in E.164, addresses and identifiers cleaned.

  2. Block

    Only records that share a key, such as email, phone, or last name and ZIP, are compared.

  3. Score

    Each field has a comparator and a weight. The score is the weighted average over shared fields.

  4. Classify

    Match, review, or no match by threshold. Hard rules override, such as never merging different tax IDs.

  5. Steward decisions

    A data steward decides the review pairs. Those decisions hold on every later run.

  6. Cluster and survive

    Matches become golden records. Survivorship picks each value, and lineage records its source.

Technical detail

Under the hood.

Vendor-neutral Python and configuration, Azure first, with tests included from the start.

Matching engine
Standardization, blocking, scoring, clustering, and survivorship in Python
Comparators
Exact, Jaro-Winkler, token set, name tokens, and numeric
Survivorship
Most recent, source priority, most frequent, and most complete
mdm evaluate
Precision and recall against a labeled sample; fails CI below a set floor
mdm explain
Field-by-field scores for any pair of records
SQL models
Warehouse landing design for golden records, cross-references, and lineage
Proof in the package

Worked example: 359 customer records from three systems.

Records from a CRM, an ERP, and an online store describe 150 real customers, with the mess real data has.

Planted in the data
  • CRM duplicates in capitals, with no email
  • ERP names as “Last, First” with nicknames
  • Store logins with +shop tags and gmail dots
  • Phones in four formats, some mistyped
  • Two different James Smiths in one ZIP code
Result

With the rules as shipped, the engine merges with no false merges. Mistyped-phone pairs go to the stewardship queue, and steward decisions raise recall on the next run. The two James Smiths stay separate.

How we run it with you

Adapted in the first cycles, handed over at the end.

  1. 01

    Label a sample

    Your people mark which of a few hundred records, hard cases included, are the same entity.

  2. 02

    Tune and measure

    We adjust weights, thresholds, and blocking, and check false merges first after every change.

  3. 03

    Set thresholds

    With the data owner, based on review capacity and the cost of a false merge.

  4. 04

    Move to a platform

    Proven rules translate to your MDM platform, and we compare its results on the same sample.

You keep the models, the tuned rules, the labeled sample, and a precision check in CI that fails any rule change causing false merges.

Where it fits

Platforms, related accelerators, and limits.

We state the limits up front, and we recommend tools based on fit. We do not resell platforms.

Works with

  • Informatica MDM, Reltio, Profisee, Semarchy, and others
  • Custom builds in your warehouse
  • Spark for large volumes

Pairs with

  • The Data Readiness Assessment finds the duplicates and join gaps
  • Golden records land on the Data Platform Foundation
  • Clean master data feeds the Metrics Layer and models

Assumptions and limits

  • The engine suits tuning and up to a few hundred thousand records; larger volumes run on the platform or Spark
  • Standardization is US-style; other markets need country rules
  • Clustering is transitive, so the scorecard flags clusters to check by hand

Take the overview with you.

Get the 10-page PDF to share with your team, or tell us the decision you want to improve and we will tell you whether the Master Data Starter fits.