Skip to content
Datos y Analítica

Master data: reaching a single view of the customer

The same customer appears three times because three systems captured them at three moments, with no shared rule about what identifies them. A single view is not solved by cleaning a database: it is solved by deciding what makes two records the same person.

What follows: why duplicates occur, what automates, what requires judgement, and how it is sustained over time.

The visible consequence is uncomfortable: a customer is offered something they already have, asked for a detail they already gave, or sent the same communication twice.

The invisible consequence is worse. When the customer count differs between areas, no figure derived from it — share, retention, value per customer — survives a review.

Why duplicates occur

The root cause is almost never carelessness. It is that each system was designed for its own purpose and captured what it needed: sales stored a contact, billing stored a taxpayer, support stored a user.

Each used a different identifier and none was obliged to agree with the others, because in its own context it worked well.

Manual capture adds to it: a name written three ways, a document number with or without separators, a personal email on one record and the corporate one on another. None of those variants is an error in itself.

What automates well

Normalisation. Standardising document, telephone, email and address formats before comparing resolves a significant share of the matches currently lost to differences in spelling.

Rule-based matching. Where a valid tax identifier exists and agrees, the decision is deterministic. That case usually covers most of the universe and needs nothing sophisticated.

Probabilistic matching. For the rest, comparing combinations of name, date, address and contact produces a similarity score. What matters is that the result is graduated rather than binary.

The review queue. Intermediate scores go to a person with both records side by side. Automating the preparation of that decision is as valuable as automating the decision itself.

What requires judgement

The identity rule is a business decision. If two people share an address and a surname, are they the same? If a company has two registered names, is it one customer or two? The right answer depends on what the single view is for.

So is the automatic merge threshold. Merging too readily produces an error that is hard to undo, because it mixes histories; merging too rarely leaves the problem intact. It is worth being conservative and widening the review queue at the start.

And there is a decision usually postponed: which system prevails when two records disagree on a field. Without that hierarchy, the single view inherits the conflict instead of resolving it.

What changes by context

In multi-country operations the tax identifier changes format and meaning, so a rule built for one does not transfer to another without adaptation.

When the customer is a company, identity also has hierarchy: parent, branches, units that buy separately. Flattening that structure simplifies the model and then prevents answering legitimate questions about concentration.

And because a single view concentrates personal data, the obligations of Habeas Data (Ley 1581) in Colombia and of the LFPDPPP in Mexico apply more strongly, not less: gathering scattered information in one place increases both the value and the responsibility.

How it is sustained over time

A one-off clean-up degrades within months, because the causes keep operating. What sustains the result is acting at the point of capture.

Validating at creation — correct format, a search for matches before creating, mandatory minimum fields — prevents the duplicate rather than correcting it afterwards, and it is far cheaper.

Measuring helps too. A simple indicator, such as the percentage of records with a valid identifier and the number of merges pending in the queue, shows whether quality is improving or whether it was cleaned once.

What has to exist first

An owner of customer data with authority to decide the identity rule, and an agreement about which system is the source for each field. Both are organisational decisions, not technical ones.

It is also worth limiting the initial scope to active customers. Cleaning the full history multiplies the effort and adds little to the decisions that motivated the project.

The order that avoids redoing the work

First agree the identity rule and write it down. It is half an hour of conversation that saves weeks, because everything built afterwards depends on it.

Second, normalise and measure the starting point: how many records hold a valid identifier, how many share an email or telephone, how many share a name. That measurement says how large the real problem is, which almost always differs from the perception.

Third, resolve the deterministic set completely before touching the probabilistic one. It usually covers most of the universe and generates no debate, so it produces a visible result while the thresholds for the rest are agreed.

Fourth, open the review queue at a volume the team can actually work. A queue of thousands is not reviewed: it is ignored, and the project loses credibility exactly when it was beginning to deliver.

Can it be solved with technology alone?

No. The mechanical part — normalising and matching — automates well, but the rule defining when two records are the same entity is a business decision someone has to take and sustain.

Should records be merged automatically?

Only where the match is deterministic, typically a valid tax identifier. For the rest a review queue is better, because a wrong merge mixes histories and is hard to undo.

Does the whole history have to be cleaned?

Usually not at the start. Limiting scope to active customers delivers the value that motivated the project for a fraction of the effort.

How are duplicates prevented from returning?

By acting at the point of capture: format validation, a search for matches before creating the record, and mandatory minimum fields. Without that, any clean-up degrades within months.

Andrés Lozada
Andrés Lozada
LinkedIn

Explore more from SUMāTO

Enterprise AI Enterprise Transformation Strategic Consulting AI Agent AI Contact Center Cybersecurity