Skip to content
Datos y Analítica

Analytics and personal data: what changes when the data is about people

An analytics programme working with data about customers, employees or patients is processing personal data, and that fact changes design decisions rather than adding a review at the end. The decisions it changes are the early ones — what is collected, what is kept, and who can see it.

Below: what the frameworks actually require of analytics, the purpose question, why anonymisation is weaker than it looks, retention, and where the practical controls go.

Data protection tends to arrive in analytics projects as a legal review shortly before launch, which is the most expensive moment for it to arrive. Everything it constrains was decided months earlier.

Treated as a design input instead, most of what it requires is cheap — and some of it improves the analytics independently of any obligation.

Nothing below is legal advice, and the specific obligations depend on the sector and on what the organisation told people at collection. What it is, is the set of design decisions that get expensive when they are made without the question having been asked.

What the frameworks require of analytics

In Colombia, Habeas Data governs the processing of personal data; in Mexico, the federal data protection law does the equivalent. The details differ and the shape does not: processing needs a purpose, the purpose has to be one the person was informed of, and the data kept has to be proportionate to it.

For analytics that translates into three concrete questions. Was this data collected for a purpose that includes analysis? Is the analysis being done on more data than the purpose needs? And can a person exercise their rights over data that has been copied into an analytical environment?

The third is the one that catches organisations out. A deletion request has to reach every copy, and analytical environments are where copies accumulate without an inventory.

The purpose question comes first

Data collected to deliver a service is not automatically available for any analysis somebody thinks of later. This is the constraint most likely to be discovered late, and it is decided at collection time by what the person was told.

The practical consequence is that consent and notice wording should be reviewed by whoever plans to use the data analytically, not only by whoever collects it. That review costs an hour and prevents an entire class of unusable dataset.

Where the purpose does not cover the intended analysis, the honest options are to obtain a basis that does, to work with data that is genuinely anonymous, or not to do it. Proceeding on the reasoning that nobody will notice is a decision the organisation is making, whether or not it is stated that way.

Anonymisation is weaker than it looks

Removing names and identifiers is de-identification, not anonymisation. A record with a postcode, a birth date and a gender is frequently unique in a population, and the standard for anonymous data is that re-identification is not reasonably possible — a considerably higher bar than removing the obvious fields.

This matters because anonymous data falls outside the framework, and teams reach for that exemption without meeting it. A dataset described as anonymised in a project document and trivially re-identifiable in practice is a worse position than never having claimed it.

The techniques that genuinely help — aggregation above a minimum group size, generalising precise values, suppressing rare combinations — all cost analytical detail. That trade is real and it should be made deliberately rather than assumed away.

Retention is where most exposure sits

Analytical environments accumulate. A dataset extracted for one analysis stays, gets copied into a notebook, gets exported to a spreadsheet, and three years later nobody knows it exists.

Each of those copies carries the same obligations as the original and none of the governance. Deleting the source does not delete them, and a rights request cannot be honoured against copies nobody has listed.

The control that works is unfashionable: a defined lifetime on every extract, enforced automatically rather than by policy, and a rule that analysis happens in a governed environment rather than on a laptop. It reduces convenience noticeably and it is the difference between a manageable position and an unmanageable one.

Third parties are part of the perimeter

Analytics rarely happens entirely in-house. A cloud platform, a visualisation tool, a modelling service, an agency running an analysis — each is processing personal data on the organisation's behalf, and the accountability does not transfer with the data.

What that requires in practice is unremarkable and frequently missing: a written arrangement stating what the provider may do with the data, where it is processed, how long it is kept, and what happens when the relationship ends. The last of those is the one nobody asks about while the relationship is going well.

Cross-border processing deserves its own line rather than being assumed. Where the data leaves the country, both the Colombian and Mexican frameworks have something to say about it, and the answer depends on the destination and the safeguards rather than on the provider's reassurance.

Where the practical controls go

Access control at the field level rather than the table level, so an analyst studying purchasing behaviour does not receive identity documents alongside it. Most tools support this and most implementations do not use it.

Logging of who accessed what, which is required in practice and is also the only way to answer the question that follows any incident. And masking applied at the point data leaves the governed environment, not at the point it is displayed.

A semantic layer helps here more than it is usually credited for, because it gives these controls one place to live rather than being re-implemented per report — which is the configuration that eventually gets one of them wrong.

What it costs and what it returns

The cost is real: less detail in some analyses, more friction in getting access, and an engineering effort that produces no visible feature. Understating that makes the programme harder to trust.

The return is not only avoided sanction. A governed analytical environment with known contents, known lineage and known retention is easier to work in, and the inventory it requires is the same inventory an organisation needs to answer any question about its own data.

Which is why the framing that works with a delivery team is not compliance. It is that the controls and the analytical hygiene are largely the same work.

That framing also decides who leads it. Run as a compliance project it produces documents; run as a data engineering project with a compliance requirement attached, it produces an environment people can actually work in.

Where to start

With an inventory of what personal data already sits in analytical environments and where it was copied from. It is usually more than expected, and the exercise takes days rather than weeks because most of it can be read from the systems themselves.

A data and analytics maturity assessment scores governance separately from source and quality, so an analytics plan can address the exposure and the capability in the same sequence rather than as competing priorities.

Frequently asked questions

What does data protection require of an analytics programme?

A purpose the person was informed of, data proportionate to that purpose, and the ability to honour rights requests across every copy — including the ones in analytical environments.

Can data collected for a service be used for analysis?

Only if the purpose the person was told about covers it. That is decided at collection time, which is why notice wording should be reviewed by whoever plans to use the data analytically.

Is removing names enough to anonymise data?

No. That is de-identification. A record with postcode, birth date and gender is often unique, and anonymous means re-identification is not reasonably possible.

Where does most of the exposure sit?

In retention. Extracts get copied into notebooks and spreadsheets, carry the same obligations with none of the governance, and cannot be reached by a deletion request nobody has listed.

Which controls matter most in practice?

Field-level access control, access logging, masking applied when data leaves the governed environment, and enforced lifetimes on extracts.

What does this cost the analytics work?

Less detail in some analyses and more friction in access. In exchange the environment has known contents, lineage and retention — which is the same inventory the analytics needs anyway.

Andrés Lozada
Andrés Lozada
LinkedIn

Explore more from SUMāTO

Enterprise AI Enterprise Transformation Strategic Consulting AI Agent AI Contact Center Cybersecurity