Skip to content
Datos y Analítica

Data architecture: what suits a mid-sized operation

A data warehouse stores information already structured and modelled to answer known questions. A lake stores data in its original form for questions not yet formulated. The choice between them does not depend on volume but on how well defined the questions are that the organisation actually needs answered.

Below: what each model solves, why volume does not decide, what each really costs, the option almost nobody considers, and what to build first either way.

Conversations about data architecture tend to arrive carrying product names before they carry business questions, and that produces expensive decisions. It is worth starting with what each model solves, which is simpler than the terminology suggests.

And it is worth saying something uncomfortable up front: a good share of mid-sized operations in the region does not need either of them yet. What they need is for their transactional systems to expose trustworthy data and an orderly query layer on top. Saying so early tends to shorten the project rather than end it.

What each model solves

The warehouse solves the repeated question. A model is defined — customers, sales, products, time — the data is transformed on the way in, and it is queried quickly and consistently. Its strength is that the answer to the same question is always the same, and that is exactly the property management requires.

The lake solves the unknown question. Data is kept as it arrived, without deciding in advance how it will be used, and transformed at the moment of analysis. Its strength is flexibility; its cost is that without governance it becomes a repository where nobody can find anything.

The two do not compete: they answer different needs. The common error is adopting the second while expecting the benefits of the first, and then blaming the tool for the mismatch.

Why volume does not decide

The justification heard most often is data growth, and it is usually the least relevant. Current tools handle the volumes of a mid-sized operation without difficulty, including the ones that feel large from the inside.

What does decide is variety and uncertainty. If the data comes from three known systems and the questions are managerial — sales, receivables, inventory — a well-modelled warehouse is cheaper, faster and far easier to govern.

If instead there are diverse, loosely structured sources — activity logs, sensors, free text — and the questions are still being formulated, the flexibility of a lake justifies its governance cost.

What each really costs

The visible cost is infrastructure and licensing, and it is the smallest of the three.

The second is modelling. A warehouse requires deciding how the business is represented, and that decision consumes the time of people who understand the business, not just technical staff. It is the investment that returns most and the one most consistently underestimated.

The third is continuous governance. Both models degrade without maintenance: the warehouse when the business changes and the model does not, the lake when data accumulates that nobody catalogued. The second degrades faster and less visibly.

A practical rule: if nobody has time assigned to govern the result, the simpler model is the right one, whatever the problem.

The option almost nobody considers

Between having nothing and building a platform there is a middle point that resolves more cases than is usually admitted: a read-only replica of the transactional systems, with a layer of views on top translating tables into the vocabulary of the business.

It is not elegant and it has clear limits: no deep history, poor consolidation of very different sources, and heavy queries competing with the operation if the replica is not properly isolated. Within those limits it delivers reliable reporting in weeks with minimal governance cost.

The reason to consider it is not the saving but the learning. Six months operating that way reveals which questions the business actually asks, and that is worth more for designing the eventual model than any requirements workshop. The architecture then gets built knowing rather than assuming.

The question that settles it faster than any comparison

Architecture debates run long because both options are defensible in the abstract. One question usually ends them: name the three decisions the business wants to make better, and say which data each one needs.

If the three are answerable from known sources with agreed definitions, the warehouse wins and the scope is suddenly small. If two of the three cannot even be phrased precisely, no architecture helps yet and the honest recommendation is to work on the questions before the platform.

The exercise also exposes something uncomfortable and useful: how often the answer is that nobody makes a recurring decision from data at all, and the request came from wanting a capability rather than from needing one. That is worth discovering in a meeting rather than after a twelve-month build.

The mistake that is hardest to undo

Copying everything into an analytics environment just in case. It produces value quickly and reverses slowly: within two years there are copies of the customer master in several places, each with its own consumers, and nobody can say how many exist.

The alternative is not to forbid exploration. It is to agree where identifiable data lives and work against that source, so whatever replicates outward is already pseudonymised unless a case justifies otherwise and somebody authorises it explicitly.

The reason this is worth deciding on day one rather than day four hundred is that every copy acquires its own consumers. Removing one later means migrating whoever depends on it, and by then the people who built it have usually moved on. What starts as a convenience becomes a dependency nobody chose.

What to build first, either way

Three things, in this order, and none depends on the architecture chosen.

First, a trustworthy source for the figures already in use. Before enabling new questions, the current ones should have a single answer. It is the least attractive work and it unblocks everything else.

Second, traceability of origin: being able to say which system and which extraction each figure came from. It costs little at build time and is impossible to reconstruct afterwards.

Third, access control, which in Colombia and Mexico is conditioned by data protection law when the set contains personal information. That is a design constraint, not a later adjustment.

How the decision is made with evidence

With an inventory of sources, a map of the questions the business needs answered, and an honest assessment of who will sustain the result. It is a matter of weeks, not months.

A data and analytics maturity assessment produces that starting point by scoring source, quality, governance, model and consumption separately, which is what allows investing in the weak stretch rather than the suspected one. From there comes the analytics plan, and where the decision touches residency it is coordinated with the cloud one rather than taken apart from it.

Frequently asked questions

What is the difference between a data warehouse and a data lake?

The warehouse stores already-modelled information to answer known questions consistently. The lake stores data in its original form for questions not yet formulated. They do not compete; they answer different needs.

Does data volume decide the architecture?

Rarely. Current tools handle mid-sized volumes without difficulty. What decides is the variety of sources and whether the questions are defined or still being formulated.

What is most underestimated when budgeting?

Modelling and continuous governance. Modelling consumes the time of people who understand the business, not just technical staff; governance is what stops the result degrading, and a lake degrades faster and less visibly.

What should be built first?

A trustworthy source for the figures already in use, traceability of where each figure came from, and access control. None of the three depends on the architecture chosen.

Is there a middle option?

A read-only replica with a view layer in business vocabulary. It has real limits, and six months of operating it reveals which questions the business actually asks — worth more for the eventual design than any requirements workshop.

What is the hardest mistake to reverse?

Copying everything into an analytics environment just in case. Within two years there are copies of the customer master in several places, each with consumers who have to be migrated before any of them can be removed.

Andrés Lozada
Andrés Lozada
LinkedIn

Explore more from SUMāTO

Enterprise AI Enterprise Transformation Strategic Consulting AI Agent AI Contact Center Cybersecurity