Skip to content
Nube

Right-sizing workloads before migrating to the cloud

A high first month in the cloud almost never comes from having migrated: it comes from having migrated at the size the server had, rather than the size the workload needs. In the data centre, over-provisioning was prudent; in the cloud it is paid every month.

What follows: why inherited sizing is expensive, what is worth measuring, what requires judgement, and how to sequence the exercise.

On owned infrastructure, sizing was done once, for three or five years, and with margin. Buying more than needed was the conservative decision, because expanding later meant another purchase.

In the cloud that same margin becomes a recurring cost nobody revisits, and it carries forward at every renewal.

Why inherited sizing is expensive

The usual practice when migrating is to replicate: if the server had sixteen cores, a sixteen-core instance is provisioned. It is quick, it reduces the risk of the change, and it is exactly where the overspend originates.

The problem is that the sixteen never described the workload: it described a purchase made years earlier, when growth was projected that may never have arrived.

To that is added the safety margin each layer contributed — the vendor, the architect, the administrator — until the final size corresponds to no measurement at all.

What is worth measuring

Real usage over a representative period. Processor, memory, disk and network across several weeks, including month-end close and any known seasonal peak. A quiet week produces a sizing that fails on the first busy day.

The percentile, not the average. The average hides the peaks and the maximum over-represents them. Sizing against a high percentile captures the real workload without paying for one unrepeatable night.

The hourly pattern. A workload that works eight hours and sleeps sixteen is a candidate for scheduled shutdown, and that changes the arithmetic more than choosing the right size does.

Dependencies between components. Measuring a server in isolation leads to moving it alone and discovering afterwards that it talked constantly to another that stayed behind, with the latency and traffic cost that implies.

What requires judgement

Some workloads carry margin as a deliberate decision, not an oversight. A system with unpredictable peaks and a high cost of unavailability is sized generously on purpose, and it is worth writing that down so nobody optimises it later out of ignorance.

There is also the question of what to do with what should not migrate as it stands. An application that only runs on a version no longer supported poses a choice — upgrade, replace or maintain — that is not an infrastructure decision.

And the choice between reserved and on-demand capacity depends on how certain the workload’s permanence is. Committing capacity for a year on something that may be replaced in six months saves on the rate and loses on flexibility.

What changes by type of workload

Stable, predictable workloads — core system databases, internal services in constant use — benefit from capacity commitments and from tight sizing.

Variable workloads — campaigns, period closes, batch processing — take advantage of elasticity, and there the expensive mistake is the opposite one: sizing permanently for the peak instead of growing and shrinking.

Non-production environments deserve their own criterion, because they tend to replicate production sizing without needing it and account for a larger share of the bill than assumed.

How to sequence the exercise

Instrument before deciding. If there is no usage history, the first step is not choosing instances but collecting data over a period that includes a full business cycle.

Group by pattern, not by application. Workloads that behave alike admit the same decision, and that reduces a long inventory to a few manageable categories.

And plan the review from the start. Sizing that is right at migration stops being right when usage changes; without a periodic review the saving erodes within months.

What has to exist first

An inventory of workloads with owners, monitoring capable of providing history, and an agreement about who authorises a change of size. Without the last, the adjustment gets proposed and not executed.

It is also worth fixing the measure by which the result will be judged. Total cost is not enough, because it grows with the business; cost per unit of work — per transaction, per user, per order — says whether efficiency improved.

The mistake that repeats in the second wave

The first migration wave is usually done carefully, because it is visible and has management attention. The second is done on momentum, replicating the first wave’s decisions without repeating the measurement.

That is where structural overspend creeps in: the sizes chosen for other workloads get adopted as the standard, and the organisation ends up with an internal instance catalogue nobody justified.

It is worth treating each wave as a new exercise, even a shorter one, and reserving the standard only for what has already been measured and behaved as expected.

How long should measurement run before migrating?

Long enough to cover a full business cycle, including month-end close and any known seasonality. Measuring one quiet week produces a sizing that fails on the first busy day.

Is it better to migrate at the same size and adjust later?

It is a defensible option when the risk of change is high, provided the adjustment is scheduled with an owner and a date. Without that, the provisional size becomes permanent.

Which indicator is worth following?

Cost per unit of work, not total cost. The total grows with the business; the unit figure says whether efficiency improved or worsened.

What about development and test environments?

They deserve their own criterion. They tend to replicate production sizing without needing it and weigh on the bill more than assumed.

Andrés Lozada
Andrés Lozada
LinkedIn

Explore more from SUMāTO

Enterprise AI Enterprise Transformation Strategic Consulting AI Agent AI Contact Center Cybersecurity