Aller au contenu
Article 9 min read

Leaving VMware: the real alternatives, and how to migrate without breaking production

Since the Broadcom acquisition, plenty of teams are looking for an exit. The four credible alternatives, what each actually costs, and the sequence that avoids an incident.

By Laurent Tulpan

Broadcom completed its acquisition of VMware at the end of 2023. Perpetual licences gave way to subscriptions, products were consolidated into broad bundles, and minimum billable thresholds went up. For a company quietly running two or three virtualized servers for the last decade, the invoice stopped bearing any relationship to actual usage.

Hence a question that keeps coming up: what do we move to, and how do we do it without breaking production. Here is the state of the ground, and the sequence that works.

First, the question to ask before leaving

Leaving on reflex is the mirror image of doing nothing, and it is just as expensive.

If your environment is heavily tooled around the VMware stack, with backup, monitoring, automation and written procedures all built for it, the real cost of exit sometimes exceeds the increase you are absorbing. This is not simply about moving machines: the whole chain has to be requalified, the operating procedures rewritten, and the team retrained.

So price the migration before deciding. That number does two things: it tells you whether leaving pays, and it gives you leverage if you choose to renegotiate instead. A renewal negotiated without a costed alternative in hand is not a negotiation.

The upheaval did produce one useful side effect. It forced a lot of organizations to ask a question they had never asked: what do we actually need? A meaningful number discovered they had been paying for years for high-availability features they never used. Start there, with an inventory of what you genuinely consume.

The four credible alternatives

Proxmox VE. Open source, with optional paid support. Today it is the most common exit for small and mid-sized companies. It assumes you will either build the internal skill or outsource the operations. Watch out for a thinner third-party tooling ecosystem, and documentation that is partly community-written and therefore uneven.

Microsoft Hyper-V. Included in the Windows Server licences many organizations already hold. It holds up well in a strongly Microsoft environment, with a directory in place and teams already trained on those tools. Watch out for Windows licensing coherence, which deserves close checking before committing, because that is where the surprises live.

Nutanix. An integrated solution combining compute and storage, supported end to end. Suits organizations that want one vendor holding the whole thing rather than an assembly of parts. Watch out for a high entry cost, and lock-in on the vendor’s stack that simply replaces the lock-in you are leaving.

XCP-ng. Also open source with optional paid support. Relevant for targeted needs and teams comfortable at the command line. Smaller community, so fewer ready-made answers when something unusual comes up.

None is better in the abstract. The deciding criterion is not technical, it is organizational: what can you operate, and what would you rather hand to someone else? A poorly monitored hypervisor with untested backups is more dangerous than a more expensive platform that is properly held.

What the migration actually costs

Four line items, three of them routinely underestimated.

Moving the machines. The visible one, and the simplest. Disk formats differ, but conversion tooling exists and works.

The drivers inside the guests. Changing hypervisor changes the input/output drivers the guest operating system sees. Every machine has to be opened, modified and restarted. Across thirty machines, that stops being trivial.

Requalifying the backups. The most dangerous item, precisely because it is invisible. Your backup tooling knows the old platform, not the new one. It has to be reconfigured, and then a real restore has to be tested. A successful migration whose backups no longer work is an incident that has not happened yet.

The operating procedures. Everything your team knows how to do becomes obsolete overnight: restarting a machine, adding disk, taking a snapshot, failing over to a degraded mode. Those runbooks need rewriting, and testing by someone who did not write them.

The sequence that works

Five steps, in this order, none skipped.

1. Inventory the real workloads. Not the machines, the workloads: what business role, what criticality, what depends on it. That inventory determines the target, not the other way round. Many migrations start from the choice of hypervisor, which is the surest way to discover mid-project that it does not fit.

2. Migrate one machine with no stakes, all the way to a restore. One machine, unimportant, but taken through the entire chain: cutover, monitoring, backup, then a tested restore. That machine validates the whole apparatus, not merely the technical feasibility.

3. Advance in waves, each with a rollback window. Each wave groups machines that depend on each other, and each has its rollback plan written before it starts. The question it answers: how do we return to the previous state, in how long, and what data is lost.

4. Requalify the backups before calling it finished. This is a milestone, not a closing formality. Until a restore has been tested on the new platform, the migration is not done, even if every machine is running.

5. Rewrite the procedures and retrain the team. The last milestone, and the one sacrificed when the schedule has slipped. Sacrificing it means creating a dependency on whoever ran the migration.

The trap of running this alongside the day job

A hypervisor migration needs sustained attention and intervention windows. Handed to an internal team already holding the daily load, it stretches, and a stretched infrastructure project accumulates intermediate states: part of the estate migrated, part not, two platforms to monitor, two sets of procedures. That is the riskiest configuration there is.

If the internal bandwidth to do this in one push is not there, it is exactly the kind of work to hand to someone on a bounded scope. That is what our legacy modernization engagements exist for.

Further reading

  • Legacy modernization : the engagement format for a bounded migration
  • Fractional CTO : when the decision itself needs someone senior on your side
  • Technical due diligence : pricing the options before committing a budget
  • Modernize or renegotiate: price the exit first : the decision above this one
  • Fractional CTO vs interim vs consultant : who should hold this call
Questions

Straight answers.

  • Should we actually leave VMware?

    Not automatically, and leaving on reflex would be its own mistake. If your environment is heavily tooled around the VMware stack, with backup, monitoring, automation and orchestration all built for it, the cost of exit can exceed the increase you are absorbing. The sound approach is to price the migration first, then decide: that number is also what gives a renegotiation any weight. Leaving is a decision. So is staying. Neither should happen by inertia.

  • Which alternative should we choose?

    It depends on what you can operate, not on an absolute ranking. Proxmox VE fits teams willing to build the skill or to outsource the operations. Hyper-V makes sense in an environment already deeply Microsoft, with a directory in place. Nutanix gives you one supported stack end to end, at a high entry cost and with lock-in that replaces the lock-in you are escaping. XCP-ng suits targeted needs and teams comfortable at the command line. The question to settle first is operational, not technical.

  • How long does a hypervisor migration take?

    Moving the machines is not the long part. Requalifying the whole chain is. Budget for inventorying the real workloads, validating one machine with no stakes all the way through to a restore, then successive waves each with a rollback window. The variable that stretches everything is requalifying backups and rewriting operating procedures, routinely underestimated because neither is visible on a dashboard.

  • What must be verified before calling the migration done?

    The backups, and a restore actually tested on the new platform. A successful migration whose backups no longer work is an incident that has not happened yet. Also confirm that monitoring sees the new machines, that your team's runbooks have been rewritten, and that the input/output drivers inside the guest operating systems have been replaced.

Next step

Recognize the situation? Twenty minutes is enough.

Describe what is happening on your side. We will say plainly whether we can help, and if another route would serve you better, we will say that instead.