Legacy system migration: how to modernize your company's software without interrupting operations
Modernizing a legacy system that handles your company's billing or logistics is a task of high technical and operational complexity. Day-to-day operational pressure and the fear that a failed migration could paralyze the business often delay necessary technical decisions, quietly increasing both technical debt and accumulated infrastructure costs over time.
Why the 'Big Bang' approach fails so often in practice
A Big Bang migration — building the entire new system in parallel and switching off the old one in one clean cut — is attractive on paper because it looks simpler to plan: a single project, a single cutover date, no need to coexist with two systems at once. In practice it fails frequently for a simple reason: a legacy system that's been in production for years has accumulated business rules, edge cases, and one-off fixes that were never documented anywhere except in the code itself. Rewriting everything from scratch means rediscovering every one of those rules under time pressure, and it's almost always the most important ones — the ones affecting a major client, or a critical accounting process — that get discovered right after the cutover, once it's already too late to revert without pain.
The Strangler Fig pattern
Instead of a 'Big Bang' migration (rewriting the entire system and switching off the old one), the better engineering practice is to replace the system incrementally. The Strangler Fig pattern lets the new architecture coexist with the old one through a routing layer (such as an API Gateway), migrating module by module until the legacy system is fully replaced. The name comes from the strangler fig tree, a plant that grows around an existing tree, gradually taking its place until the original tree is no longer needed to hold it up — the metaphor is exact: the new system grows by wrapping around the old one, instead of replacing it in a single blow.
How the routing layer works in practice
The component that makes this pattern possible is a middle layer — typically an API Gateway or a purpose-configured reverse proxy — that decides, for every incoming request, whether the legacy system or the new module that's already replaced it should handle it. At first, that layer sends practically all traffic to the old system, with just a single route pointing to the newly migrated module. Over time, more routes get redirected to the new system, until the legacy system stops receiving any traffic at all and can be switched off without risk, because by that point it has no active responsibility left.
How to choose which module to migrate first
The right criterion isn't 'the easiest' or 'the most important' — it's the module with the best ratio of low risk to demonstrable value. An isolated module, with few dependencies on the rest of the system and not critical to daily operations (report generation or a read-only catalog, for example), is ideal for validating that the pattern works without putting the business at risk. Migrating the most critical or complex module first — however tempting because 'it's what hurts most' — usually gets expensive, because the team still doesn't have real experience with the coexistence pattern between the two systems.
Strategy for a safe migration
- Identify and isolate the simplest yet valuable module (e.g. user registration or report generation).
- Build an interoperability layer to sync databases in real time between the new and old structures.
- Implement automated regression tests to ensure the new API responds exactly like the legacy endpoint.
- Progressively migrate network traffic using controlled rollouts (canary releases).
- Keep a documented and tested rollback plan for each migrated module, not just for the migration as a whole.
Hidden costs of postponing the migration
The decision to postpone modernizing a legacy system almost never gets made explicitly — it happens by omission, because there's always something more urgent on the roadmap. The cost of that postponement doesn't show up directly on any financial statement, but it's real: every new integration built on top of the old system is one more dependency that will have to be migrated later, every team member who leaves takes with them undocumented knowledge of how the legacy system actually works, and every platform or language version left un-upgraded (because the legacy system depends on a specific one) increases accumulated security risk. None of these costs show up in a quarterly report, but all of them eventually surface as a bigger, riskier migration than it would have been two years earlier.
How to communicate the migration to the organization, not just the technical team
A well-executed Strangler Fig migration is, by design, nearly invisible to the rest of the organization — and that's exactly what makes it hard to communicate upward. Without dramatic results to show overnight, it's easy for leadership to perceive the effort as slow or low-impact, precisely because it's working the way it's supposed to: without visible disruption. It's worth defining simple metrics in advance that can actually be communicated clearly — how many modules have already migrated, what percentage of traffic the new system already handles, how many incidents were avoided compared to the old system — so real progress stays visible beyond the engineering team, instead of depending on someone noticing the absence of problems.
The role of automated regression tests, in detail
The third item in the strategy below — regression tests that compare the new system's behavior against the legacy one — deserves more detail because it's the piece that makes it possible to trust the migration without relying on exhaustive manual review. The most effective technique is capturing real traffic (or a representative sample) from the legacy system, sending it in parallel to the newly migrated module, and automatically comparing both responses without the end user ever seeing a difference — the new system responds 'in shadow,' without making any real decisions yet. Once the responses consistently match over a reasonable period, that's when it makes sense to start redirecting real traffic to the new system, not before.
It's worth resisting the temptation to switch off the legacy system immediately after migrating the last active module. A grace period of several weeks, with the old system stopped but not yet deleted (receiving no real traffic, but available for a one-off lookup if something unexpected comes up), gives a reasonable safety margin before declaring the whole process closed and permanently decommissioning that infrastructure.
An additional benefit of the incremental approach, rarely mentioned, is that it reduces risk for the development team itself, not just the business. Rewriting an entire system in one shot is a long project with no visible results for months, and that kind of project tends to lose priority against urgent business needs — it's common for a full rewrite to get abandoned halfway because 'something more important came up,' leaving the company with two half-maintained systems. Migrating module by module delivers visible value every few weeks, which makes it much easier to sustain organizational support throughout the entire process.
It's also worth documenting each migrated module not just with its code, but with the reasoning behind the decisions made during that specific migration: what hidden business rules were discovered in the legacy system, what edge cases had to be replicated and why. That documentation, generated as a natural byproduct of the migration process itself, ends up far more valuable than any attempt to document the entire legacy system before starting — because it captures exactly the knowledge that was actually needed, at the moment it became clear it was needed, instead of a generic inventory written without that real-use context.
This staged modernization, carried out with discipline and patience, reduces organizational anxiety and ensures that key integrations with local accounting ERPs (such as Siigo or SAP) stay operational without interruption throughout the entire transition, from the first migrated module to the last.
