Architecture
11 min

Modernizing Legacy Systems with AI Agents

Legacy modernisation was too expensive to justify for decades: millions of lines nobody understands any more, no tests, no documentation. AI agents change that arithmetic, because they can read and explain unfamiliar code at scale. But 'feasible' is not 'automatic' — and the naive AI migration produces a new, worse kind of legacy.

The System Nobody Wants to Touch

Almost every company older than ten years has a system people only discuss in lowered voices. It runs. It posts payments, calculates premiums, manages inventory — and has done for fifteen, twenty, sometimes thirty years. And nobody fully understands it any more. The authors retired long ago or moved on to other firms. The documentation was either never written or describes a version that hasn't existed since 2011. Tests? A few, around the edges, for the harmless parts. The core — the millions of lines of COBOL, the bloated Java from the early days of J2EE, the PHP monolith that grew with the business like a coral reef — is a black box that does what it's meant to, as long as you leave it alone.

Everyone knows this system is a risk. It blocks new features, it chains the company to a dying technology, it needs specialists you can barely hire any more. And yet, year after year, it doesn't get modernised. Not out of inertia, but out of a perfectly rational calculation: modernisation was always more expensive than the pain of living with the old thing.

And that calculation was correct. Anyone who wants to replace a legacy system has to understand what it does first — not roughly, but exactly, down to every special case some clerk requested by phone twelve years ago and that has lived quietly in the code ever since. Reconstructing that understanding, with no tests, no documentation, no original authors, was an archaeological dig that took months and stayed patchy anyway. Every migration was therefore a bet: you rewrite what you only half understand, and hope the other half wasn't the part half the business depended on.

What Changed in the Arithmetic

This is exactly where AI agents come in — and not where the marketing places them. The value isn't that a model writes new code quickly. The value is that it can read and explain old code at scale.

That is the capability classic modernisation always foundered on. An agent can work its way through 400,000 lines of an unfamiliar language and, in hours, write down what a team of humans would have painstakingly assembled over weeks: which modules call which? Where is that one premium formula calculated, and at which seven places is it invoked? Which database table is dead ballast and which is the beating heart? What does this 900-line batch job actually do, the one nobody has opened since 2014?

Concretely, AI shifts four cost blocks that made migration uneconomic:

  • Understanding large codebases. An agent reads and summarises, maps dependencies, traces the data flow. Months of archaeology become days — and the result is more complete, because no human would have the patience to follow every dead path.
  • The missing tests. And this is the real lever: an agent can generate a test harness around untested code that pins down the current behaviour — before a single line is migrated. Exactly the safety net that was missing before.
  • Translating idioms. Mapping COBOL constructs onto Java patterns, turning a procedural PHP structure into something testable — the dull, error-prone grind that wears humans down and breeds careless mistakes.
  • Mapping dependencies. What hangs off what? An agent finds the implicit couplings no diagram ever recorded.

So the economics really have changed. Migrations that would have failed a cost-benefit comparison three years ago are now something you can put numbers to. That is no small thing — it's the shift that thaws a whole class of frozen projects.

The Right Order

But "feasible" is not "automatic." The most expensive mistake you can make now is to draw the wrong conclusion from the new economics — namely to instruct the agent to simply rewrite the whole thing. Do that, and you haven't learned the lesson of thirty years of failed rewrites; you've merely accelerated it.

The order that works is the same one Michael Feathers described twenty years ago in Working Effectively with Legacy Code — AI doesn't change the method, it only lowers the cost of each step:

  1. Characterise. First you lace up a safety net. You don't write tests that prove the code is correct — you write characterisation tests that capture what it actually does, quirks and all. The bug that's sat in the system since 2015, the one three downstream processes now rely on, is no longer a bug; it's behaviour you have to preserve. This is exactly where AI is priceless: it generates that harness around untested code at a pace that makes the exercise practicable in the first place.
  2. Understand. With the agent, you work out what the system does and why. Not to admire it, but to reconstruct the intent behind the code, which is written down nowhere.
  3. Migrate incrementally. You swap out one piece — a module, an endpoint, a batch job — for the new implementation. Small enough that a human can review the change in full.
  4. Verify. The safety net from step one tells you immediately whether the new piece behaves exactly like the old one. If it diverges, the divergence was either intended — or you've just found a defect before it reached production.

The crux: the tests come first, not the new code. Without the net, every migration is the old bet all over again, only with a model at the wheel that produces its mistakes with the same confidence as its hits. With the net, the bet becomes a controlled, reversible process.

Strangler Fig, Now with Agents

There's a pattern invented for exactly this kind of rebuild, long before AI: the strangler fig, named after the vine that slowly grows over a tree until it has entirely replaced its frame, while the tree stands the whole time. You build the new around the old, redirect traffic piece by piece, and only switch the old off once no path runs through it any more. No cut-over date where everything flips at once. No weekend where half the team prays.

Agents are what finally make this pattern practicable. The classic weakness of the strangler fig was the façade — the understanding work at the boundary between old and new, mapping every call that has to be intercepted and rerouted. That is exactly the work an agent speeds up: it finds the seams, explains what happens on both sides, generates the characterisation tests for the next piece you carve off. The fig grows faster. But it grows — branch by branch, each one signed off by a human — it doesn't drop from the sky as a finished tree.

The New Legacy Nobody Has Read

And here lurks the trap that can turn the whole exercise into its opposite. It's seductive, because it feels like progress.

An agent can translate a COBOL monolith into Java and produce code that compiles, runs, and looks plausible. Tens of thousands of lines, cleanly formatted, conventionally named. It's tempting to log that as done. But if no human has ever read and understood that code, you haven't solved the original problem — you've reproduced it. You've moved from a black box nobody understands to a new black box nobody understands. Except the old one is thirty years battle-tested in the field, and the new one isn't. You've reset the legacy's age to zero and sacrificed the one property that made it valuable: that it demonstrably works.

This is the most expensive kind of technical debt — not code that looks bad, but code nobody ever held in their head. A machine-translated monolith the team knows only from scrolling past is worse than the original, because it has lost the original's trustworthiness and kept its impenetrability. The translation is not the goal. The human understanding on the other side is the goal — the translation is merely the tool that gets you there.

When to Leave It Alone

Now the honest other side, because here the new economics tempt you into two fallacies.

The first: that AI finally makes the big-bang rewrite safe. It doesn't. The temptation to just let the model rewrite everything and skip the whole laborious incremental dance is precisely the classic rewrite disaster — only reached faster. The reason big-bang rewrites fail in droves was never typing speed. It was that you never fully capture the old behaviour, never fully test the new against reality, and on cut-over day discover a truth you should have known months earlier. An agent that spits out 200,000 lines in a week doesn't shrink that problem — it enlarges it, because now even more unchecked code flips at once. Speed without the net is not a solution; it's the same cliff with a longer run-up.

The second fallacy is subtler: that everything that can be migrated should be. It shouldn't. Some legacy is best left alone. A stable, running, change-averse system — the batch job that has driven the nightly settlement without complaint for twelve years, the module nobody touches any more because it simply works — is not a problem waiting for its solution. It's a solved problem. Modernising it just because it's old trades a proven system for the risk of a new one, without anyone drawing any benefit from the swap. The honest trigger for a migration is never age alone, but pain: the system concretely blocks a feature, it tears open security holes, it chains you to a platform that is genuinely disappearing. Absent the pain, standing still is the superior strategy.

Both fallacies come down to the same point. The economics of modernisation have improved — but the discipline matters more as a result, not less. Incremental, not in one go. Test-backed, not on hope. Human-verified, not machine-rubber-stamped. AI lowers the cost of every single one of those steps dramatically. It does not license you to skip one.

Conclusion

For decades, legacy modernisation was trapped behind a single, brutal cost factor: nobody could afford to understand old code thoroughly enough to replace it safely. AI agents dissolve exactly that factor. They read at scale, they explain, they string up the safety net of tests that was missing before. A whole class of projects that seemed frozen forever has suddenly become something you can put numbers to.

But the shift is in the cost, not the method. The right order remains the same as it was before AI — pin down what the system does with characterisation tests, understand what sits behind it, migrate piece by piece, verify against the net. Replace that order with "let the model rewrite everything" and you haven't built yourself a modern system; you've built a newer, untested legacy that nobody has read.

This is exactly how we approach legacy at NH Labs. We use agents to understand codebases no human still holds in their head, and to generate the missing tests before we move a single line. Then we migrate incrementally, in pieces a human can review and stand behind — never as a blind big-bang rewrite. What comes out the far end is not a machine-translated monolith nobody understands, but a system a team genuinely commands. That is the whole difference between a modernised system and a new legacy that merely looks younger.