Legacy firmware has a specific way of becoming a problem. Nobody sets out to leave a system undocumented – it happens gradually, as the people who understood it move on, the reasoning behind its design decisions goes unrecorded, and the code itself becomes the only remaining record of what the system actually does. Eventually something needs to change, and there’s no one left who can explain why the code was written the way it was.
At that point, the honest answer to “can this be fixed?” is usually yes. But it needs a structured approach, not a hopeful one.
Why legacy firmware is harder to touch than legacy hardware
Legacy hardware, however old, can usually be inspected directly: you can measure it, probe it, and observe what it does. Legacy firmware hides its logic inside code that may reference assumptions, timing relationships, and hardware quirks that are no longer written down anywhere. A change that looks safe in isolation can break a dependency nobody remembers existed, because the person who would have remembered it left the company two developers ago.
This is what makes legacy firmware support genuinely risky rather than simply inconvenient. Every undocumented system carries invisible constraints, and the only way to find them safely is to go looking deliberately, rather than discovering them the hard way through a regression in the field.
What “no documentation” actually means in practice
“No documentation” rarely means literally nothing exists. It usually means what exists is partial, outdated, or was accurate for an earlier version of the system that has since been modified without the documentation being updated to match. The result is often worse than having no documentation at all, because it’s easy to trust a description that no longer reflects what the code actually does.
The practical implication is that documentation, where it exists, should be treated as a hypothesis about the system’s behaviour, not a reliable statement of it – something to be checked against the code and against observed behaviour, rather than assumed correct.
A structured approach to understanding code you didn’t write
Isolate before you investigate
The first step in understanding an unfamiliar codebase safely is separating the parts that matter for the change you need to make from the parts that don’t. Most embedded applications are smaller, in terms of the logic that actually matters for a given fault or feature, than their full codebase suggests – a great deal of the code is I/O handling, communication protocol implementation, or boilerplate that can be set aside while you focus on the specific logic in question.
Mock the hardware so you can run it anywhere
Once the relevant logic is isolated, the physical hardware interfaces – GPIO reads and writes, timers, communication buses, on-chip storage – can typically be abstracted behind mock implementations that run on a desktop rather than the original target. EEPROM reads and writes become file accesses. Hardware timers become desktop OS timers. Bus messages become message queues. This isn’t a simplification of the problem; it’s what makes the problem tractable, because it lets you run the application repeatedly, inject arbitrary conditions, and control timing precisely – all without needing the original physical product in hand, and without the risk of damaging or tying up hardware that may be scarce, in service, or expensive to access. This is the same technique behind the fire safety curtain investigation described below, where mocking the hardware let us fully diagnose and fix a legacy fault without ever touching the physical product.
Diagnose against the design’s original intent, not just its current behaviour
Once the code can be run and manipulated freely, the diagnostic process is about comparing what the system currently does against what it was clearly designed to do – inferred from the parts of the logic that are still internally consistent – rather than simply patching the symptom that’s currently visible. A fault introduced by a later, undocumented change is often best understood by finding where the code stops matching its own original design pattern, since that’s usually where an assumption was broken without anyone realising it.
What this looked like on a real project
We were brought in to investigate a fault on a fire safety curtain system – roller shutter equipment installed on client sites at significant cost – where the curtain would deploy but occasionally fail to retract, leaving users trapped and unable to reach the release button. The client’s own contact had described the problem as “intractable,” and previous fault investigation attempts had been unsuccessful.
The application ran across two 8-bit CPUs, communicating over a combination of GPIO and CAN bus, with settings stored in on-chip EEPROM. The code had originally been well structured, but it was clear that either the original requirements no longer applied, or a different developer had since become involved – and the affected feature was exactly the one the customer had flagged as problematic.
Rather than schedule a site visit to test against physical hardware, we abstracted the I/O layer: EEPROM reads and writes became file accesses, GPIO signals were scripted against a desktop timer, and CAN bus messages became OS message queues. With both MCU applications now running on a desktop, we could inject any error condition, at any timing, without touching the physical product at all – handy when the product in question weighs hundreds of kg and requires a reinforced floor. That setup revealed the actual fault quickly: a later change had introduced a lengthy blocking section that prevented a critical release signal from being polled correctly, breaking the system’s ability to recover from a brief activation. Moving that check back into an interrupt service routine, in line with the original design’s intent, resolved the issue.
When this can be done without physical hardware at all
The fire safety curtain investigation is a useful illustration of a broader point: a properly isolated and mocked legacy codebase can often be fully diagnosed and fixed without needing a site visit, without transporting or reinstalling physical equipment, and without the delay and expense either of those options carries. This isn’t always possible – some faults genuinely require physical hardware to reproduce – but it’s more often possible than teams assume, and it’s worth ruling out before committing to a more expensive investigation route. Hardware-in-the-loop testing and rigorous integration testing practices, applied to legacy code rather than only to new development, are what make this approach viable at all.
Deciding whether to fix, rewrite, or replace
Not every legacy firmware problem should be solved by patching the existing code, and it’s worth reading up on how to vet a partner for this kind of work before committing to either route. Sometimes a system is undocumented enough, or has drifted far enough from any coherent original design, that a structured rewrite of the affected subsystem is more reliable than continuing to build on top of it. That decision should follow from the investigation, not precede it – you need to understand a system properly before you can judge whether it’s worth preserving. Design verification practices applied to a legacy system, even one that predates any formal verification process, are what make that judgement a reasoned one rather than a guess.
This kind of work sits alongside our wider embedded software and firmware development services, and you can find more examples in our case studies.
If you’re supporting a product whose original firmware team has moved on and you need to make changes without the risk of introducing regressions into code nobody currently understands, book a free 30-minute consultation with one of our engineers.
FAQs
Can legacy firmware be maintained without the original developer?
Yes, in most cases. It requires a structured approach – isolating the relevant logic, abstracting hardware dependencies so the code can be run and tested independently, and diagnosing against the system’s original design intent – rather than assuming the code is unreadable without the person who wrote it.
What does it mean to abstract hardware dependencies in legacy firmware?
It means replacing direct hardware interfaces – GPIO, timers, communication buses, on-chip storage – with equivalent implementations that run on a desktop, so the application logic can be exercised, tested, and debugged repeatedly without needing the physical hardware or product on hand.
Is it safe to modify undocumented legacy firmware?
It’s safe when changes are made after a proper investigation, since undocumented systems often carry invisible dependencies that aren’t obvious from reading the code in isolation. A structured diagnostic process, rather than a direct patch to the visible symptom, significantly reduces the risk of introducing a regression.
Can a legacy firmware fault be diagnosed without a site visit?
Often, yes. If the relevant application logic can be isolated and its hardware interfaces mocked, the resulting code can typically be tested and debugged on a desktop, without needing physical access to the original equipment or installation site.
How do you know if legacy firmware should be fixed or rewritten?
That decision should follow a proper investigation of the existing system rather than precede it. Once the code’s actual behaviour and design intent are understood, it becomes possible to judge whether targeted fixes are reliable or whether the affected subsystem has drifted too far from a coherent design to safely build on.
What should I look for in a partner to support legacy firmware?
Evidence of having diagnosed and resolved faults in undocumented or partially documented codebases before, a structured process for isolating and testing legacy logic independently of physical hardware, and a track record of doing so without introducing new issues into code they didn’t originally write.
Other services we offer
- Embedded Design Verification
- Embedded Testing
- Failure Mode and Effects Analysis
- Embedded Software and Firmware Development Services
Industries we serve
Related blogs
- When a Firmware Engineer Leaves Mid-Project: Protecting Product Knowledge Before It Walks Out the Door
- How to Choose an Embedded Software Development Partner (And the Questions to Ask)
- The Firmware Decisions That Delay Your Product Launch
- Reliability Testing in Embedded Systems Explained
- What FMEA Means for Embedded Firmware Development
- The Missing Link Between Requirements Documents and Hardware-in-the-Loop Testing
- Embedded Bootloaders Explained
- Introduction to Integration Testing
- Yocto vs Buildroot: Choosing an Embedded Linux Build System