Skip to content

Making the case for technical debt without using the words technical debt

“We should rebuild this properly” never gets funded, and that is rational. Three ways to convert technical debt into delay, risk and money, so it competes with the requests that do get funded.

By Laurent TulpanPublished 9 min read

The sentence exists in every engineering team, and it never gets anything: we should rebuild this properly.

It gets nothing because it asks an executive to arbitrate on a criterion they have no way to evaluate. Set spending that produces revenue against spending that produces tidiness, and the outcome is decided before anyone sits down.

The problem is not that the business does not understand engineering. It is that the request is denominated in a unit nothing else on the agenda uses.

What the debt actually costs, in three places the finance system already tracks

Technical debt is not paid in one installment. It is paid in small amounts, continuously, across three lines that already exist in your accounts without being attributed to it.

Extra time on every change. This is the largest and the easiest to measure. Take the last five change requests and compare the time actually spent with what the same change would have cost on a sound part of the system. That difference is not a projection. It is money invoiced this year, and it is sitting in your tickets.

Time spent not breaking things. On a fragile system, part of the work is verifying that a change broke nothing elsewhere, because nobody can establish that by reading the code. This time is invisible in estimates, where it is folded into the total, and it grows faster than everything else.

Concentrated knowledge. When one person is the only one who can safely touch a component, that is not a skills problem, it is an operational risk. It prices out through one question: what happens if that person is unavailable for three weeks. If the answer is that work stops, the debt carries an insurance cost nobody has provisioned. That question deserves its own conversation, and it usually turns out to apply to more than one component.

Three translations into units that compete

Into delay. The most legible for a general manager. “A change to invoicing takes three weeks; the same change to ordering, which we rebuilt last year, takes four days.” Both figures come from your own history, neither requires expertise to understand, and together they make the gap tangible.

Into risk, expressed as hours of downtime. Not “this is risky”, but “if this component fails we have no recovery procedure and restoring service depends on one person”. Hours of downtime multiply by the cost of an hour of downtime, and the result belongs in a budget table. If nobody has ever costed an hour of downtime, that is the first thing to establish, because without it every reliability investment looks expensive: you are comparing it to zero.

Into opportunity cost, when the opportunity is real. “We cannot open the portal to customers because the system cannot handle external access.” That only counts if the portal is genuinely planned. Invoking hypothetical opportunities weakens the case, because the person across the table can tell when an argument has been constructed for the occasion.

The map that makes the decision possible

A list of debts does not prioritize itself. Two axes are enough, and they can be filled in during one meeting with the people who touch the system.

Does it change? Is this part receiving change requests, or has it been stable for years?

What happens if it fails? Does the business stop, or can people work around it for a few days?

Crossing the two gives four situations, and only one of them justifies immediate funding.

It changes and failure stops the business. The priority, without discussion. Every change here costs too much and the exposure is maximal.

It changes but failure stops nothing. Improve it in passing, on each change, rather than launching a program. This is the cheapest debt repayment that exists.

It does not change but failure stops the business. The issue is not the debt, it is recovery. What you need is a tested restart procedure and a second person able to run it. Far cheaper than a rebuild, and it addresses the actual risk. This is also where an untested backup usually surfaces, which is a different problem wearing the same clothes.

It does not change and failure stops nothing. Do nothing. That is a decision, not an oversight, and it deserves to be written down so it is not re-litigated every six months.

That fourth box is the one missing from most remediation plans a vendor will propose. In those plans everything is a priority, which is the same as saying nothing is.

Why the full rewrite is almost always the wrong answer

The argument is attractive: the system is old, we start clean, it will be faster than untangling.

It fails for three reasons, and they recur.

The new system must do everything the old one did, including behaviors nobody can explain and which surface on cutover day, when a user reports that the special case for one particular customer no longer works.

The old system keeps living during the rewrite. Urgent changes get made twice, or the old one freezes and the business waits.

Cutover is a single high-risk event. Everything changes on one day, and if you have to roll back, you roll back to a system that has not evolved for months.

Replacing in waves avoids all three. One function at a time, both systems coexisting for the duration of each switch, each wave delivered and verified before the next begins. That is exactly how we rebuilt more than ten management applications for an international group, and the reason was not comfort: an outage in invoicing was not an option.

The sentence to replace

“We should rebuild this properly” asks for nothing specific, so it gets nothing.

“Over the last six months, four requests touching this area took three times the quoted time, one person is the only one who can work on it, and two changes planned for next year go through it. Bringing it up to standard costs this much, and the alternative is continuing to pay the difference on every request” asks for something specific, in a unit that compares.

The second sentence does not speak engineering better. It speaks the language of the person deciding, which is the only way to get a decision rather than a deferral.

Further reading

The questions this raises, answered straight.

Why does technical debt work never get approved?

Because it is presented in a unit the person deciding cannot evaluate. Faced with spending that produces revenue and spending that produces tidiness, the choice is made before the meeting starts.

The problem is not that executives fail to understand engineering; it is that the request arrives denominated in code quality, and nothing else on the agenda is priced that way.

Should we pay all of it down?

No, and trying to is the surest way to pay none of it down. Debt in a part of the system that no longer changes costs nothing: it is stable, it works, and leaving it alone is the rational call.

What deserves funding is debt sitting on the path of the changes you already intend to make, because that is where the interest is actually charged.

Is a full rewrite ever the right answer?

Rarely, and almost never for the reason given. A rewrite replaces an imperfect system you understand with an unknown one, and the new one reproduces a good share of the old one's oddities because business constraints created them, not carelessness.

Replacing function by function, with both systems live during the switch, costs more on paper and much less in reality.

How do I get the numbers if nobody tracks them?

From the last five change requests. Compare time actually spent with what the same change would have taken on a healthy part of the system.

That is not an estimate, it is money already invoiced this year, and it is sitting in tickets and timesheets you already have. An afternoon of digging produces a defensible figure.

Recognize your situation? 20 minutes is enough to tell.

Describe what is happening on your side, and where your data and customers are. You get a plain answer on whether we can help, and a better route if there is one.