Skip to content
Delivery practice

Technical Debt: What It Is, How to Measure It, How to Pay It Down

By Sebastiaan Jansen · · 8 min read

Technical debt is the extra cost a team pays later because code, design, or process shipped in a form that is harder to change than it should be. Old code doesn't automatically count, and neither does every bug. Real debt is a trade you can name and measure, and then pay down on purpose.

What technical debt is

Ward Cunningham coined the phrase as a money metaphor. You borrow speed now, then pay interest every time someone works around the shortcut: slower changes, more defects, longer reviews, and estimates that slip.

The interest is the part to watch. A shortcut nobody touches again is a one-time cost. A shortcut on the path of every feature is a loan with a weekly payment. Teams get stuck when they track the principal ("that module is messy") and ignore the interest ("every ticket there takes three extra days").

Debt can be deliberate or accidental, and reckless or prudent. Martin Fowler's explanation of the metaphor separates a conscious shortcut with a payback plan from a design the team never examined. Both cost time, but only the first was a decision.

What does not count

A missing feature is not debt, and neither is a product decision you dislike. An unfashionable library doesn't count if it is maintained and still easy to change. If a fix would not make a named future change cheaper, safer, or faster, it is a preference and doesn't belong on the schedule.

A technical debt definition your team can share

You want a definition that holds up in planning: a known gap between the system you have and the system that would make a stated class of changes cheap and safe. The gap has an owner, a place in the code or the process, and a cost you can see in delivery.

Write it down where tickets get written. Otherwise tech debt turns into a label for work nobody wants to explain. Require four fields:

  • Where it lives: a module, service, pipeline, or process step.
  • Which change it makes more expensive.
  • How you noticed: cycle time, defects, review comments, or failed deploys.
  • What "paid" looks like, in a check a reviewer can apply.

That last field is a definition of done. Nobody can pay off "Clean this up." Someone can pay off "Split billing so a price-rule change touches one package and ships behind the existing flag."

How to measure technical debt

No single score means the same thing in two companies. Measure the interest, inventory the principal, and leave out any dollar figure you cannot defend.

Inventory the principal

Keep the register as normal tickets rather than a side spreadsheet. Each item names the gap, the area, and the payback. Review it with the backlog, and archive or rewrite items that won't help anyone in the next two quarters. If the register only ever grows, it isn't a plan.

Measure the interest

You already have the interest in your delivery data.

Signal What it suggests How to read it
Cycle time in one area vs the rest That area costs more to change Same ticket type, same team
Change-fail rate after touches there The design is brittle Pair with defects, not with blame
Review time and revert rate The code is hard to reason about Rising review time is interest
Extra manual release steps Process debt, not only code Count gates on the last five releases
Cycle share spent on rework Interest is crowding out scope Rework means repeating a change

Use percentiles instead of averages. A p50 of two days and a p85 of eleven means a few areas cause most of the pain, and a forecast from that history shows a wide date range. When dates are unpredictable, the cause is often unpaid interest rather than bad estimation.

Linter scores and "hours of debt" plugins are inputs. Complexity only matters if it sits on next quarter's roadmap, so weight the list by planned work. A messy module you will not change is cheap; one in the next release is expensive. Coverage, dependency age, and TODO counts only point to candidates. Confirm them with cycle time and with the people who last shipped there.

A worked example

Northwind hard-codes three shipping rules inside OrderService.price() so a partner launch can make Friday. The ticket says to revisit them before a fourth carrier.

By August there are six carriers, and every rule change edits that method. Across seven tickets, checkout cycle time is a p50 of four days. The rest of the storefront, with the same team and ticket size, is a p50 of a day and a half. Two tickets failed review once on a tax edge case. That works out to about two and a half extra days per change, plus rework.

One paydown ticket covers it: extract a data-driven ShippingRules table, keep the old method as a wrapper for one release, and add a characterization test for the six carriers. Done means the wrapper is gone, those carriers pass the same cases, and adding a carrier is a data change.

They schedule the extract with the seventh carrier, because that carrier pays for the work, and they don't rewrite checkout. The next two carrier tickets close in under two days. The hours are an illustration, not a benchmark. The method is what carries over: name the shortcut, count the extra time on real tickets, and pay it down on the next change that hits the same code.

Managing technical debt while you keep shipping

The hard part is deciding what not to fix. Editing the code is easier.

Set a visible budget

Reserve a stable share of each cycle. Fifteen to twenty percent of capacity is a common heuristic, though nobody enforces it. With a fixed share, paydown doesn't have to beat a customer request every week. If one area carries almost all the interest, spend the whole share there until its cycle time matches the rest. Ten tiny cleanups can feel fair and still pay off almost nothing.

Put the share on the board before you pull features. Whatever time is left at the end of the cycle usually rounds to zero.

Rank by interest, then by roadmap

Ask how often you will touch this next quarter and how many extra days each touch costs. Multiply expected changes by extra days. That ranks better than "how ugly is the file." A scary module off the roadmap ranks below a dull script that every release runs by hand.

Attach the top item to a feature already committed. If the paydown has no line in a product requirements document, it is optional, and optional cleanup loses. A constraint such as "the next carrier must not edit order pricing" funds the extract. When the gap spans several tickets, a requirements traceability matrix shows which requirement still depends on the wrapper, so you don't call the release done too early.

Keep shipping

A cycle of nothing but paydown usually fails. There's no user-facing check, a catch-up cycle follows, and the same register is back next quarter. Pay the loan down alongside work already in flight. Reducing the debt without stopping features applies the same idea at sprint level.

Release management belongs in the loop too. A manual checklist you repeat because nobody trusts the pipeline is process debt, so automate one line of it each release.

Catch new items in the ticket

Most new debt arrives when the story says "make it work for Friday" and never names the shortcut. Before merge, ask what will be harder next time and log a specific answer. Give the paydown its own ticket with a clear boundary. Drive-by refactors inside a feature branch can look done while the interest stays flat. Stay inside the ticket, whether a person or an agent writes the change.

How to pay an item down

If you skip a step, the cost comes back under a new file name.

Start by characterizing current behavior, awkward edges included. The tests don't endorse that behavior; they stop the extract from changing what customers already rely on. If you can't characterize it, start with a log, a trace, or a fixture.

Then make the change smaller. Ship a wrapper and delete it in the same item. Two small releases beat one long branch, and a wrapper that lives for a year is a second loan.

After release, read the next tickets in that area. Cycle time and change-fail rate should move. If they don't, the bottleneck is probably review, data setup, or a manual gate. Record why you took the shortcut and why this paydown was enough.

Rewrite only when you cannot characterize the behavior, the module is on the roadmap, and a parallel path costs less than an extract. Disliking the code isn't a reason. Run both paths until the new one matches the cases you care about, then delete the old path in the same quarter.

Where the register should live

Debt that only lives on a slide doesn't get paid. Tickets do, if they have an owner, a stage, and a link to the feature they unblock. A note stays a candidate until someone confirms the interest. A recurring cycle can reserve the paydown share and roll unfinished work forward under review, so a half-done extract doesn't vanish when the cycle ends.

If you build the roadmap from tickets, the paydown sits next to the feature it serves. The carrier work includes the extract, and without it the date range stays wide. Cumulative flow and cycle-time percentiles show whether the share is working. If the debt column never shrinks, either the share is too small or the tickets are too vague.

Falrow runs that loop in one workspace. Chat, call notes, and production errors become tickets with custom stages, and a cycle can reserve a paydown share with reviewed rollover. The ticket workflow holds the register, and the delivery loop is how a request becomes a reviewed change. Another tracker works too, as long as the register, the share, and the interest check stay in one place.

FAQ

What is a simple way to define this?

It is a known gap between the system you have and the system that would make a specific change cheap and safe. You pay extra time whenever you cross it. If you can't name the change that would get cheaper, you don't have an item yet.

Is messy or old code always debt?

No. It only charges interest while you still have to change it. A module you won't touch is cheap to carry, and a small function on the critical path can cost more than a large file off the roadmap. Rank by expected touches and extra days.

How much of a cycle should go to paydown?

Many teams reserve a steady share, often about fifteen to twenty percent of capacity, and spend it on the highest-interest items. That is a common choice rather than a standard. Raise it when one area is clearly slower, and lower it once later tickets there match comparable work.

Can a team get to zero?

You can pay individual items down completely. A codebase that is still changing will keep producing new shortcuts, and some of them are worth taking. Aim for a register you trust, a capacity share that actually gets spent, and cycle time that doesn't creep up where you're about to ship.

More on delivery practice