When one person is your only backup, and nobody put it on the risk register
Every small IT setup has one person who alone can fix a particular thing. That is not a staffing problem but an availability risk, and it shrinks in weeks rather than by hiring.
By Laurent TulpanPublished 9 min read
Every organization running technology at low headcount has one. The person who is the only one who can restart that thing, or who knows why that particular setting must never be changed.
It is rarely written on any risk register, for a simple reason: nothing is going wrong. The system works, the person is there, and the arrangement has held for years.
It is also the exposure that turns an ordinary incident into a bad week.
Why it forms, and why it is nobody’s fault
The knowledge concentrates by default, not by neglect.
Somebody builds a system under delivery pressure. They know how it works because they built it, so writing that down produces no value they can see, and nobody asks. Over years the system grows, the operating knowledge grows with it, and it lives in one head because there was never a moment where it had to live anywhere else.
Framing this as a documentation failure by an individual is both unfair and counterproductive: it guarantees the conversation goes badly and nothing gets fixed. In our experience, the expert is the person most conscious of the problem. Being the only one who can fix something means never truly being off duty, and it is a heavier burden than it looks from outside.
Pricing it, so it competes with everything else
An unpriced risk loses every budget argument. This one prices out with three questions.
What stops if this person is unavailable for three weeks? Not unavailable forever, which nobody plans for. Three weeks: a leg in plaster, a family emergency, a resignation with notice. If the honest answer is that a category of work stops, you have a number, because a stopped category of work has a cost per week.
What is the longest procedure only they can perform? Usually a restore, a deployment, or a recovery after a specific kind of failure. Time it, then ask what happens if that procedure is needed while they are away.
What is the current recovery time, and what would it be without them? The gap between those two numbers is the risk, expressed in hours of downtime, which multiplies by the cost of an hour of downtime. That gives the figure that belongs in a budget table, in the same unit as the rest of the technical debt argument.
What actually reduces it, in order of cost
One. Have somebody else run the procedure. This is the highest-value exercise available, it costs an afternoon, and almost nobody does it.
The second person performs the critical procedure alone, following whatever notes exist, while the expert watches and does not touch the keyboard. Not a demonstration by the expert: an attempt by the other person.
Everything missing surfaces inside an hour, and there is always something missing. Usually a credential nobody wrote down, or a step so habitual to the expert that it never occurred to them to record it. On one engagement the missing piece was a configuration value that existed only in the head of the person who had set it three years earlier.
Two. Write the runbook from that attempt, not before it. A runbook written by the expert describes what the expert remembers. A runbook written from a failed attempt by somebody else describes what a person actually needs. The second is worth several times the first, and it takes less effort because the gaps have already been found for you.
Three. Put the credentials somewhere the organization controls. Not for lack of trust: because an account in a personal name, or a hosting contract signed by an individual, is a dependency that only reveals itself at the worst possible moment. This one is administrative rather than technical, and it is often the fastest item on the list.
Four. Rehearse once, on a schedule. Once a year, the second person performs the procedure again. Systems change, and a runbook that has not been exercised since it was written is an assumption. This is the same discipline as testing that a backup restores rather than that it ran, and it fails for the same reason when skipped.
Five. Hire, if the volume of work justifies it. Last, deliberately. A second person only reduces the risk if the procedures exist and have been rehearsed. Hiring first produces two people who depend on the same undocumented knowledge, one of whom happens to hold it.
The variant that hides better: the external single point of failure
The same exposure exists outside your organization, and it attracts less attention because it looks like a contract rather than a person.
If one vendor is the only party who can deploy your system, holds the hosting account, and has never transferred the operating documentation, you have the same risk with an invoice attached. It is worse in one respect: you cannot ask a vendor to train their own replacement out of goodwill, and by the time you want to, your negotiating position is gone.
That is why the transfer of source code, credentials and operating documentation belongs in the contract at signature, when you can still walk away. We treat it as a deliverable rather than a courtesy, which is also what made a sovereign program workable when the institution required control of everything.
The conversation to have this quarter
It fits in one meeting and needs no preparation beyond willingness.
Name the two or three procedures that only one person can perform. Not systems, procedures: what has to be done, by hand, when something specific happens.
For each, schedule the attempt by somebody else, with a date. A risk with a date attached gets addressed; a risk on a list gets re-read.
And say out loud that the point is not to reduce anyone’s importance. The expert stays the expert. What changes is that the organization stops depending on their calendar, which is a benefit to them before it is a benefit to anyone else.
Further reading
- Making the case for technical debt: the same risk, in a budget unit
- Your backups run. Do they restore?: the procedure nobody has rehearsed
- Handing over the code and the team: transfer treated as a deliverable
- Who owns the code your contractor writes: the external version of the same exposure
- Post-launch support: what a light watch covers, and what it does not
The questions this raises, answered straight.
Is key person risk really a technology problem?
It shows up as one and it is created by how the work was organized.
A system that only one person can operate is usually a system whose operating knowledge was never written down, because the person who built it never needed to write it down. That is not a character flaw, it is the default outcome when nobody is asked for documentation and delivery pressure is real.
Can we solve it by hiring a second engineer?
Sometimes, and it is the slowest and most expensive route.
A second person only reduces the risk if they can actually perform the critical procedures, which requires those procedures to exist and to have been rehearsed. Hiring without doing that first produces two people who both depend on the same undocumented knowledge, one of whom happens to know it.
What is the cheapest thing that reduces it?
A runbook written by somebody else. Have a second person perform the critical procedure, alone, following the notes, while the expert watches without touching the keyboard.
Everything missing surfaces within the hour, and there is always something missing. That exercise costs an afternoon and it converts an assumption into a verified fact.
How do we raise this without making the person feel accused?
By framing it as what it is: a risk the organization carries, not a failure the individual caused.
In practice the expert is usually the person most aware of the problem, and often relieved that somebody finally wants to address it, because being the only one who can fix something means never being genuinely off duty.
Recognize your situation? 20 minutes is enough to tell.
Describe what is happening on your side, and where your data and customers are. You get a plain answer on whether we can help, and a better route if there is one.

