Operations
How to Find the Single Points of Failure in a Small Team
A single point of failure in a business is any person or system whose absence stops work. How to find yours, rank them, and fix the ones that matter.

A single point of failure in a business is any person, system or piece of knowledge whose absence stops work from moving. Not slows it. Stops it. In a team under fifty people there are usually between three and eight of these, and most of them have a name and sit twenty feet away.
Everybody half-knows who they are. What almost nobody has is the list, written down, ranked, with the cost of each one next to it. That gap between half-knowing and knowing is the whole problem, because a worry produces anxiety and a list produces decisions.
Why they form, and why nobody notices
Single points of failure are not a sign of bad management. They are the natural result of a business growing faster than it documents itself.
Somebody joins early and learns how the billing works because they built it. Three years later the billing is more complicated, they are still the only person who understands the exception cases, and nobody made that decision. It accumulated.
The reason it stays invisible is that a single point of failure produces no symptoms while the person is present. Everything works. The process is efficient, in fact, because one person doing all of something is faster than three people coordinating. You only find out on the Tuesday they are in hospital.
That is what makes this hard to prioritise. Every other operational problem announces itself continuously. This one is silent right up until it is a crisis, so it loses every week to things that are on fire now.
Four places to look
You do not need an audit to find them. You need four questions, asked of the right people.
Who is the only person who can approve this? Approval bottlenecks are the most common kind and the easiest to fix. Often the approver has not looked at the substance in a year and would happily delegate below a threshold if anybody asked.
Whose login is this? Shared credentials, personal accounts holding company data, the domain registrar in one person's name, the payment processor tied to a director's phone for two-factor. This category is unglamorous and produces the worst outages, because it is not just knowledge that leaves, it is access.
What happens when this goes wrong? The exception path is where dependency hides. Twelve people can run the happy path. One person knows what to do when the payment fails, the customer disputes it, and the finance system will not reverse the entry.
Who do people ask? Not the org chart. Watch where the questions actually go. There is usually one person who fields questions from three departments they do not sit in, and they are load-bearing in a way no document reflects.
Rank them before you fix anything
The instinct after finding six is to fix all six. Do not. Score each one on two axes and you will find that two matter and four can wait.
How often does the work run? A daily process with a single dependency is a live risk. An annual one is an inconvenience.
How long would recovery take? This is the axis people underweight. Some dependencies are recoverable in an afternoon by reading a document. Others involve knowledge that exists nowhere except in one head, and recovery means rediscovering it under pressure while customers wait.
Multiply the two. The top of that list is your actual risk register, and it is usually shorter than you feared and different from what you assumed.
The fixes are duller than you expect
Three things resolve most of them, and none of them are a project.
Write down the exception path, not the happy path. Sit with the person, ask them to talk through the last three times it went wrong, and record it verbatim. This takes an hour and removes more risk than any amount of process redesign.
Put a second pair of hands on it once. Not training. Have someone else actually do it, with the expert watching, one time. The gap between "I read the document" and "I have done it" is where recovery time lives.
Move access out of individuals and into the company. Shared vaults, company accounts, a second admin on everything that matters. This is the cheapest fix and the one most often skipped because it is nobody's job.
The concession
Some single points of failure are worth keeping, and pretending otherwise makes this exercise feel like a compliance chore.
Concentrating a specialised job in one capable person is often the right call in a small team. The cost of spreading genuine expertise across three people who each do it occasionally is real: they will each be slower, worse, and more likely to make errors than the one person who does it every day. Redundancy is not free.
The distinction is between concentration you have chosen and concentration you have drifted into. A deliberate one comes with a written fallback, even if the fallback is "we accept two weeks of disruption and here is who we would call". A drifted one comes with nothing, and the difference only shows up on the day it matters.
Seeing them all at once
The reason this stays a worry rather than a list is that no single person can see the whole business at the same time. You know your own dependencies. You do not know the ones in a team you do not sit with, and the people in that team do not know theirs, because from the inside it just looks like how the work gets done.
That is the case for drawing the whole thing out. abi. Clone builds a working model of your organisation covering departments, roles, systems and the processes running between them, and dependencies become visible as a shape rather than a suspicion. It is free, no card, and a first pass by hand takes about ten minutes.
But a list on paper counts. The map is a faster route to the list, not a substitute for having one. What matters is that the four questions get asked, the answers get written down, and the top two get fixed this month rather than after the Tuesday.
View more articles
Learn actionable strategies, proven workflows, and tips from experts to help your product thrive.


