What happens when this fails?
The question I ask of every system before it goes live. Break an invoice run four ways and watch what a good answer does.
4 min read
Every system I build gets the same question before it goes live: what happens when this fails?
It is not a trick question. A demo shows the path where everything works. The people who use the system meet the other paths first. A supplier rings about an invoice nobody paid. A finance lead finds the same bill paid twice. Nobody saw an error, because nothing raised one.
Break it yourself
Here is a small invoice run. An invoice arrives, somebody approves it, then it is paid. The top lane was built with the question asked. The bottom lane was built for the demo. Pick how it breaks and where, then send an invoice through both.
Try it
Break the invoice run
An invoice arrives, gets approved, then paid. Pick how it breaks and where, then send one through. The top lane was built for failure. The bottom one was built for the demo.
How it breaks
Where it breaks
Built for failure
Owner Quiet
Send an invoice to watch both lanes
- Landed
- 0
- Person
- 0
- Lost
- 0
Built for the demo
- Landed
- 0
- Person
- 0
- Lost
- 0
Send a clean one first. Both lanes look the same. That is the trouble with demos: the difference only shows when something breaks. By then the system is live.
Four moves and one rule
Break it a few ways and the top lane keeps making the same four moves.
- Retry with Waiting a little longer after each failed try, so a struggling service gets room to recover instead of a flood of calls. when a service is down for a moment. The wait grows each time so the retries do not make it worse.
- Fall back to the last good copy when the missing piece is reference data, like the approval rules. The run carries on and says so.
- Route to a person with the reason when only a person can fix it. A missing total or a new bank account is not a technical fault. The person gets the invoice and the one sentence that explains it.
- Tell the owner whenever the system did something unusual on its own. Nobody should learn about a fallback from an audit.
And one rule: money never runs on an old copy. When the bank is not answering, the payment waits. It does not guess.
The bottom lane makes none of these moves. Worse, it says nothing. Its log is empty or it says paid. Watch the tally: the top lane loses nothing and asks a person only when a person is needed.
A good answer names the step that stops and the person who decides.
Ask it while you scope
The question belongs in scoping, not the week before handover. The answer changes the design. It decides where a check goes, which step can stop the run and where a person still decides.
It also changes what done means. Here is one requirement, written before and after someone asked.
Before
If the payment fails, log an error.
After
If the bank does not answer, retry with a growing wait, hold the payment and tell the finance owner. Pay once it answers. Never pay twice for one invoice.
The second version is longer. It is also the only one a team can test, support and hand over.
What a useful answer looks like
A useful answer is specific. For each failure case I want four things written down:
What the person using it sees
Not a spinner that never ends. A status that says what happened and what comes next.
Which step stops the run
And the state everything else is left in, so a retry never pays twice.
Who is told and who decides
A named role, not a shared inbox. A backup for when that person is away.
What the team owns after handover
The alerts, the queue a person works through and the fallback copies that need to stay fresh.
If nobody can answer, the system is not ready. That is a finding, not a failure.
The duplicate is worth a second look. The top lane sees a second approval or a second payment request and does nothing twice. That property has a name, Safe to run twice: the second run changes nothing.. It is what makes a retry safe. Without it, every retry is a fresh chance to pay twice.
What happens when this fails?
The question has failure modes of its own. I watch for three:
- Answers that stay on paper. A failure plan nobody has tried is a guess. Break the system on purpose before launch, the way the drill above does.
- Everything goes to a person. Routing to a person is the safe move, which makes it the lazy one. If the queue fills with work a machine could retry, the design is wrong. A 401 and a blurry receiptA 401 and a blurry receipt. A pipeline that never throws can hide an outage inside the review queue. Keep the never-throw design and make every failure carry its reason. shows what that looks like.
- Alerts nobody reads. Telling the owner only works if the owner can act. Alert on what needs a decision, not on every retry.
I also say whether an example is a demo, a pilot or a production system. The drill above is a simulation. A useful experiment is not proof that something is ready for your team.
When something does fail, I explain the impact, the response and the lesson. Then I end the way I always do.
Next
Here is what I would do next
- Pick one workflow that moves money or promises something to a customer.
- Write the four answers for every step: what the person sees, what stops, who decides and what the team owns.
- Break it on purpose before it goes live: a service down, bad input, a person away and a duplicate.
- Check that every break leaves a log line and reaches someone who can act.