All notes

A 401 and a blurry receipt.

A pipeline that never throws can hide an outage inside the review queue. Keep the never-throw design and make every failure carry its reason.

A finance team photographs receipts all day. A model reads each one and pulls out the merchant, the date and the total. Anything it is unsure of goes to a review queue, where a person checks it by hand. That queue is the safety net. It is also the most expensive part of the system, because it runs on people.

I like these pipelines to never throw. One bad photo should not stop a batch of five hundred good ones. So every document produces a result. A failure is folded into that result as low confidence. The batch finishes. The person sees what needs a person.

Then the question: what happens when this fails? Not when a photo is bad. When the system is.

The morning the key was revoked

Someone rotates a credential and the extraction service starts answering with a . Every call fails. The pipeline does exactly what it was built to do. It never throws. It marks each receipt low confidence and moves on.

The review queue fills with receipts that are perfectly sharp. The reviewers work faster. Nobody is paged, because nothing crashed. From the outside the outage looks like a busy day.

A blurry photo and a revoked key arrive at the same place, wearing the same label. Try it below: break the pipeline both ways, then turn reasons on and do it again.

Try it

Break the pipeline

Receipts flow into extraction. Revoke the key or feed a blurry photo, then turn reasons on and do it again.

InboxExtractSort

Filed

0

Review

0

Held

0

On callQuiet

Nobody has been paged.

What the reviewer sees

  • Nobody is waiting on a person

A simulation of the pattern, not a real system.

Keep never-throw, add a reason

The fix is not to start throwing again. The never-throw design is right. What it lost was one piece of information: whose problem is this?

A failure is either about the document or about us. A blurry photo, a torn receipt or the wrong kind of file is about the document. A person can fix it, so it belongs in the queue. A rejected key, a or a service that timed out is about us. No reviewer can fix a 401 by looking harder at a receipt.

Low confidence is a statement about the document. A revoked key is a statement about us.

So every failure still comes back as a result, never an exception. It just carries a reason alongside the confidence. The router reads the reason before it reads the confidence:

  • a document reason goes to the review queue, as before
  • an operational reason raises an alert and holds the document for a retry
  • the queue only fills with work a person can actually do

When the key comes back, the held documents go through again and most of them file themselves. Nobody had to review a receipt that was never the problem.

How to tell if you have this

You probably do if your pipeline folds errors into confidence. Look for these signs:

  • The review queue grows faster than the documents coming in.
  • Reviewers say the items look fine and approve them unchanged.
  • Queue spikes line up with deploys, credential changes or a provider's status page.
  • The alert history is quiet on the days the queue was loudest.

What happens when this fails?

The reason can be wrong too. I plan for three cases:

  • An error nobody named. When the pipeline cannot tell whose problem it is, I treat it as ours. It pages. A person in the queue cannot fix an unknown error. An engineer can.
  • Alert noise. Paging on every failed call trains people to ignore the pager. I alert on the rate of operational failures, not on each one.
  • Held work that never comes back. Holding documents is only safe if something re-runs them. I count what is held and alert if the count stops falling after the fix.

This does not make review smaller on a normal day. It keeps review for what it was built for: documents a person needs to read.

Next

Here is what I would do next

  • Find every place your pipeline turns an exception into low confidence.
  • Add a reason to that result, with two families: the document or the system.
  • Route system reasons to an alert and a hold, never to the review queue.
  • Watch the queue against the intake for a week and look for spikes with no cause.

Have a workflow like this?