All notes

The dumbest check I trust most.

One wrong fact in an AI draft and I throw the whole draft away. Retrying until it passes ships wrong numbers that look right.

A team wants a model to write the messages it sends every day. Invoice reminders, clinic notices, the short notes nobody enjoys writing. The model is good at tone. The team is happy with the demo.

Then I ask the question I ask of every system: what happens when this fails? Not when the tone is flat. When the model writes the wrong due date and the message goes out anyway. A bland message costs nothing. A confident wrong date costs a phone call, an apology and some trust.

The check

The check I trust most is almost embarrassing. Every fact in the source record has to appear in the draft, exactly as written. The date, the amount and the reference. If one is missing, the whole draft is thrown away and the message goes out as the copy we already had: the plain template the team used before the model arrived, filled from the same record.

There is no partial fix. There is no second attempt. It is all or nothing. Either way the customer gets a correct message.

Below is the same idea as a toy. Let the model slip and watch both lanes.

Try it

Let the model slip

One slipped fact goes to both lanes. The left lane throws the draft away. The right lane asks again until its check says yes.

Source

Due date
14 March 2027
Amount
$1,280.00
Invoice
INV-20417

AI draft

Accounts · Harbour Print

To Sam

Payment reminder

Hi Sam, a quick reminder that invoice INV-20417 for $1,280.00 is due on 14 March 2027. Reply here if anything looks off.

All or nothing

Every source fact must appear in the draft exactly.

  • 14 March 2027found
  • $1,280.00found
  • INV-20417found

Press a button above to run it

Wrong facts shipped

0

0 drafts discarded

Retry until it passes

A second model grades the draft. Ask again until it says yes.

Press a button above to run it

Wrong facts shipped

0

0 runs

Press the slip button to send one bad fact through both lanes.

A simulation of the pattern, not a real system.

Why I never ask again

The popular fix goes the other way. When a check fails, send the error back to the model and ask again. Guardrail tools are built around that . It feels responsible. It is the opposite choice.

A retry loop needs a check it can pass. An exact match leaves the model nothing to negotiate with, so the loops that retry lean on softer checks. Another model grades the draft or a rule checks its shape. Does this look like a date? Is the amount formatted? Each retry is a fresh chance for the model to produce something that clears the bar.

A retry loop stops at the first draft that passes. Not the first draft that is true.

That is how a loop converges on a wrong number that looks right. The first attempt says May and gets caught. The second says the fifteenth of March, which reads perfectly and is one day off. The grader nods. It ships.

Throwing the draft away is boring on purpose. The fallback copy was already approved, so the worst case is a plainer message, never a wrong one.

What the check costs

The guard is strict, so it throws away drafts that were nearly right. A missing currency code or a date written a different way is enough. The reader gets the plain copy that day. I count that as the price of the check, not a bug in it.

It is also a signal. I track how often drafts are discarded. If the rate climbs after a model update or a prompt change, something moved. I would rather learn that from a counter than from a customer.

What happens when this fails?

The check has limits. I say them out loud:

  • It only checks facts that are in the source. If the model invents an extra claim, a late fee or a new deadline, presence alone will not catch it. That needs its own rule or a template narrow enough to leave no room.
  • It trusts the source. A wrong record produces a wrong message on either path.
  • It needs a fallback worth sending. If the team has no approved copy, the guard has nothing to fall back to. A person has to decide before anything goes out.

That last point is where I keep a person in the loop. The model drafts. The check decides whether the draft is safe to send. A person owns the fallback copy and reviews the discard rate. Nobody reviews every message. Nobody has to.

Next

Here is what I would do next

  • Pick the facts that must never be wrong in your messages: dates, amounts, references, names.
  • Write the plain fallback copy first and get it approved before the model writes a word.
  • Add the presence check, all or nothing, with no retry.
  • Count the discards every week and treat a jump as a finding.

Have a workflow like this?