What Retrying Costs — Not the Side Calling, the Side Receiving

7 min
GoalWriting about retries is mostly written from the caller's side. This piece counts from the receiving side. Four callers got through a 700-millisecond outage two different ways, and both eventually succeeded. What differed is how many times the gateway was hit meanwhile — 64 and 25.

At the end of Orders with the Till Gone we wrote this.

There is no automatic retry. The harness reconnected explicitly. A real app would try periodically in the background, and that period and backoff were not decided here.

This piece decides that period. And while deciding it, measures what happens when you decide it badly.

This is not a specific business. We constructed the situation of a payment gateway going down briefly. The code, screens and logs actually ran.

Counting from the receiving side

Most retry writing is the caller's story — how do I get my request through.

This sample's server stands on the other side. It counts how many attempts arrived, and how many of those landed inside the same 200 milliseconds.

/// That second number is the one that matters. A gateway that is briefly down
/// does not care that you tried again. It cares how many of you tried again at
/// the same instant, because that is what keeps it down.

A briefly-dead gateway is not interested in the fact that you tried again. What it cares about is how many of you tried again at the same instant. That is what keeps it dead.

What gets written when you are in a hurry

/// Try again immediately. This is what gets written when the retry is added in
/// a hurry, and it is the strategy that turns a short outage into a long one.
class NoBackoff extends Backoff {
  @override
  int waitBefore(int n) => n == 1 ? 0 : 20;
}

Retry immediately — 64 attempts, 18 of them inside one 200ms window. The list shows every attempt with the wait that preceded it

RETRY IMMEDIATELY: 4 callers, 64 attempts total, 4 accepted, peak 18
  caller 1 attempt 1: straight away -> DOWN at 12ms
  caller 1 attempt 2: after 20ms -> DOWN at 58ms
  ...
  caller 1 attempt 16: after 20ms -> ACCEPTED at 708ms

It got through. It survived a 700-millisecond outage and all four were charged.

And the gateway was hit 64 times in the meantime.

Double each time, then shake it

@override
int waitBefore(int n) {
  if (n == 1) return 0;
  final base = 60 * (1 << (n - 2)); // 60, 120, 240, 480 ...
  // Full jitter: anywhere in [0, base]. Spreading matters more than being
  // punctual — nobody is waiting on an exact millisecond here.
  return _rng.nextInt(base + 1);
}

Exponential backoff with jitter — 25 attempts for the same outage, worst window 13. The line below sets the two strategies' cost side by side

EXPONENTIAL + JITTER: 4 callers, 25 attempts total, 4 accepted, peak 13
  caller 1 attempt 1: straight away -> DOWN at 13ms
  caller 1 attempt 2: after 53ms  -> DOWN at 70ms
  caller 1 attempt 3: after 87ms  -> DOWN at 159ms
  caller 1 attempt 4: after 143ms -> DOWN at 305ms
  caller 1 attempt 5: after 389ms -> DOWN at 697ms
  caller 1 attempt 6: after 449ms -> ACCEPTED at 1150ms

Same outage, same people, same result. All four were charged.

Attempts went from 64 to 25. Times 2.6 fewer.

Doubling and shaking are different jobs

This is what this piece wants to say.

Everybody knows exponential backoff — 60, 120, 240, 480. Jitter often gets left out. But with one caller jitter makes no difference, and with several, jitter is everything.

The reason is simple. If four failed at the same instant, then without jitter four come back at the same instant. Four at 60ms, four at 180ms, four at 420ms. The gateway keeps taking the same spike that just knocked it over.

Look at the log above: the waits are 53, 87, 143, 389, 449 — none of them 60, 120, 240 or 480. They are shaken values.

The verification confirms that.

# Backing off must actually cost the gateway less, otherwise the whole point
# is lost: everybody still gets through, but fewer of them arrive at once.
assert polite["attempts"] < hammer["attempts"], (polite["attempts"], hammer["attempts"])
assert polite["peak"] < hammer["peak"], (polite["peak"], hammer["peak"])

It also checks that both got through

Without this the check lies. The easiest way for backoff to reduce load is to give up.

echo "$HAMMER" | grep -q "4 accepted" || { echo "   hammering did not get everybody through"; exit 1; }
echo "$POLITE" | grep -q "4 accepted" || { echo "   backoff did not get everybody through"; exit 1; }

Only then does it compare load.

   attempts 64 -> 25
   peak in one 200ms window 18 -> 13
   same outage, both got through · 64 -> 25 attempts · spike 18 -> 13

It looks at both the total and the peak. With only the total you cannot tell it apart from "it just went slower"; the peak has to come down for the gateway to actually have an easier time.

The screen nearly lied in green

This came out of building it.

At first there was a verdict sentence at the bottom of the screen. It is a warning — "the gateway was hit harder than the number of callers" — and it rendered in green.

The colour was set to change with state, but that sentence's colour was hard-coded as #166534 in the screen definition. The log was accurate; only the pixels were wrong.

The same kind of defect that was caught in Put the Reason Next to the Answer came back. There too, an answer with zero grounding was green.

Fixing it, we removed the verdict altogether.

// The screen reports what was counted. It does not grade it — the
// comparison is the article's job, and a screen that grades tends to
// grade in a colour that disagrees with its own sentence.

The screen states what was counted; the article does the comparing. A screen that grades tends to disagree with itself in colour, and when it disagrees, the colour wins.

The range of this sample

The peak difference is not large (18 → 13). Total attempts fell by 2.6× while the peak fell by 1.4×. There is a reason — the stdio transport serialises requests, so the four never truly arrive at the same instant. It is the same constraint met in Told Yes Twice. On a real HTTP gateway this difference would be much larger, and this sample did not show that.

Load does not lengthen the outage. This gateway comes back after 700 milliseconds no matter what. In reality the harder you hammer the later it recovers, and that feedback is what makes this a real problem — and it is not in the model.

There is no circuit breaker. Stopping attempts entirely after repeated failure and probing once after a while is the next step, and it was not built.

Idempotency is already in another piece. For a retry to be safe the same request must be able to arrive twice, and that is what Orders with the Till Gone covered. Read only this piece and add a retry and you get double charges.

Four callers is all. A real spike is thousands. A different order of magnitude is a different story.

What is left

A retry is easy to add. One line will do.

And that one line lengthens somebody else's outage. What keeps a service that would have been down briefly down for a long time is usually the retry loop of whoever uses it.

Increasing how long you wait and waiting differently from each other are different jobs, and in a system with many callers the latter matters more.

For the price of one more line, it is cheap.

Practice task

Send the same booking twice in a row. Show that the second one does not create a second booking, and explain which side made sure of it.

Related articleWhat Retrying Costs — Not the Side Calling, the Side Receiving