Told Yes Twice — What One await Decides

7 min
GoalTwo people take the same slot at the same moment. Described in words, the two handlers do exactly the same thing — look whether it is free, and if it is, take it. One of them confirms to both people; the other does not. The difference is one await in the middle.

This is not a specific business. We constructed the shape exactly as it works in a small place that takes bookings — a salon, a study room, a shared kitchen. The names and times are invented; the code, screens and logs actually ran.

The worst failure in a booking tool is not that the booking fails.

It is sending "you're booked" to both people. If it fails you try again; if both are told yes, what comes next is somebody sorting it out on the phone.

Two handlers, the same description

This sample has two handlers that take a slot. Described in words they are identical — look whether it is free, and if it is, take it.

One is this.

// 1. Look.
final free = slot.takenBy == null;

// 2. Anything at all in here — a database round trip, a payment check, a
//    log write, a lookup of the customer's name — hands control to
//    whatever else was in flight.
await Future<void>.delayed(Duration.zero);

// 3. Write, using what was true in step 1.
if (free) { slot.takenBy = who; ... }

The other is this.

if (slot.takenBy != null) {
  return '$at is already taken by ${slot.takenBy}';
}
slot.takenBy = who;
slot.confirmations.add(who);

await Future<void>.delayed(Duration.zero); // the slow part, after the claim

The same await, in a different position. In the first it sits between the look and the write; in the second it sits after the write.

And that position produces the double booking.

Make it actually happen

unsafe -> Smith:  "CONFIRMED 10:30 for Smith"
unsafe -> Jones: "CONFIRMED 10:30 for Jones"
unsafe result: 1 slot confirmed to more than one person

After the race — two people were told yes for the one 10:30 slot. TOLD YES TWICE reads 1, and the row still carries both names

The screen shows exactly the real-world symptom. The slot has one owner, Jones, and two people were told they had it.

This is the shape it takes in a real shop. The book says one; two turn up at the door.

So the server remembers everybody it confirmed to.

/// Everybody who has ever been told they got this slot. In a correct
/// booking system this never has more than one name in it, which is exactly
/// why it is worth keeping.
final confirmations = <String>[];

Correct means always a one-name list. A list that must always have one entry is exactly the list worth counting.

And then make it not happen

Same two people, same slot, only the gap closed.

--- cleared, same two people, same slot, gap closed ---
safe -> Smith:  "CONFIRMED 10:30 for Smith"
safe -> Jones: "10:30 is already taken by Smith"
safe result: nobody was told yes twice

The same race with the server taking one at a time — 10:30 belongs to Smith alone and TOLD YES TWICE is 0. The foot line reads "One slot, one yes"

The refusal is not "that failed" but "Smith has it". What the second person needs to know is not that they failed but that the slot is gone, so that they pick another time.

Check only the good case and it passes

This sample's verification is unusual. It checks that the bug happened.

unsafe = s.call("book.stampede", {"at": "10:30", "first": "Smith", "second": "Jones", "mode": "unsafe"})
assert unsafe["doubleCount"] == 1, unsafe      # the race must actually happen, or this run proves nothing
safe = s.call("book.stampede", {"at": "10:30", "first": "Smith", "second": "Jones", "mode": "safe"})
assert safe["doubleCount"] == 0, safe

The reason is simple. If a booking test only checks "exactly one was confirmed", it passes against a server with the bug still in it. If the race does not actually happen, the requests are handled in order and the result looks fine.

So the order is this.

  1. With the unsafe one, confirm the race actually happened
  2. With the safe one, confirm it does not happen in the same race

Without the first, the second means nothing.

There is one more. If the gap disappears, it fails.

# The two handlers must stay near-identical apart from where the await sits.
# If somebody "fixes" the unsafe one, this sample stops teaching anything.
grep -q "await Future<void>.delayed(Duration.zero);" booking_server/bin/server.dart \
  || { echo "   the gap in the unsafe handler is gone — nothing is demonstrated"; exit 1; }

Building it, the race did not happen

On the first attempt the bug would not appear.

The harness fired two tool calls at once and the result was correct.

take_unsafe -> Smith:  "CONFIRMED 10:30 for Smith"
take_unsafe -> Jones: "10:30 is already taken by Smith"
take_unsafe result: nobody was told yes twice

The transport layer was serialising. stdio reads one request line, runs its handler to completion, and then reads the next. There is no interval in which to overlap.

Two ways to go. One is to shrug — "it does not happen on this transport, so there is no problem". The other is to make the overlap and check.

We took the second. Switch to HTTP or SSE and that overlap arrives without anybody making it. What a transport happens to be protecting is not protected.

// The overlap is made here, inside the server, and that is on purpose.
// ...
// On an HTTP or SSE transport the overlap arrives by itself and nobody
// has to arrange it. So this tool arranges what that transport would have
// handed the server anyway.

This is a first of its kind in this series. Every defect in the earlier pieces was in my code; this one is a case of the environment hiding the defect. And an environment that hides things can change at any time.

The range of this sample

I made the overlap. Two handlers are started concurrently inside the server. It is not a race arising from genuinely concurrent requests coming in over the transport. That the race does not reproduce over stdio is what this piece found out, so this sample shows "what it would have been over HTTP". It was not confirmed on an HTTP transport.

Only the race inside a single process is covered. Grow to two servers and this remedy is void entirely. Then you need a unique constraint or a lock in the database, and guaranteeing "nothing between the look and the write" outside the process is much harder. Not covered.

No cancellation and no waiting list. Releasing a slot you took, handing a released slot to somebody who was waiting — neither exists. Half of a real booking tool is over there.

No payment. Attach payment and a state appears where the slot is held but the payment failed, and how long to hold that state becomes a new problem.

What is left

The code in this piece is under ten lines. And those ten lines cannot be told apart in words.

"Look whether it is free and if it is, take it" — both handlers are described that way. Both would pass code review, and both would be written the same way in a specification.

The only place it shows is execution. So this kind of defect is not found by reading; it is found by making it happen.

And once you have made it happen, that becomes the check.

Practice task

Send the same booking twice in a row. Show that the second one does not create a second booking, and explain which side made sure of it.

Related articleTold Yes Twice — What One await Decides