The Tool That Refuses to Read — Keeping Judgment Outside a Clinic Flow

6 min
GoalIn 2026-04 we ran a piece about a doctor building a clinic flow tool. That piece stated in prose what the tool does not do. This time we drew it in code — this server has no tool that produces a judgement, and it throws if it tries to emit a word that sounds like a verdict.

Two things stated up front.

First, in April 2026 this magazine ran a piece about a doctor building a clinic flow tool. This is not a specific practice. We constructed the shape exactly as it actually works, and we do not name where. What follows was built to that shape, and the procedure and the numbers are invented.

Second, this is not medical advice and what we built is not a medical device. This series did no clinical validation and is not in a position to. What it deals with is one software boundary — what keeping judgment outside the tool looks like in code.

A line drawn in prose gets erased

The original had that line too. "Leave the diagnosis to the person; only organise the flow of the visit onto a screen."

Correct. The problem is that it existed only as a sentence.

A line that exists as a sentence gets erased. Specifically it gets erased like this: months later somebody says "a one-line summary would be handy" and adds a field. Into that field goes a string like broadly within range. Nobody notices the moment. That is the moment the tool started diagnosing.

So this time we drew it in code. Twice.

The first line — there is no such tool

tools offered: flow.state, flow.done, readings.list

Three. Show the routine, mark one step done, return the readings as they are. There is no tool that judges, classifies or interprets.

The verification starts from the tool names.

WORDS = ["diagnos", "assess", "judg", "risk", "abnormal", "recommend"]
for n in s.tool_names():
    assert not any(w in n.lower() for w in WORDS), n

This is an application of the principle confirmed in an earlier piece. The tool list is the surface, and a capability not on the list cannot be reached by a person pressing or a model choosing.

The second line — sound like a verdict and it throws

The list alone is not enough, because a judgement can slip inside the strings flow.state returns.

/// Words that would turn a report into a verdict. Anything this server is
/// about to say is checked against them.
///
/// A list of words is a crude guard and it is meant to be. It cannot stop a
/// determined author, but it does stop the ordinary way this line gets
/// crossed: someone adds a helpful-sounding summary field months later and
/// nobody notices that the tool started diagnosing.
static const forbidden = [
  'diagnos', 'likely', 'suggests', 'consistent with', 'probable',
  'abnormal', 'normal', 'healthy', 'concerning', 'severe', 'mild',
  'recommend', 'should take', 'prescribe',
];

static String _guard(String s) {
  final lower = s.toLowerCase();
  for (final w in forbidden) {
    if (lower.contains(w)) {
      throw StateError('clinic_server tried to emit a judgement word: "$w"');
    }
  }
  return s;
}

Every string this server emits passes through that function. Note that normal is on the list — the most harmless-looking word, and the one that most often crosses the line.

It is a crude guard. It will not stop somebody determined. But the way this line actually gets erased is not determination but inattention, and inattention is stopped by this much.

And the verification sweeps the whole run.

no judgement vocabulary in anything the server returned

What the tool actually does

Having drawn the line twice, look at what is left.

It shows the routine in order. Exactly the order the clinician wrote.

routine: intake* -> vitals* -> review -> exam -> plan -> note

The * marks what is finished. What appears at the top of the screen is what comes next — and even that is not advice.

// "Next" is position in a list the clinician wrote. It is not advice.
'next': _guard(next.label),

It is a position in a list. The distinction looks trivial, but "do this next" and "the next slot in the order you wrote is this" put responsibility in different places.

Today's routine — each step says what it needs before it can be ticked. Real render capture

Mark one step done and only the position moves.

After one step — NEXT IN THE LIST moves on. "Next" is a position in the list the clinician wrote, not advice

Readings are given as they are, with the reference range beside them.

readings: Blood pressure 148/92mmHg (clinic uses <130/80)
        | Heart rate 78bpm (clinic uses 60-100)
        | Temperature 36.8C (clinic uses 36.1-37.2)

<130/80 sits next to 148/92. But it does not say "high". That judgment is not this tool's.

There is a reason the range has to be there at all. With only a number, the reader fills in the verdict themselves. And that verdict is recorded nowhere. With the range beside it, at least what it was compared against stays on the screen.

The verification looks at that too.

readings = s.call("readings.list")["readings"]
assert len(readings) == 3 and all(r["usual"] for r in readings)

Readings — each number sits beside its usual range, and there is no verdict anywhere

The range of this sample

There is no doctor. And the routine and reference values above are invented. A value like clinic uses <130/80 is this sample's number, not any clinical guideline.

What this piece verified reaches one property of the software. That this tool does not produce judgments. That does not mean it is a good clinical tool. Whether it is clinically useful, safe, or meets regulatory requirements is all outside this series, and we did not do the verification those judgments require.

Medical device regulation is not addressed either. At what point software like this falls under regulation differs by jurisdiction, and this piece did not examine that line.

The guard is a word list. It can be routed around if you want to. It is a device against inattention, not against malice.

Writing what you do not do into the code

The original's sentence was right. Leave the diagnosis to the person.

Building it, we add a line. That sentence has to exist somewhere in the code in an executable form. If it lives only in a document, nobody remembers it six months later, and a principle nobody remembers is not a principle.

In this server that form was three things.

  • The absent tool — the list is the surface
  • The throwing guard — sound like a verdict and it raises
  • The failing verification — if either of the two collapses, it does not pass

None of the three is impressive engineering. But all three are code, not prose. That difference is what survives six months.

Practice task

Pick one answer from the plant example and name the record it cites. What would the tool do if no record matched?

Related articleThe Tool That Refuses to Read — Keeping Judgment Outside a Clinic Flow