Put the Grounds Next to the Answer — Wiring an LLM That Asks the Plant

13 min
GoalA person standing in front of a machine asks in words, and the LLM calls the plant's tools to answer. What matters is not the answer but that what it actually queried appears on screen beside it. An answer that called no tool is marked as such. The whole mcp_client + mcp_llm wiring, the run log, four real render captures.

A maintenance technician standing in front of a machine usually wants a few things. Is this machine past its service interval? Is vibration within limits? What do I check before I touch it?

All of this already exists somewhere in the plant. In the maintenance history database, in the sensor collection system, in the safety checklist document. What is missing is a way for the person standing in front of the machine to ask.

So the thought of attaching an LLM comes naturally. And that is where the real problem starts. If there is no way to tell apart a machine that says "conveyor 3 is past its service interval" from a machine that says so having queried nothing, then the person who trusts that advice and puts a hand inside the machine is in danger.

This article is a record of building that distinction onto the screen. Beside the answer sits the tool call that produced it. And an answer that called no tool is marked as such.

The result first

Asked about a machine. Below the answer sits one line: the tool call that produced it.

Asked about CONV-03 — under the answer, "grounded in 1 call to the plant" in words, and under that the call itself, verbatim. Real render capture

Asked for the checklist. It comes out in the order the plant engineer wrote it.

Asked for the checklist — the call log under the answer holds checklist.get, verbatim

And this is the most important screen in the article. A question the plant cannot answer.

A question the plant cannot answer — no calls at all, and the screen says "no tool call behind this answer — it stands on the model alone"

Zero tools were called, and the screen says so. It does not wear the same face as the two answers before it.

The whole picture

plant_server (mcp_server)      assistant (client + server)          tablet
  equipment.list      ◀──MCP──   a client of the plant              ──MCP──▶  ui://assistant
  equipment.read                 a server to the screen                        answer + call record
  checklist.get                  the model in the middle

Let me state what the middle piece is up front. The model slot in this sample holds a deterministic stub. Anyone has to be able to run and check this without an API key. And that is the least important part of this article — what is worth reading is the wiring on either side of it, and that wiring is the same whether the middle is a stub or Claude. The swap point is shown below as it is.


① The plant server does not judge

The tool side first. One thing is deliberately not done here — the server does not say "this machine is fine / dangerous."

handler: (args) async {
  final id = (args['id'] as String?)?.toUpperCase();
  final m = _machines[id];
  if (m == null) { /* ... */ }

  // The server states facts, and how those facts compare to their limits.
  // It does not say the machine is "fine" — that word belongs to the person
  // holding the checklist.
  final overdue = (m['runHours'] as int) > (m['serviceEveryHours'] as int);
  final vibrationOver =
      (m['vibrationMm'] as num) > (m['vibrationLimitMm'] as num);
  return _json({
    'id': id,
    ...m,
    'serviceOverdue': overdue,
    'vibrationOverLimit': vibrationOver,
  });
}

serviceOverdue: true is a fact. safe: false is a judgement. The server emits only the former.

Same for the checklist. The tool description carries "these steps are set by the plant engineer and must not be paraphrased." A tool description is text the model actually reads.

server.addTool(
  name: 'checklist.get',
  description:
      'Get the plant safety checklist for a machine type (press, conveyor, welder). '
      'These steps are set by the plant engineer and must not be paraphrased.',
  /* ... */
);

② The wiring — where the tools reach the model

This is the substance of the article. Three things joined.

  final llm = McpLlm()..registerProvider('bench', BenchProviderFactory(bench));

  final client = await llm.createClient(
    providerName: 'bench',
    config: LlmConfiguration(model: 'bench-1'),
    mcpClient: mcpClient,
    systemPrompt: assistantSystemPrompt,
  );

  // 3. Ask.
  for (final q in questions) {
    stdout.writeln('\n> $q');
    final response = await client.chat(q, enableTools: true);
    stdout.writeln(response.text.trim());
  }

  // 4. What the plant was actually asked. An assistant's answer is only worth
  //    what the record behind it is worth.
  final audit = await mcpClient.callTool('audit.log', const {});
  final first = audit.content.first;
  if (first is TextContent) {
    final calls = (jsonDecode(first.text) as Map<String, dynamic>)['calls'] as List;
    stdout.writeln('\n# tool calls the plant actually received (${calls.length}):');

The mcpClient: line is the whole of the wiring. chat(enableTools: true) hands the tool list to the model, executes tool calls over MCP when the model calls them, attaches the results and asks once more for the final answer.

Swapping in a real model happens right here too. Left in the sample as a comment.

//      llm.registerProvider('claude', ClaudeProviderFactory());
//      final client = await llm.createClient(
//        providerName: 'claude',
//        config: LlmConfiguration(apiKey: Platform.environment['ANTHROPIC_API_KEY'],
//                                 model: 'claude-sonnet-5'),
//        mcpClient: mcpClient,
//        systemPrompt: systemPrompt,
//      );
//
//    Nothing below this point changes.

Two lines. Below that, not a character changes.

③ The system prompt — every sentence has a reason

Written short. Each sentence is there because removing it produces a specific bad answer.

You help a maintenance technician standing in front of a machine.

Rules:
- Every number you state must have come from a tool result in this conversation.
  If you do not have it, call the tool. Never estimate a reading.
- Safety checklist steps are the plant engineer's. Quote them in order and do
  not paraphrase, shorten or reorder them.
- You do not decide whether a machine is safe to work on. You report what the
  readings are, how they compare to their limits, and what the checklist says.
- If the plant has no tool that answers the question, say so.
  • Remove the first and you get invented readings. A plausible vibration value is indistinguishable from a real one.
  • Remove the second and you get a summarised safety procedure. A four-step checklist shortened to three is not a checklist.
  • Remove the third and you get a judgement. "You may work on it" is not this system's to say.
  • Remove the fourth and it pretends to know what it does not.

But a prompt is a request, not a guarantee. Which is why the next section is needed.

④ Where the grounds get counted

Say "use the tools" in a prompt and never look at whether they were used, and an ungrounded answer looks exactly like a grounded one on screen. So only the tool calls that one question triggered are cut out precisely.

            ? 'No tool was called. Treat this as the assistant talking about '
              'itself, not about the plant.'
            : '';
        return _state();
      },
    );

    server.addTool(
      name: 'assistant.state',
      description: 'Current question, answer and the tool calls behind it',
      inputSchema: const {'type': 'object', 'properties': {}},

And the counting side has to exclude its own calls.

    // audit.log is itself a tool call, but it is ours, not the assistant's —
    // counting it would inflate every answer's evidence by one.

⑤ Fixed while building — zero grounds, rendered green

Taking the first captures, the screen with zero grounds showed "Built from 0 tool call(s)" in green. The sentence is correct. But to someone glancing at it on a factory floor, green reads as "checked, all clear." The opposite of what it means.

            'style': {'fontSize': 13, 'fontWeight': 'bold', 'color': _accent},
          },
        },
        {
          'type': 'box',
          'width': 720,
          'height': 110,
          'child': {
            'type': 'list',
            'items': '{{toolCalls}}',
            'shrinkWrap': true,
            'spacing': 4,

The kind of defect that could not have been caught without looking at rendered pixels. The log said grounded=false calls=0 perfectly correctly.

⑥ The run log

[+ 2240ms] assistant up — it is a client of the plant and a server to us
[+ 2253ms] assistant offers: assistant.ask, assistant.state
[+ 2722ms] Q: how is CONV-03 doing?
[+ 2723ms]    grounded=true calls=1 equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
[+ 2723ms]    A: CONV-03 needs attention — service is overdue (9310 h against a 8000 h
              interval) and vibration is above limit (5.2 mm against 4.5 mm).
[+ 2894ms] Q: what is the checklist before I work on CONV-03?
[+ 2894ms]    grounded=true calls=1 checklist.get(machineType: conveyor) -> 4 steps
[+ 3028ms] Q: what is the weather like?
[+ 3028ms]    grounded=false calls=0
[+ 3029ms]    A: I can look up machines, their current readings, and the plant checklist
              for a machine type. Ask me about one of those.

The same chain runs from the CLI too, and that side shows every call the plant actually received.

# tools offered by the plant: equipment.list, equipment.read, checklist.get, audit.log

# tool calls the plant actually received (4):
#   equipment.list(line: A) -> 2 machines
#   equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
#   checklist.get(machineType: conveyor) -> 4 steps
#   equipment.read(id: PRESS-01) -> overdue=false vibrationOver=false

# provider decisions (9):
#   chose equipment.list(line: A)
#   answer from tool result (141 chars)
#   ...
#   no tool matched — declined to guess

The build passes like this.

$ dart analyze     # plant_server
No issues found!
$ dart analyze     # assistant
No issues found!
$ bash verify.sh
   [player] open in AppPlayer, ask three questions
plant-assistant: two grounded answers with their calls on screen, one ungrounded answer flagged

The verification script sets its pass conditions like this. An answer coming out is not a pass.

# an answer about a machine must have reached the plant
reading = s.call("assistant.ask", {"question": "how is CONV-03 doing?"})
assert reading["grounded"], "a question about a machine must reach the plant"

# an unanswerable question must not have invented tool calls
weather = s.call("assistant.ask", {"question": "what is the weather like?"})
assert not weather["grounded"] and weather["notice"], weather

# and the checklist reaches the screen as a quotation
ap.wait_text("checklist.get")

Measurements, and their range

Value
Tools the plant offered 4
Tool calls triggered by 3 questions 2 (one each for the two grounded answers)
Tool calls for the unanswerable question 0
Round trip to rendered screen about 120–150 ms per question (with the stub model)

The response-time figures mean little. The stub model answers instantly, so the 120–150 ms recorded here is effectively two MCP round trips plus render cost. Attach a real model and inference latency lands wholesale in between — that was not measured.

What was not measured. How accurately a real model picks these tools, how well the system prompt is followed, and what happens to selection accuracy when tools grow to dozens. None of the three can be measured without a real model, and this article does not claim to have measured them. What it showed is the wiring and the grounding-display structure.

The range of this sample

The model is a stub. No real LLM was attached.

Let me not blur that sentence. The stub's "choosing" a tool is a regular expression finding a machine name in the question, and it is different from a real model's judgement. What this article proved is that the wiring really runs and that a structure counting grounds and putting them on screen holds up — not that a model picks tools well. To claim the latter, it must be measured again with a real model.

And the plant data is simulated too. The machine names, running hours, vibration values and checklists are invented by this sample, not any plant's real data.

⑧ Run it yourself

The sample is self-contained in content/sample/plant-assistant/.

( cd plant_server && dart pub get )
( cd assistant && dart pub get )

# ask from the CLI
cd assistant
dart run bin/ask.dart                        # five prepared questions
dart run bin/ask.dart "how is PRESS-01 doing?"

The client targets the AppPlayer standard edition. Pro is not required.

What to take

  1. The three wiring steps — steps 1, 2 and 3 in bin/ask.dart. MCP client → register provider → createClient(mcpClient:) → chat(enableTools: true)
  2. Grounds attribution — the part that cuts out per-question calls by differencing _auditCalls() before and after, own-call exclusion included
  3. Tool design that does not judge — handlers that emit facts and limit comparisons and never emit "safe"

Where to change it for your own

A real model — the two lines registerProvider and createClient in bin/server.dart. Swap BenchProviderFactory for ClaudeProviderFactory and put the key and model name in LlmConfiguration. Below that nothing changes.

Your plant data — the tool handlers in plant_server/bin/server.dart. Today they read a constant map; change them to query the maintenance database or the collection system. As long as the tool names and return shapes match, the assistant side stays.

More tools — one addTool block. And writing the description well matters more than the code — that sentence is what the model uses to choose the tool.

So what this structure sells

There is a sentence in the expert-knowledge-system proposal — "systematically transplanting know-how that lived in an expert's head into a knowledge graph." This article pays the step before that. Before know-how is moved, there has to be a structure that shows whether what was moved was actually queried.

What is sold to a maintenance technician is not "AI will tell you." It is that "this answer came from querying this value on this machine" sits on the screen beside it. Without that, it does not get used on the floor.

Let me note what was not recovered. The consulting-fee payment integration and knowledge graph accumulation the proposal also mentioned are not covered here. And as stated above, the model's tool-selection accuracy is outside this article's measurement.

Practice task

Pick one answer from the plant example and name the record it cites. What would the tool do if no record matched?

Related articlePut the Grounds Next to the Answer — Wiring an LLM That Asks the Plant