把依据摆在答案旁边 — 接一个会去问设备的 LLM

13 分钟
目标站在机器前的人用话来问,LLM 就去调用设备的工具来答。要紧的不是答案,而是这个答案究竟查了什么才出来的,要和答案一起显示在画面上。一个工具都没调用的答案,会被标出来。mcp_client + mcp_llm 的全部接线、运行日志、4 张真实渲染截图。

站在工厂设备前的维修工想知道的通常就那么几件:这台机器过保养周期了没有?振动在限值内吗?动手之前要先确认什么?

这些信息其实早就都在厂里某处。在保养履历数据库里,在传感器采集系统里,在安全点检表文档里。只是站在机器前的人,没有办法去问。

于是「接个 LLM 不就行了」这个念头很自然地冒出来。而真正的问题从那里才开始。如果没法区分「说出『三号传送带过保养期了』的机器」和「这么说但什么都没查过的机器」,那么信了这句建议、把手伸进机器里的人就危险了。

这篇文章记录的,是把这个区分做到画面上为止。答案旁边,会一起出现造出这个答案的那次工具调用。 而一个工具都没调用的答案,会被标出来。

先看结果

问了机器的情况。答案下面附着 造出这个答案的那一行工具调用。

问 CONV-03 —— 答案下面用文字写着 "grounded in 1 call to the plant",再下面原样附上那次调用。真实渲染截图

问了点检表。按工厂工程师写的顺序原样出来。

询问检查清单 —— 答案下面的调用记录里原样留着 checklist.get

而这是本文最重要的一张画面。抛了一个工厂答不了的问题。

与设备无关的问题 —— 调用为 0,屏幕写着 "no tool call behind this answer — it stands on the model alone"

工具被调用了 0 次,画面就这么说。它没有摆出和前面两个答案一样的脸。

全景

plant_server (mcp_server) assistant (client + server) 平板
 equipment.list ◀──MCP── plant 的客户端 ──MCP──▶ ui://assistant
 equipment.read 画面的服务端 答案 + 调用记录
 checklist.get 中间是模型

先把中间那块的身份说明白。这个样例里模型的位置放的是一个确定性桩。 因为任何人都得能在没有 API key 的情况下跑起来验证。而这恰恰是本文最不重要的部分 —— 值得读的是 它两侧的接线,而那套接线不管中间是桩还是 Claude 都一样。替换点在下面原样可见。


① 设备服务端不做判断

先看工具这侧。这里有一件事是刻意不做的 —— 服务端不说「这台机器没问题/危险」。

handler: (args) async {
  final id = (args['id'] as String?)?.toUpperCase();
  final m = _machines[id];
  if (m == null) { /* ... */ }

  // The server states facts, and how those facts compare to their limits.
  // It does not say the machine is "fine" — that word belongs to the person
  // holding the checklist.
  final overdue = (m['runHours'] as int) > (m['serviceEveryHours'] as int);
  final vibrationOver =
      (m['vibrationMm'] as num) > (m['vibrationLimitMm'] as num);
  return _json({
    'id': id,
    ...m,
    'serviceOverdue': overdue,
    'vibrationOverLimit': vibrationOver,
  });
}

serviceOverdue: true 是事实。safe: false 是判断。服务端只出前者。

点检表也一样。工具的说明文里写了 「这些步骤由工厂工程师制定,不得改写」。工具说明文是模型真的会读的文本。

server.addTool(
 name: 'checklist.get',
 description:
 'Get the plant safety checklist for a machine type (press, conveyor, welder). '
 'These steps are set by the plant engineer and must not be paraphrased.',
 /* ... */
);

② 接线 —— 工具够到模型的地方

这是本文的正题。把三块接起来。

// 1. Attach to the plant. Not a privileged channel — an ordinary MCP client.
final connected = await McpClient.createAndConnect(
 config: McpClient.simpleConfig(name: 'Plant Assistant', version: '1.0.0'),
 transportConfig: const TransportConfig.stdio(
 command: 'dart',
 arguments: ['run', 'bin/server.dart'],
 workingDirectory: '../plant_server',
 ),
);
final mcpClient = connected.get();

// 2. Register a provider and make the client that joins the two.
final bench = BenchProvider();
final llm = McpLlm()..registerProvider('bench', BenchProviderFactory(bench));

final client = await llm.createClient(
 providerName: 'bench',
 config: LlmConfiguration(model: 'bench-1'),
 mcpClient: mcpClient, // ← the tools go in here
 systemPrompt: assistantSystemPrompt,
);

// 3. Ask. Passing the tool list, executing the calls and feeding the results
// back all happens inside this one line.
final response = await client.chat(question, enableTools: true);

mcpClient: 这一行就是接线的全部。chat(enableTools: true) 把工具清单递给模型,模型调用工具时经 MCP 执行,再把结果附上去问一次拿到最终答案。

换成真实模型也是在这个位置。 样例里以注释留着。

// llm.registerProvider('claude', ClaudeProviderFactory());
// final client = await llm.createClient(
// providerName: 'claude',
// config: LlmConfiguration(apiKey: Platform.environment['ANTHROPIC_API_KEY'],
// model: 'claude-sonnet-5'),
// mcpClient: mcpClient,
// systemPrompt: systemPrompt,
// );
//
// Nothing below this point changes.

两行。它下面一个字都不变。

③ 系统提示词 —— 每一句都有理由

写得短。每一句之所以在那儿,是因为把它抽掉就会出现某种特定的坏答案。

You help a maintenance technician standing in front of a machine.

Rules:
- Every number you state must have come from a tool result in this conversation.
 If you do not have it, call the tool. Never estimate a reading.
- Safety checklist steps are the plant engineer's. Quote them in order and do
 not paraphrase, shorten or reorder them.
- You do not decide whether a machine is safe to work on. You report what the
 readings are, how they compare to their limits, and what the checklist says.
- If the plant has no tool that answers the question, say so.
  • 抽掉第一句,就会出现 编出来的数值。 听着合理的振动值和真实振动值分不出来。
  • 抽掉第二句,就会出现 被概括过的安全流程。 四步压成三步的点检表,就不是点检表了。
  • 抽掉第三句,就会出现 判断。「可以动手了」不是这个系统该说的话。
  • 抽掉第四句,它就会 不懂装懂。

不过提示词是请求,不是保证。所以才需要下一节。

④ 数依据的地方

在提示词里说「要用工具」,却不去看到底用没用,那么没用工具的答案和用了的在画面上长得一模一样。所以要把 一个问题所引发的工具调用 精确地切出来。

// Mark the current point in the plant's audit log, so only the calls this
// question triggered can be attributed to it — not everything since boot.
final before = await _auditCalls();
final response = await llm.chat(question, enableTools: true);
final after = await _auditCalls();

_answer = response.text.trim();
_toolCalls = after.sublist(before.length);
_notice = _toolCalls.isEmpty
 ? 'No tool was called. Treat this as the assistant talking about '
 'itself, not about the plant.'
 : '';

而数的那一侧要把自己的调用排除掉。

// audit.log is itself a tool call, but it is ours, not the assistant's —
// count it and the grounds under every answer inflate by one.
return calls.cast<String>().where((c) => !c.startsWith('audit.log')).toList();

⑤ 做的时候改掉的 —— 0 次却是绿色的

第一次截图出来一看,依据为 0 的那张画面上 "Built from 0 tool call(s)" 是绿色的。 句子没错。可对在车间地面上瞥一眼的人来说,绿色读作「已确认,无异常」。意思正好相反。

// A grounding count must not wear a reassuring face when there is no
// grounding. Green on "0 tool calls" reads at a glance as "checked, all
// clear", which is the opposite of what it means.
{
 'type': 'conditional',
 'condition': '{{grounded}}',
 'then': { /* 绿色,N 次 */ },
 'else': {
 'type': 'text',
 'content': 'Not built from any plant data',
 'style': {'fontSize': 13, 'fontWeight': 'bold', 'color': _accent},
 },
},

这类缺陷,不亲眼看渲染出来的像素是抓不到的。日志里明明白白写着 grounded=false calls=0。

⑥ 运行与校验日志

[+ 2240ms] assistant up — it is a client of the plant and a server to us
[+ 2253ms] assistant offers: assistant.ask, assistant.state
[+ 2722ms] Q: how is CONV-03 doing?
[+ 2723ms] grounded=true calls=1 equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
[+ 2723ms] A: CONV-03 needs attention — service is overdue (9310 h against a 8000 h
 interval) and vibration is above limit (5.2 mm against 4.5 mm).
[+ 2894ms] Q: what is the checklist before I work on CONV-03?
[+ 2894ms] grounded=true calls=1 checklist.get(machineType: conveyor) -> 4 steps
[+ 3028ms] Q: what is the weather like?
[+ 3028ms] grounded=false calls=0
[+ 3029ms] A: I can look up machines, their current readings, and the plant checklist
 for a machine type. Ask me about one of those.

用 CLI 也能跑同一条链,那边会把设备实际收到的调用全都列出来。

# tools offered by the plant: equipment.list, equipment.read, checklist.get, audit.log

# tool calls the plant actually received (4):
# equipment.list(line: A) -> 2 machines
# equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
# checklist.get(machineType: conveyor) -> 4 steps
# equipment.read(id: PRESS-01) -> overdue=false vibrationOver=false

# provider decisions (9):
# chose equipment.list(line: A)
# answer from tool result (141 chars)
# ...
# no tool matched — declined to guess

构建是这样通过的。

$ dart analyze     # plant_server
No issues found!
$ dart analyze     # assistant
No issues found!
$ bash verify.sh
   [player] open in AppPlayer, ask three questions
plant-assistant: two grounded answers with their calls on screen, one ungrounded answer flagged

校验脚本把通过条件定成了这样。光是出了答案,不算通过。

# an answer about a machine must have reached the plant
reading = s.call("assistant.ask", {"question": "how is CONV-03 doing?"})
assert reading["grounded"], "a question about a machine must reach the plant"

# an unanswerable question must not have invented tool calls
weather = s.call("assistant.ask", {"question": "what is the weather like?"})
assert not weather["grounded"] and weather["notice"], weather

# and the checklist reaches the screen as a quotation
ap.wait_text("checklist.get")

实测值及其范围

值
设备提供的工具 4
3 个问题引发的工具调用 2(两个有依据的答案各 1)
答不了的问题的工具调用 0
到画面渲染的往返 每个问题约 120–150 ms(以桩模型为准)

响应时间的数字意义有限。 桩模型是即刻作答的,所以这里记下的 120–150 ms 实际上是两次 MCP 往返加上渲染开销。接上真实模型,推理时延会整块塞进中间 —— 那个没有测。

没能测的。 真实模型挑这些工具挑得多准、系统提示词被遵守到什么程度、工具增到几十个时选择精度如何。三样没有真实模型都测不了,而 这篇文章不主张自己测过那一部分。 本文展示的是接线和依据显示的结构。

这个样例的范围

模型是桩。没有接真实的 LLM。

这句话不含糊。桩「挑」工具靠的是从问题里找机器名的正则,和真实模型的判断不是一回事。这篇文章证明的是 接线真的跑得通 和 数依据并把它摆上画面的结构成立,而不是模型挑工具挑得好。要主张后者,得用真实模型重新测。

而且设备数据也是模拟的。机器名、运行时长、振动值、点检表都是这个样例编出来的,不是哪家工厂的真实数据。

⑧ 自己跑一遍

样例自包含地放在 content/sample/plant-assistant/ 里。

( cd plant_server && dart pub get )
( cd assistant && dart pub get )

# ask from the CLI
cd assistant
dart run bin/ask.dart                        # five prepared questions
dart run bin/ask.dart "how is PRESS-01 doing?"

可以带走的

  1. 接线三步 — bin/ask.dart 里的第 1、2、3 步。MCP 客户端 → 注册供应商 → createClient(mcpClient:) → chat(enableTools: true)
  2. 依据归属 — 用 _auditCalls() 前后差分切出每个问题的调用的那一段,包括排除自己的调用
  3. 不做判断的工具设计 — 只出事实与限值比较、不出「安全」的处理器

换成自己的要改哪里

换成真实模型 — bin/server.dart 里 registerProvider 与 createClient 两行。把 BenchProviderFactory 换成 ClaudeProviderFactory,在 LlmConfiguration 里填上 key 和模型名。它下面不变。

换成自己的设备数据 — plant_server/bin/server.dart 里的工具处理器。现在读的是常量映射,改成去查保养数据库或采集系统。工具名和返回形状一致,助手那侧就原样不动。

要加工具 — 一块 addTool。而且 把说明文写好比代码更要紧 —— 模型挑工具的依据就是那句话。

所以这个结构卖的是什么

专家知识系统提案书里有这么一句 —— 「把原本在专家个人脑中的诀窍,体系化地移植到知识图谱。」 这篇文章还的是它前面那一格。在移植诀窍之前,得先有 一个能看出「移过去的东西是否真的被查过」的结构。

卖给维修工的不是「AI 会告诉你」。而是 「这个答案是查了这台机器的这个值才出来的」 和答案一起摆在画面上。没有这一条,现场就不会用。

把没有兑现的写下来。提案书一并提到的 顾问费支付对接 与 知识图谱积累,这一篇没有涉及。而如前所述,模型挑工具的准确度 在本文的测量范围之外。

练习任务

从设备示例中选一个回答,指出它引用的记录。如果没有匹配的记录,工具应该怎么做?

相关文章Put the Grounds Next to the Answer — Wiring an LLM That Asks the Plant