答えの横に根拠を付ける — 設備に問い合わせる LLM をつなぐ
13 分工場設備の前に立つ整備士が知りたいのはたいてい数点だ。この機械は整備周期を過ぎているか。振動は限界内か。手を入れる前に何を確認すべきか。
この情報はすでに工場のどこかに全部ある。整備履歴の DB に、センサー収集システムに、安全点検表の文書に。ただ、機械の前に立つ人に問う手段が無い。
だから LLM をつなげばいい、という発想が自然に出てくる。そしてそこから本当の問題が始まる。「コンベア 3 号は整備周期を過ぎています」と言う機械と、そう言いながら何も照会していない機械を区別する方法が無ければ、その助言を信じて機械の中に手を入れる人が危険になる。
この記事はその区別を画面に載せるところまで作ってみた記録だ。答えの横に、その答えを作った道具の呼び出しが一緒に出る。 そして道具がひとつも呼ばれなかった答えは、そうと表示される。
まず結果
機械について問うた。答えの下に その答えを作った道具の呼び出しが一行付いている。

点検表を問うた。工場のエンジニアが書いた順序のまま出てくる。

そしてこれがこの記事でいちばん重要な画面だ。工場が答えられない質問を投げた。

道具が 0 件 呼ばれ、画面がそう言っている。前の二つの答えと同じ顔をしていない。
全体像
plant_server (mcp_server) assistant (client + server) tablet
equipment.list ◀──MCP── a client of the plant ──MCP──▶ ui://assistant
equipment.read a server to the screen answer + call record
checklist.get the model in the middle真ん中の部品の正体を先に明かしておく。このサンプルのモデルの席には決定論的なスタブが入っている。 API キー無しで誰でも回して確かめられなければならないからだ。そしてそれがこの記事でいちばん重要でない部分である — 読む値打ちがあるのは その両脇の配線であり、その配線は真ん中がスタブでも Claude でも同じだ。差し替え地点は下にそのまま見える。
① 設備サーバーは判断しない
まず道具の側。ここで意図的にひとつやらない — 「この機械は大丈夫/危険」をサーバーが言わない。
handler: (args) async {
final id = (args['id'] as String?)?.toUpperCase();
final m = _machines[id];
if (m == null) { /* ... */ }
// The server states facts, and how those facts compare to their limits.
// It does not say the machine is "fine" — that word belongs to the person
// holding the checklist.
final overdue = (m['runHours'] as int) > (m['serviceEveryHours'] as int);
final vibrationOver =
(m['vibrationMm'] as num) > (m['vibrationLimitMm'] as num);
return _json({
'id': id,
...m,
'serviceOverdue': overdue,
'vibrationOverLimit': vibrationOver,
});
}serviceOverdue: true は事実だ。safe: false は判断だ。サーバーは前者だけを出す。
点検表も同じだ。道具の説明文に **「この手順は工場のエンジニアが定めたもので、言い換えてはならない」**を入れた。道具の説明文はモデルが実際に読むテキストである。
server.addTool(
name: 'checklist.get',
description:
'Get the plant safety checklist for a machine type (press, conveyor, welder). '
'These steps are set by the plant engineer and must not be paraphrased.',
/* ... */
);② 配線 — 道具がモデルに届く場所
ここがこの記事の本論だ。三つをつなぐ。
final llm = McpLlm()..registerProvider('bench', BenchProviderFactory(bench));
final client = await llm.createClient(
providerName: 'bench',
config: LlmConfiguration(model: 'bench-1'),
mcpClient: mcpClient,
systemPrompt: assistantSystemPrompt,
);
// 3. Ask.
for (final q in questions) {
stdout.writeln('\n> $q');
final response = await client.chat(q, enableTools: true);
stdout.writeln(response.text.trim());
}
// 4. What the plant was actually asked. An assistant's answer is only worth
// what the record behind it is worth.
final audit = await mcpClient.callTool('audit.log', const {});
final first = audit.content.first;
if (first is TextContent) {
final calls = (jsonDecode(first.text) as Map<String, dynamic>)['calls'] as List;
stdout.writeln('\n# tool calls the plant actually received (${calls.length}):');mcpClient: の一行が配線の全部だ。chat(enableTools: true) が道具の一覧をモデルに渡し、モデルが道具を呼べば MCP で実行し、結果を付けてもう一度問うて最終の答えを受け取る。
実際のモデルに変えるのもこの場所だ。 サンプルにコメントで残してある。
// llm.registerProvider('claude', ClaudeProviderFactory());
// final client = await llm.createClient(
// providerName: 'claude',
// config: LlmConfiguration(apiKey: Platform.environment['ANTHROPIC_API_KEY'],
// model: 'claude-sonnet-5'),
// mcpClient: mcpClient,
// systemPrompt: systemPrompt,
// );
//
// Nothing below this point changes.二行だ。その下は一文字も変わらない。
③ システムプロンプト — 文ごとに理由がある
短く書いた。各文がそこにある理由は、その文を抜くと特定の悪い答えが出るからだ。
You help a maintenance technician standing in front of a machine.
Rules:
- Every number you state must have come from a tool result in this conversation.
If you do not have it, call the tool. Never estimate a reading.
- Safety checklist steps are the plant engineer's. Quote them in order and do
not paraphrase, shorten or reorder them.
- You do not decide whether a machine is safe to work on. You report what the
readings are, how they compare to their limits, and what the checklist says.
- If the plant has no tool that answers the question, say so.- 一行目を抜けば でっち上げの数値が出る。もっともらしい振動値は実際の振動値と区別が付かない。
- 二行目を抜けば 要約された安全手順が出る。4 段階を 3 段階に縮めた点検表は点検表ではない。
- 三行目を抜けば 判断が出る。「作業して構いません」はこのシステムが言う言葉ではない。
- 四行目を抜けば 知らないことを知ったふりをする。
ただしプロンプトは依頼であって保証ではない。だから次の節が要る。
④ 根拠を数える場所
プロンプトで「道具を使え」と言っておいて実際に使ったかを見なければ、使わなかった答えと使った答えが画面で同じ顔をする。だから 質問ひとつが誘発した道具呼び出しだけを正確に切り出す。
? 'No tool was called. Treat this as the assistant talking about '
'itself, not about the plant.'
: '';
return _state();
},
);
server.addTool(
name: 'assistant.state',
description: 'Current question, answer and the tool calls behind it',
inputSchema: const {'type': 'object', 'properties': {}},そして数える側は自分の呼び出しを除かねばならない。
// audit.log is itself a tool call, but it is ours, not the assistant's —
// counting it would inflate every answer's evidence by one.⑤ 作りながら直したもの — 0 件なのに緑だった
キャプチャを最初に撮ってみたら、根拠 0 件の画面に "Built from 0 tool call(s)" が緑色で出ていた。文としては正しい。ところが工場の床でちらりと見る人にとって、緑は「確認済み、異常なし」と読める。意味が正反対だ。
'style': {'fontSize': 13, 'fontWeight': 'bold', 'color': _accent},
},
},
{
'type': 'box',
'width': 720,
'height': 110,
'child': {
'type': 'list',
'items': '{{toolCalls}}',
'shrinkWrap': true,
'spacing': 4,レンダーされたピクセルを目で見ていなければ捕まえられなかった種類の欠陥だ。ログには grounded=false calls=0 と正確に出ていたのだから。
⑥ 実行・検証ログ
[+ 2240ms] assistant up — it is a client of the plant and a server to us
[+ 2253ms] assistant offers: assistant.ask, assistant.state
[+ 2722ms] Q: how is CONV-03 doing?
[+ 2723ms] grounded=true calls=1 equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
[+ 2723ms] A: CONV-03 needs attention — service is overdue (9310 h against a 8000 h
interval) and vibration is above limit (5.2 mm against 4.5 mm).
[+ 2894ms] Q: what is the checklist before I work on CONV-03?
[+ 2894ms] grounded=true calls=1 checklist.get(machineType: conveyor) -> 4 steps
[+ 3028ms] Q: what is the weather like?
[+ 3028ms] grounded=false calls=0
[+ 3029ms] A: I can look up machines, their current readings, and the plant checklist
for a machine type. Ask me about one of those.CLI でも同じ連鎖を回せて、そちらは設備が実際に受けた呼び出しを全部見せる。
# tools offered by the plant: equipment.list, equipment.read, checklist.get, audit.log
# tool calls the plant actually received (4):
# equipment.list(line: A) -> 2 machines
# equipment.read(id: CONV-03) -> overdue=true vibrationOver=true
# checklist.get(machineType: conveyor) -> 4 steps
# equipment.read(id: PRESS-01) -> overdue=false vibrationOver=false
# provider decisions (9):
# chose equipment.list(line: A)
# answer from tool result (141 chars)
# ...
# no tool matched — declined to guessビルドはこう通る。
$ dart analyze # plant_server
No issues found!
$ dart analyze # assistant
No issues found!
$ bash verify.sh
[player] open in AppPlayer, ask three questions
plant-assistant: two grounded answers with their calls on screen, one ungrounded answer flagged検証スクリプトは通過条件をこう置いた。答えが出たというだけでは通過ではない。
# an answer about a machine must have reached the plant
reading = s.call("assistant.ask", {"question": "how is CONV-03 doing?"})
assert reading["grounded"], "a question about a machine must reach the plant"
# an unanswerable question must not have invented tool calls
weather = s.call("assistant.ask", {"question": "what is the weather like?"})
assert not weather["grounded"] and weather["notice"], weather
# and the checklist reaches the screen as a quotation
ap.wait_text("checklist.get")実測値とその範囲
| 値 | |
|---|---|
| 設備が提供した道具 | 4 |
| 質問 3 件が誘発した道具呼び出し | 2(根拠のある答え 2 件に各 1) |
| 答えられない質問の道具呼び出し | 0 |
| 画面レンダーまでの往復 | 質問あたり約 120〜150 ms(スタブモデル基準) |
応答時間の数値は意味が限られる。 スタブモデルは即座に答えるので、ここに出た 120〜150 ms は実質 MCP の往復二回とレンダーの費用だ。実際のモデルをつなげばその間に推論の遅延が丸ごと入る — それは測っていない。
範囲の外。 実際のモデルがこれらの道具をどれだけ正確に選ぶか、システムプロンプトがどれだけ守られるか、道具が数十個に増えたとき選択の精度がどうなるか。三つとも実際のモデル無しには測れず、この記事はその部分を測定したとは主張しない。 この記事が示したのは配線と根拠表示の構造だ。
このサンプルの範囲
モデルはスタブだ。実際の LLM はつないでいない。
この文をぼかさない。スタブが道具を「選ぶ」のは質問から機械名を探す正規表現であり、実際のモデルの判断とは違う。この記事が証明したのは 配線が実際に回ることと 根拠を数えて画面に載せる構造が成立することであって、モデルが道具をうまく選ぶことではない。後者を主張するには実際のモデルで測り直さねばならない。
そして設備データもシミュレートだ。機械名・運転時間・振動値・点検表はこのサンプルがでっち上げたものであり、どこかの工場の実データではない。
⑧ 自分で動かす
サンプルは content/sample/plant-assistant/ に自己完結で入っている。
( cd plant_server && dart pub get )
( cd assistant && dart pub get )
# ask from the CLI
cd assistant
dart run bin/ask.dart # five prepared questions
dart run bin/ask.dart "how is PRESS-01 doing?"クライアントは AppPlayer 標準版が基準だ。Pro は要らない。
持ち帰るもの
- 配線 3 段階 —
bin/ask.dartの 1・2・3 段階。MCP クライアント → プロバイダ登録 →createClient(mcpClient:)→chat(enableTools: true) - 根拠の帰属 —
_auditCalls()の前後差分で質問ごとの呼び出しだけを切り出す部分。自分の呼び出しの除外まで - 判断しない道具設計 — 事実と限界との比較までを出し、「安全」は出さないハンドラ
自分のものに変えるにはどこを直すか
実際のモデルへ — bin/server.dart の registerProvider と createClient の二行だ。BenchProviderFactory を ClaudeProviderFactory に変え、LlmConfiguration にキーとモデル名を入れる。その下は変わらない。
自分の設備データへ — plant_server/bin/server.dart の道具ハンドラだ。いまは定数のマップを読むが、整備 DB や収集システムを照会するように変える。道具の名前と返す形が同じなら、アシスタント側はそのままだ。
道具を増やすには — addTool のひと塊。そして 説明文をうまく書くことがコードより重要だ — モデルが道具を選ぶ根拠がその文だからである。
だからこの構造が売るもの
専門家知識システムの提案書にこういう一文がある — 「専門家個人の頭の中にあったノウハウを知識グラフへ体系的に移植。」 この記事はその手前の一段を返す。ノウハウを移す前に、移したものが実際に照会されたかが見える構造が先に無ければならない。
整備士に売るのは「AI が教えてくれる」ではない。「この答えはこの機械のこの値を照会して出た」 が画面に一緒にあることだ。それが無ければ現場では使われない。
回収できていないものを書いておく。提案書が併せて語った 顧問料の決済連携と 知識グラフの蓄積 はこの記事では扱わなかった。そして先に述べたとおり モデルの道具選択の精度 はこの記事の測定範囲の外である。
練習課題
設備の例から答えをひとつ選び、それが引用している記録を指してください。該当する記録がなければツールは何をすべきですか。