Foundation · 6 of 9
Tools, authority, and trust
A tool definition makes an action expressible. It does not make the action valid, permitted, confined, successful, or wise.
§1
A tool has a model face and an effect face
The model face is descriptive: a name, a purpose, and a machine-readable argument schema. The effect face is operational: code running with some filesystem, process, network, credential, and user authority. The harness is responsible for connecting the two without letting a plausible proposal masquerade as permission.
Seven questions should remain separate: Is the capability registered? Is it visible to this model call? Are the arguments valid? Is this use authorized? Is execution confined? What did the implementation return? What effect can be verified afterward? Interpretation
A system may answer “yes” to one question and “no” to the next. A shell tool can be installed but withheld from a read-only subagent. A file-write call can match its schema but target a denied path. An authorized command can run inside a sandbox and still exit with an error. A successful API response can still fail to establish the business outcome the user intended.
§2
The trust pipeline
| Stage | Question | Typical mechanism | Evidence emitted |
|---|---|---|---|
| Registration | Can the harness resolve this capability? | Built-in registry, plugin service, or extension registration. | Tool identity, version, implementation/provider. |
| Presentation | Should the model know about it in this step? | Scope filtering, task profile, tool selection, compact tool descriptions. | The exact model-visible name, description, and schema. |
| Validation | Does the proposal match the contract? | JSON parsing, schema checks, type normalization, semantic guards. | Typed arguments or a structured validation error. |
| Authorization | May this principal perform this operation here? | Allow/deny/ask rules, workspace boundaries, identity and intent checks. | Policy decision, rule, scope, and approval reference. |
| Confinement | What can the process reach if it behaves unexpectedly? | Sandbox, container, filesystem allowlist, network controls, reduced credentials. | Effective execution boundary. |
| Execution | What did the tool attempt and return? | Process/API/filesystem adapter with timeout and cancellation. | Status, output, diagnostics, affected identifiers. |
| Verification | What effect can we establish? | Read-after-write, test, diff, health check, receipt, or user confirmation. | Postcondition evidence or an explicit unverified status. |
§3
Schema-valid can still be wrong
A schema protects the implementation from malformed structure. It can require path to be a string and line to be a positive integer. It cannot prove that the file is relevant, the edit matches the user’s intent, or the path falls inside the current authority.
Semantic guards handle domain invariants: a resolved path must remain under the workspace; a command must not contain an unsupported mode; a destination must belong to the selected account; a patch must match the file version it was based on. Guards should run before effects and fail closed with a precise reason.
When the model can repair a validation failure, return a bounded, structured error through the agent protocol. The model can repair an omitted field if it sees which contract failed. It should not receive a success-shaped empty result, nor a raw internal exception that leaks unrelated data.
§4
Approval transfers a bounded decision
An approval prompt is useful when a human can understand the proposed action and its scope. “Allow this exact command in this directory once?” is a bounded decision. “Let the agent do whatever is necessary?” is difficult to evaluate and easy to overread.
Authority comes from the user or operating environment, not from the model’s confidence. The harness should preserve the subject, operation, target, scope, and duration of an approval. Reusing approval beyond those bounds turns a safety surface into a confused-deputy risk.
A confused deputy appears when the harness has power that the content requester should not control. For example, a document may contain text asking the agent to send private workspace material elsewhere. The document is data, not an authority-bearing user instruction. The harness must keep source and precedence visible.
§5
Tool output is evidence and untrusted input
Files, web pages, issue bodies, terminal output, and API responses can contain instruction-like text. That text may be relevant evidence, stale guidance, or an attempt to redirect the run. When the harness places it beside trusted instructions without clear provenance, the model may follow the wrong authority.
Prompt injection is a system problem. A bad sentence can trigger it, but the available authority and runtime boundaries determine the possible damage. Useful defenses include labeling source and trust level, limiting which tools a context can invoke, keeping credentials out of model-visible text, requiring approval for consequential actions, and verifying effects independently. No single prompt can enforce a boundary that the runtime leaves open.
§6
Preserve the difference between request, attempt, and effect
A complete record may contain: the model’s proposed call, the harness’s normalized arguments, the policy decision, any user approval, the effective sandbox, the execution attempt, the raw or referenced result, a bounded model-visible observation, and postcondition evidence. These are not redundant. They answer different audit and recovery questions.
Error propagation should retain structured status while keeping the next context usable. A tool timeout can include the operation ID, elapsed limit, partial output reference, and retry policy. The model can then decide whether a retry is justified. A generic “something went wrong” blocks recovery; an uncapped log dump can consume the remaining context and hide the useful line.
Verification should match consequence. A read-only search may need no postcondition. A file edit deserves a diff or reread. A code change deserves the smallest relevant test. An external message deserves a receipt or draft review. Making each consequential transition observable and bounded matters more than trying to make the model certain.
§7
Verify completion outside the implementation story
Post-effect verification asks whether one operation produced its intended result. A completion gate asks a different question: does the finished artifact satisfy the original acceptance criteria? Re-read the task, inspect the current artifact, and run evidence-producing checks. The implementing agent's declaration is useful state, but it is not completion evidence.
A clean-context evaluator or separately assigned reviewer can reduce dependence on the sequence of guesses and repairs that produced the artifact. Independence does not make the evaluator correct; its report still needs tests, browser flows, static checks, traces, receipts, or other inspectable evidence tied to the original criteria. Interpretation
Measure a tool through four separate gates
01 · Available
Was there a fair opportunity?
Was the capability registered, presented in the relevant scope, and described clearly enough to choose?
02 · Adopted
Did the model select it?
Count legitimate opportunities separately from actual calls. An installed tool that is never selected has not entered the workflow.
03 · Outcome
Did the task improve?
Measure task success, evidence quality, error rate, or another downstream result. Use alone is not benefit.
04 · Cost
What did the use consume?
Track latency, tokens, money, failures, and variance. A useful tool can still be the wrong trade at its operating cost.
This sequence prevents three common category errors: availability is not adoption, adoption is not task benefit, and benefit does not settle cost. Nuanced's vendor-run LSP evaluation is a useful concrete example because it measured eligible opportunities, actual use, task results, duration, tokens, and cost separately; its particular LSP result should not be generalized to every tool. Interpretation