Foundation · 3 of 9
Two kinds of machinery
Agent runs mix ordinary program control with probabilistic proposals and an unruly outside world. Reliability begins by naming which kind of uncertainty lives where.
§1
“Deterministic or not?” is too coarse
It is tempting to split the system in two: deterministic harness, nondeterministic model. The split points in the right direction, but it leaves out two important layers. Harness behavior depends on configuration and policy, and tool execution depends on an environment that can change between runs.
A more useful control spectrum has four regions: deterministic mechanism, configurable policy, probabilistic model inference, and environmental uncertainty. Each region has a different owner, test strategy, and failure mode. Interpretation
“Deterministic” here means that a defined function should return the same result when its relevant inputs and version are fixed. It does not mean “infallible,” “secure,” or “immune to concurrency.” A deterministic parser can consistently accept the wrong grammar. A deterministic policy can consistently authorize too much.
§2
The control-spectrum matrix
| Region | Typical owner | Repeatability | Enforcement role | Characteristic failure |
|---|---|---|---|---|
| Deterministic mechanism | Harness code | Expected for fixed inputs, code, and dependencies. | Parse messages; validate schemas; select files; enforce counters; append events. | Bug, invariant violation, stale cache, race, or incorrect error handling. |
| Configurable policy | User, operator, or product configuration interpreted by the harness | Repeatable once rules, precedence, identity, and request are fixed. | Allow, deny, ask, constrain, budget, route, retry, or stop. | Ambiguous precedence, unsafe default, scope leak, or a rule that does not match the intended authority. |
| Probabilistic proposal | Model inference under provider and sampling settings | Variation is expected; identical-looking calls may still differ across runs or service versions. | Suggest text, plans, arguments, classifications, and the next action. | Fabrication, omission, malformed structure, unstable choice, or confident misuse of evidence. |
| Environmental outcome | External system plus the tool implementation | Depends on mutable files, processes, networks, clocks, services, and concurrent actors. | Return observations and realize authorized effects. | Timeout, partial write, stale read, rate limit, permission error, or state changed by someone else. |
§3
What the harness can make crisp
A harness can define contracts around uncertain behavior. It can require a tool call to match a JSON schema; normalize provider events into one message type; cap a loop at twenty steps; reject a path outside a workspace; append every effect attempt to an event log; or ensure that a failed tool result is distinguishable from a successful empty result.
These mechanisms are excellent targets for unit and property tests. Given the same event stream, does the session projection reconstruct the same state? Given the same path and policy, does authorization return the same decision? Given a malformed argument object, does the validator refuse execution before side effects?
Deterministic code also makes replay and debugging practical. If the model output and tool observations are recorded, maintainers can replay the harness’s parsing and transition logic without paying for another inference or touching the external system. The replay cannot prove that the original environment was correct; it can show how the controller interpreted the record.
§4
Policy is code with authority behind it
Policy often executes deterministically, yet its meaning comes from configuration and authority. The same shell proposal might be automatically allowed in a disposable container, require approval in a local repository, and be denied in a protected workspace.
Good policy surfaces make inputs visible: requested capability, arguments, current scope, identity, workspace, prior approvals, and rule precedence. Good logs record the decision separately from the tool result. “Denied by policy” is different from “permitted but failed,” and both are different from “the model never proposed it.”
Policy can also govern non-security behavior: context budgets, model routing, retry limits, parallel tool execution, compaction thresholds, or whether a subagent receives fresh versus inherited context. Those choices may be configurable while the code that applies them remains testable.
§5
Use the model for proposals, not guarantees
The model is valuable precisely where a fixed ruleset is brittle: interpreting an underspecified goal, choosing a search strategy, relating unfamiliar code, drafting a patch, or compressing a long history. These are judgments over open-ended inputs. The harness should treat them as proposals whose downstream risk determines the required checks.
A low-risk proposal, such as choosing which read-only search to run, may be executed after schema and scope validation. A high-impact proposal, such as changing a production database or sending a message, may require explicit human approval, stronger confinement, and post-effect verification. The same model uncertainty is present in both; the harness changes the consequence path.
Structured output narrows syntax, not truth. A model can produce perfectly valid JSON naming the wrong file. A type checker can prove the argument has a path string, not that the path is relevant or authorized. Each guard answers one question; reliability comes from composing the right questions in the right order.
§6
The environment gets the last word
A tool call is an attempt to observe or alter external state. Between planning and execution, a branch can advance, a file can be edited, a token can expire, or a service can reject a request. Even a successful exit status may establish less than the agent assumes.
Reliable harnesses preserve environmental evidence: exit code, standard error, response status, affected identifiers, file hashes, timestamps where relevant, or a follow-up read. The needed evidence depends on the claim. “The command ran” may require an exit status. “The deployment is healthy” requires a health observation after deployment, not merely a successful deploy command.
This yields a practical design rule: spend deterministic machinery on boundaries and evidence, and use probabilistic inference for open-ended judgment. Keep uncertainty from crossing a high-consequence boundary without an explicit contract.