Foundation · 2 of 9
One turn, end to end
A tool-using turn is a small protocol. Messages cross the model boundary; proposed calls cross the authority boundary; observations cross back into context.
§1
A turn can contain several model calls
A user experiences one turn: ask a question, wait, receive a response. Inside the harness, that turn may contain many steps. The model asks to inspect a file; the harness returns it; the model asks to run one test; the harness returns the failure; the model finally explains the cause.
The reusable pattern is an alternation between model inference and harness-mediated action. Each new observation changes the next model request. The loop ends only when a stopping rule says it should. Interpretation
while turn is active:
request = build_model_request(state)
proposal = call_model(request)
if proposal contains tool calls:
results = validate_authorize_execute(proposal.calls)
state = record(state, proposal, results)
else:
return finalize(proposal)
The pseudocode omits streaming, cancellation, parallelism, retries, and provider-specific formats. It keeps the ownership visible: the model proposes; the harness decides what crosses into execution and what becomes the next request.
§2
A successful trace
- User request: “Find the failing test and explain the cause.”
- Harness assembly: record the message; select instructions, history, tools, model, and settings.
- Model inference: produce a continuation from the assembled request.
- Tool proposal: emit a structured search or test request.
- Harness mediation: resolve the tool; parse arguments; validate schema; apply policy; obtain approval when required.
- Environment execution: run the bounded operation and return status, output, and diagnostics.
- Observation recording: normalize and store the tool result; decide how much enters the next context.
- Next inference: ask the model again with the new evidence.
- Stop handling: accept a terminal response or continue if another valid call is proposed.
- User response: finalize the turn and preserve its durable record.
§3
Three boundaries in every tool call
1. Model boundary: messages in, proposals out
A model adapter translates harness messages and tool definitions into the provider’s wire format. Streaming output may arrive as text deltas, reasoning events, partial JSON, completed tool calls, usage data, or errors. The harness must assemble those fragments into its own message and event model.
2. Authority boundary: proposal to permitted operation
A valid tool name and valid JSON arguments establish shape, not permission. The harness may apply allowlists, path constraints, approval rules, sandbox modes, rate limits, or task-specific policy. A denial is itself an observation that the loop must represent honestly.
3. Reality boundary: attempted operation to observed effect
Execution can fail after authorization: a file moved, a process exited nonzero, a network call timed out, or an API accepted a request but did not create the intended state. The tool result should preserve enough status and evidence for the harness and model to distinguish “requested,” “attempted,” and “verified.”
§4
The branches a real loop must represent
| Branch | What the harness does | What should enter state | Possible next step |
|---|---|---|---|
| Invalid arguments | Rejects a missing field, wrong type, malformed JSON, or value outside the tool contract. | A structured validation error; no claim that execution occurred. | The model repairs the call or chooses another tool. |
| Unknown tool | Fails lookup or refuses a name that was not exposed in this request. | The unresolved name and bounded error. | The model uses an available capability or stops. |
| Permission denied | Applies policy or records the user’s refusal without executing. | A denial event and, where appropriate, its policy reason. | The model proposes a safer route or asks the user for a decision. |
| Tool failure | Captures nonzero exit, exception, timeout, cancellation, or provider error. | Status, diagnostics, and any partial result clearly marked as partial. | The model diagnoses, retries within policy, or reports the blocker. |
| Successful result | Normalizes output, applies size limits, records provenance, and updates state. | The observation plus a reference to any larger artifact kept off-context. | The model reasons from the evidence or requests another operation. |
The error channel is part of the protocol. Converting a timeout or denial into an empty “success” teaches the model a false world. Good harnesses make failure legible enough that the next inference can respond to the real condition.
§5
Stopping is a harness decision too
A model response with no tool call is a common terminal signal, but production loops need more: maximum steps, elapsed-time budgets, token or cost limits, cancellation, repeated-call detection, unrecoverable provider errors, user interruption, and explicit “needs input” states. The harness converts those conditions into a durable outcome.
There is a subtle asymmetry. The model can propose “I am done,” yet the harness decides whether the protocol accepts that proposal. The harness may require a final schema, a verification step, or a user confirmation. Conversely, the harness may stop a model that would prefer to continue because the run has exhausted time, authority, or context.