Foundation · 8 of 9
How to read a harness
Follow one turn through contracts and state transitions. A directory tour tells you where files live; a turn trace tells you how the system works.
§1
Choose one concrete scenario
Use a request that requires at least one tool and one second inference: “Read the configuration, find the active timeout, and explain where it is enforced.” Record the expected arrows before opening code: request, context assembly, provider call, streamed proposal, validation, authorization, execution, result recording, next provider call, stop.
Pin the repository version. Generated architecture docs and code graphs can accelerate navigation, but they are maps, not proof. Verify each important edge at the pinned source. If a graph labels an edge INFERRED, treat it as a hypothesis until the implementation confirms it.
§2
The eleven-stop reading order
| Stop | Find | Question |
|---|---|---|
| 1 · Entry point | CLI command, SDK method, HTTP handler, or UI action. | What creates the session and starts a turn? |
| 2 · Main loop | The function that alternates model calls, tools, and stopping. | What are its inputs, events, and terminal states? |
| 3 · Messages | Internal message/event types and conversion functions. | Which information is durable, model-visible, or UI-only? |
| 4 · Model adapter | Provider request builder and stream normalizer. | Where do provider-specific roles and tool formats enter? |
| 5 · Tools | Registry, schema, lookup, preparation, dispatch, result type. | When does a proposal become executable input? |
| 6 · Context | Token estimates, selection, pruning, retrieval, and compaction. | How is the next working set constructed? |
| 7 · Persistent state | Session store, event log, snapshots, artifacts, branches. | What survives restart and what can be replayed? |
| 8 · Permissions | Allow/deny/ask rules, approval UI, sandbox and path guards. | Which layer enforces authority before effects? |
| 9 · Failure semantics | Timeouts, retries, cancellation, partial output, invalid calls. | How does each failure reach the next inference and the user? |
| 10 · Extensions | Plugins, hooks, skills, providers, custom tools, subagents. | Which contracts may extensions observe or replace? |
| 11 · Observability | Typed events, tracing IDs, logs, metrics, and debug views. | Can one effect be traced back to its prompt and policy decision? |
§3
Trace data, not only calls
A call graph shows that runTurn() invokes executeTool(). Trace the data at that edge: raw model JSON or validated arguments; tool name or resolved implementation; user request or derived authority; direct output or a bounded observation. Write the type and transformation beside every arrow.
Search for append, emit, commit, project, convert, transform, prune, summarize, authorize, approve, and finalize. Those verbs reveal state transitions that a top-level loop may hide behind services. Read tests around the transition; they often state invariants more precisely than prose.
§4
Read the unhappy path before admiring the abstraction
Ask what happens when the model streams half a tool call and disconnects, a user cancels during execution, approval arrives after the state changed, a tool returns fifty megabytes, compaction fails, a child agent times out, or persistence succeeds only partly. Failure paths reveal who truly owns lifecycle and cleanup.
Look for broad catches, default-success fallbacks, swallowed errors, and state changes that happen before authorization. Then look for idempotency or reconciliation: can the run resume, retry, or establish whether an external effect already occurred?
§5
Printable inspection checklist
One-page worksheet
- Repository and pinned revision:
- Concrete user scenario:
- Turn entry point and main loop:
- Internal message/event types:
- Provider adapter boundary:
- Tool registration → validation → authorization → execution path:
- Context construction and token budget:
- Durable state and replay boundary:
- Failure, cancellation, and retry semantics:
- Extension and delegation contracts:
- Trace identifiers and post-effect evidence:
- Facts verified in source versus interpretations still open:
Complete the worksheet before drawing conclusions from feature names. It becomes the anatomy ledger you can compare across systems.