§1

Look for roles, not one canonical diagram

Most tool-using harnesses must solve a recurring set of problems: speak to a model provider, assemble model-visible context, run a turn loop, expose and execute tools, apply authority, retain state, and offer integration seams. One repository may place each role in a package; another may combine several in one session class. Interpretation

The map below is therefore functional rather than prescriptive. A box says “some part of the system must own this contract.” It does not say every harness should have a directory with the same name.

MODEL EDGE

Adapter + messages

Translate the harness’s internal vocabulary into a provider request and normalize the streamed response.

CONTROL CORE

Context + loop

Choose the working set, schedule inference and tools, and represent terminal and interrupted states.

EFFECT EDGE

Tools + authority

Describe capabilities to the model and mediate proposals before they touch an environment.

DURABLE EDGE

State + events + extensions

Preserve a record, derive views, expose observability, and admit controlled variation.

These four boxes group contracts by boundary. The control-plane view later in this chapter regroups the same roles by when and why they act. They are two projections of one anatomy, not competing subsystem lists.

§2

The twelve-part anatomy ledger

ComponentRoleInputOutputOwnerNot the same as
Model adapterTranslate provider-independent requests and streamed events.Internal messages, tools, model settings.Provider request; normalized assistant events.Harness integration layer.The model itself or the turn loop.
Instruction assemblyResolve system, project, task, skill, and runtime instructions with precedence.Configuration, files, task metadata, policy.Ordered model-visible instruction content.Context/configuration layer.Conversation history or authorization.
Message modelRepresent user, assistant, tool, control, and custom events without losing meaning.UI events, provider events, tool results.Typed internal messages and provider-ready projections.Core runtime.The displayed transcript.
Tool registryName capabilities and publish descriptions plus argument contracts.Built-ins, plugins, extensions, current scope.Lookup table and model-visible tool definitions.Capability layer.Permission to execute a registered tool.
Execution loopAlternate model calls, tool handling, state transitions, and stopping.Turn request, state, events, budgets.Events, effects, terminal or suspended outcome.Control core.A model-generated plan.
Context managerConstruct the bounded working set for the next inference.Transcript, instructions, summaries, retrieval, budgets.Ordered model request content.Control/context layer.Durable state or the full transcript.
State storeRetain session events, artifacts, settings, checkpoints, and derived records.Committed runtime events and external references.Replayable history, snapshots, or queryable views.Persistence layer.The model’s context window.
Permission layerDecide whether a proposed capability may run under current authority.Identity, request, arguments, scope, rules, approvals.Allow, deny, ask, or constrained execution decision.User/operator policy enforced by harness.Schema validation or sandbox confinement.
Event streamExpose ordered lifecycle facts to UI, logs, plugins, and projections.Model, tool, policy, state, and control transitions.Typed events with ordering and identifiers.Runtime/observability layer.Plain-text logs or user-visible chat alone.
User interfaceCollect input and approvals; render streaming progress, tool activity, and outcomes.Events, session views, user actions.Commands, cancellations, approvals, displays.Product surface.The agent runtime, even when bundled together.
DelegationCreate bounded child work with defined context, tools, authority, and return channel.Task description, parent state, policy, child configuration.Child events, artifacts, summary, or result.Scheduler/subagent layer.A model merely suggesting subtasks in prose.
Extension seamsAllow supported variation in tools, hooks, services, providers, prompts, or lifecycle.Plugin/extension definitions and host contracts.Registered behavior within a declared lifecycle.Host architecture plus extension author.Arbitrary monkey-patching or undocumented internals.

§3

The loop is small; its control plane is not

The request-cycle kernel can remain understandable while explicit layers around it own context, policy, durable state, coordination, observability, and cross-run recovery. This is another view of the same twelve responsibilities, not a thirteenth component or a required package layout. Interpretation

Small kernel, explicit layersKeep the model/tool protocol visible. Put production concerns in named contracts around it.
[5] Execution loopRequest-cycle kernel
  1. Build request
  2. Call model
  3. Inspect proposal
  4. Execute authorized tools
  5. Record and continue
01 · Request cycleOne inference and its immediate tool work
02 · Work cycleA turn or run with one or more request cycles
03 · Task lifecycleEligibility, retries, stalls, cancellation, reconciliation, and workspace ownership across runs

The scale boundary matters because retrying a failed provider call, continuing a tool loop, and rescheduling a stalled task are different decisions. OpenAI's Symphony specification is one concrete outer-control-plane design: a single orchestrator owns eligibility, bounded concurrency, retries, cancellation, and reconciliation around coding-agent sessions. Other harnesses may place the same responsibilities elsewhere.

§4

The model edge: preserve meaning across formats

Providers differ in message roles, tool-call formats, reasoning events, streaming protocols, usage accounting, and error shapes. The model adapter keeps those differences from infecting the whole loop.

A useful internal message model is usually richer than the provider vocabulary. It may retain UI-only messages, branch summaries, approval events, or custom tool results. At the inference boundary, a transformation selects and converts only the messages the provider can accept. That projection should be explicit: silently dropping an unsupported message can change what the model knows.

Instruction assembly belongs near this edge but is a separate responsibility. It decides precedence and ordering among product instructions, repository guidance, skills, task constraints, and temporary runtime notes. It does not grant authority. Text saying “never write outside this directory” is useful guidance; a path guard or sandbox is the enforcement mechanism.

§5

The control core and effect edge

The execution loop owns the protocol described in Chapter 2. The context manager feeds it a bounded model request. The registry tells it which tools exist in this scope. The permission layer tells it which proposed uses are allowed. The loop dispatches an allowed operation to the registered tool implementation and records its typed observation. “Runtime” is therefore implementation machinery inside these contracts, not a thirteenth anatomy role.

These roles may be adjacent, but merging their concepts creates security bugs. Registration answers “can the system resolve this name?” Visibility answers “was the model told this capability exists?” Validation answers “do the arguments match the contract?” Authorization answers “may this principal use it here?” Confinement answers “what can the process reach if the tool misbehaves?” Execution answers “what did the implementation attempt?” Verification answers “what effect can we now establish?”

A harness can expose fewer tools than it has registered, authorize a smaller argument range than the schema accepts, and run an authorized tool inside a still narrower sandbox. Those are layered controls, not duplication.

§6

The durable edge: events, views, and replay

Agent work unfolds over time. Streaming model deltas, complete assistant messages, tool proposals, approvals, results, cancellations, summaries, branches, and artifacts all need representation. An event stream provides a common observation surface.

Some systems store events as the source of truth and derive transcript or UI views from them. Others store messages and emit transient lifecycle events around them. Either design should answer: which facts survive restart, which projections can be rebuilt, what ordering is guaranteed, and how partial operations are represented.

Observability is downstream of this contract. A log line can aid debugging, but typed events with turn, step, tool-call, and parent-child identifiers support stronger questions: which prompt produced this call, which policy decision preceded it, which observation entered the next model request, and which child task returned this artifact?

§7

Extension design is about where change is allowed

Evaluate an extension system by the contracts it can alter and the lifecycle phases where it can act. Can it register a tool, add instructions, observe an event, veto execution, replace the session store, supply a sandbox, spawn a child agent, or run out of process?

Broader seams make more parts replaceable, but they also create more lifecycle and failure cases. A small callback API is easier to reason about but may force product-specific forks. A service/plugin kernel can replace major subsystems but requires clear rules for lifecycle, dependencies, scope, and failure under composition.

Created by Varma Chanderraju. Built with Codex, Claude and Gemini. · Glossary · Sources