Foundation · 9 of 9
Comparing harness designs
A feature checklist rewards breadth. An architecture comparison asks where control lives, what evidence supports the claim, and which product goal shaped the tradeoff.
§1
Compare only what the evidence can carry
DeepSeek Harness and Pi are open-source codebases in this study, pinned to exact commits. Their rows can describe source implementation. Codex and Claude Code are products whose current official documentation exposes behavior and configuration surfaces but not the complete internal implementation.
For Codex and Claude Code, this chapter says DOCUMENTED when official pages state behavior and UNKNOWN when the implementation is not public evidence. Similar product behavior does not establish a shared algorithm or code structure. Boundary
§2
One loop, three vocabularies
The useful invariant is a request cycle, not the word “turn.” DeepSeek calls one model request plus its tool work a step inside a turn. Pi calls that unit a turn inside an agent run. Keep the topology fixed, then change the labels.
Abstract loop
- Loop
while continuation is owed- Request
build_model_request(state): instructions, visible history, tool schemas and call settings.- Model boundary
call_model(request)crosses the provider adapter and returns an assistant proposal.- Tool branch
validate_authorize_execute(calls), record results, then begin another request cycle.- Final branch
- Finalize only when tools, queues and policy owe no continuation.
DeepSeek Harness
- Composition
- Cordis plugins install services, typed event listeners and reversible effects.
- Loop and state
ctx.agentLoopdrivesReactLoopAgent;ctx.sessionsstores durable facts while live Cordis events coordinate work in flight.- Request
ctx.systemPromptand tool schemas assemble first;agent/pre-stepadmits messages;deriveMessages()projects history;agent/requestadjusts call configuration.- Provider
ctx.llm.streamcrossesllm/streamto the selectedLlmAdapter.- Tools
tools/pre-execute→ctx.approvalwhen asked → guards →tools/execute→tools/post-execute.- Vocabulary
- One request cycle is a step; zero or more steps form a turn.
Pi
- Product shell
AgentSessionowns coding-product resources, model/auth, retries, compaction, extensions and session UX.- Loop and state
Agent.prompt()/continue()callrunAgentLoop()/runAgentLoopContinue(), then sharedrunLoop(), overAgentContext.messagesand anAgentEventstream.- Request
transformContext→convertToLlm→{systemPrompt, messages, tools}.- Provider
- The injected
streamFunction(model, llmContext)is the core seam; Pi AI supplies defaults and the coding layer may replace them. - Tools
prepareToolCall→ argument validation →beforeToolCall→AgentTool.execute→afterToolCall.- Vocabulary
- One request cycle is a turn; one agent run may contain several turns.
These labels describe ownership in the pinned implementations, not identical package boundaries. DeepSeek distributes the loop across Cordis services and live/durable event planes. Pi separates an injected reusable loop from the coding product that prepares and may continue the run. Sources
§3
The bounded comparison ledger
| Dimension | DeepSeek Harness | Pi | Codex | Claude Code |
|---|---|---|---|---|
| Evidence available here | SOURCE Pinned code and repository docs. | SOURCE Pinned code and repository docs. | DOCUMENTED Official product docs; internals UNKNOWN. | DOCUMENTED Official product docs; internals UNKNOWN. |
| Visible architectural emphasis | Plugin-composed Cordis services, events, scopes, and replaceable capability seams. | Three separable packages: provider API, reusable agent core, and coding-agent integration. | CLI/agent workflow with documented approvals, sandbox, project instructions, skills, and subagents. | Documented tools, permissions, hooks, skills, subagents, and context-management surfaces. |
| Context surface | Source exposes session projection, token metering, compaction, pruning, and spill components. | Source exposes message transformation, session history, read/output limits, cut points, and compaction. | Exact internal selection and compaction implementation is UNKNOWN in this primer. | Official docs describe context occupancy and documented compaction behavior; exact summarizer internals remain UNKNOWN. |
| Tool and permission surface | Source separates runtime, policy decisions, approval, sandbox/capability providers, and effect handling. | Core/coding layers expose tools and extensions; the pinned README explicitly says no built-in permission system. | Official docs describe approval and sandbox controls. | Official docs describe allow/ask/deny rules, permission modes, tool names, and pre-tool hooks. |
| Extensibility surface | Cordis plugins and capability seams can compose or replace major services. | Extensions, custom tools, resources, skills, and library-level reuse. | Officially documented project instructions, skills, and configurable subagents. | Officially documented hooks, skills, custom subagents, tool and permission configuration. |
| Version warning | The pinned README calls the project a developer preview and expects compatibility-breaking changes. | The former repository URL redirects to the current canonical repository; this primer uses one pinned commit. | Current product documentation can change; checked 2026-08-14. | Current product documentation can change; checked 2026-08-14. |
§4
Scope explains many apparent gaps
DeepSeek Harness presents itself as an extensible harness architecture. Its many seams make composition and replacement explicit, while increasing the number of lifecycle and policy contracts a reader must track. Pi exposes a smaller reusable core and a coding-agent product layer; it is easier to follow end to end and leaves permission confinement to the embedding environment.
Codex and Claude Code are compared as product surfaces because that is what public evidence supports. Their documented controls cover approval, sandboxing, instructions, skills, hooks, and child contexts. A product control does not prove a particular internal class, event store, or compaction algorithm.
“Has more built in” is not a universal score. A library may deliberately leave security policy to its caller. A product may bundle it for an interactive workflow. A plugin kernel may optimize for replacement. Ask whether the boundary is explicit and appropriate for the deployment.
§5
What the public product docs do establish
OpenAI’s current Codex documentation describes a CLI agent workflow, approval and sandbox controls, hierarchical project instructions, reusable skills, and subagent workflows whose results return to the main task. This supports a product-surface comparison, not a private implementation claim. Documented
Anthropic’s current Claude Code documentation describes named tools, permission rules and modes, hooks, skills, fresh-context subagents, and context-window behavior. The internal dispatcher, message serialization, and exact summarizer prompt remain outside this primer’s evidence. Documented
§6
Questions to take into a design review
- Which layer is the product: library, embeddable runtime, interactive coding agent, or plugin platform?
- Can one turn be reconstructed from durable events and model/tool inputs?
- Where are registration, authorization, confinement, and verification separated?
- How does context policy expose loss, retrieval, and durability?
- What can an extension replace, and which invariants remain enforced by the host?
- How do child agents inherit context, tools, and authority?
- Which component compensates for a current model weakness, and what evidence would justify keeping, changing, or retiring it?
- Which claims come from source, official documentation, observation, or interpretation?
The next two tracks apply these questions to DeepSeek Harness and Pi at pinned commits. Read either one independently; read both to see how the same anatomy can produce very different code shapes.