There is a lot of activity around agent harnesses in general. Industry is continuing to address challenges around state, isolation, shared infrastructure, evaluation, cost, and context management.

DeepSeek Harness Design

DeepSeek Harness makes models, tools, skills, sessions, sandboxes, storage, scheduling, and the user interface replaceable components via a plugin design.

Its open implementation gives AI Engineers and developers building AI-enabled applications a concrete system to inspect. You can trace how the harness assembles model context, controls tool use, records outcomes, and decides whether another model call is needed. Those design choices are what turn repeated model calls into a controlled application workflow.

For a deeper look, Understanding Agent Harnesses examines these principles and includes source-guided tours of DeepSeek Harness and Pi.

Coding agents and Enterprise Deployments

According to Spotify, Xirp “is a vendor-neutral agentic development environment born out of a concrete engineering need.” It is designed to manage sessions across Claude Code, Gemini CLI, and Codex. Its capabilities are enhanced by connecting it with Portal, which serves as a catalog for connecting agent sessions with component architecture, system design, and other contextual and governance information. Spotify reports thousands of internal users and more than 36,000 sessions.

Xirp treats coding agents as an internal engineering platform. Once agents are used across a company, teams need identity, permissions, organizational context, isolated workspaces, observability, and administration. The model and coding agent remain relevant but are essentially fungible in this world. The value comes from the enterprise tooling built around them to make them usable across teams and organizations.

AgentENV supports Kimi K3’s post-training

Moonshot AI’s Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model with native vision and a one-million-token context window. Along with the model weights and technical report, Moonshot open-sourced some of the infrastructure used to train it.

AgentENV is an open-source sandbox platform developed by Moonshot AI in collaboration with KVCache.AI. During Kimi K3’s reinforcement-learning post-training, the model proposes actions, those actions run inside isolated Firecracker environments, and the results are returned to the model for the next step. Snapshot, resume, and fork support allow many of these training runs to proceed in parallel. AgentENV does not train or serve the model itself; it manages the environments in which the agent acts.

Testing whether harness improvements carry over

The HarnessCompass paper asks whether an automated system can improve an agent harness in ways that still work on tasks it has not seen. HarnessCompass reviews agent runs, proposes changes to prompts, tools, middleware, or memory, and rejects changes that contain clues tied to a particular repository, test, or task.

The researchers used 50 SWE-bench Verified tasks to improve the harness and kept another 450 tasks out of that process. With the same GPT-5.4 base model, the updated harness increased the success rate on those unseen tasks from 51.6% to 60.4%. The useful discipline for an AI Engineer is to validate a harness change on different repositories and tasks, not only the examples that exposed the original problem. That helps distinguish a broader improvement in agent behavior from a workaround fitted to a few known cases.

Better context engineering can mean less context

Anthropic says it removed more than 80% of the Claude Code system prompt for newer models with no measurable loss on its coding evals. Its article on context engineering for newer Claude models recommends shorter tool descriptions, progressive disclosure, and moving conditional procedures into skills rather than loading them into every request.

More instructions do not automatically produce more control. Old, duplicated, or conflicting guidance can make the model less predictable and consume context that could be used for the task. This also makes a good case for treating prompts, skills, and tool descriptions as maintained software rather than an append-only collection of instructions.

Memory as a maintained shared artifact

Chroma describes Foundation as a system that consumes coding-agent sessions and company data to build a durable, wiki-like record. Its design includes versioning, lineage, citations, access control, and support for concurrent updates.

The interesting memory problems here go beyond storing and retrieving vectors. The system has to decide what should survive, show where it came from, and allow several agents to update shared knowledge without corrupting it. Chroma’s transaction design also accounts for the cost of throwing away agent reasoning when concurrent writes conflict, instead of assuming that aborting and retrying is cheap. Chroma uses what it calls optimistic concurrency control (OCC) to support atomic chunk updates in its distributed storage system.