Routing tasks between models
Cognition’s Fable and Sidekick results are a useful example of routing different parts of an agent run to different models. Cognition reports that its paired setup costs 54% less than using Fable alone while producing a similar score.
The useful part for me is making model selection part of the orchestration logic. Routine steps can go to a cheaper model, while harder decisions can still use the stronger one. The cost and quality of the completed task are the measures that matter.
Building tests from real traces
LangChain’s Eval Engineering Skill reads the codebase and existing agent traces, suggests behaviors worth testing, asks the user which ones matter, and creates executable Harbor tasks.
Starting with real traces makes sense. They show where the agent was confused, brittle, or unsafe in actual use. A human still has to decide which behaviors matter, but the tool can do much of the work required to turn those examples into repeatable tests.
Separating agent behavior from runtime infrastructure
Ankur Goyal describes an agent runtime as two layers: agent behavior, such as context compaction and delegation, and runtime infrastructure, such as event storage, sandboxes, and turn coordination.
This seems like a useful design boundary. A team should be able to change how an agent manages context or delegates work without rebuilding storage, execution, and isolation. It also makes failures easier to separate. A poor decision by the model is different from losing an event or allowing a sandbox to leak.
Rebuilding an AI-generated system
The exe.dev team says it built a distributed DNS server in about a week, rebuilt it twice on purpose, shipped it, and saw no DNS incidents during the following month.
While the one-week timeline is impressive, the two rebuilds are actually more interesting. AI is making code generation cheaper, so throwing away a weak or shaky implementation becomes a viable engineering choice. Review remains necessary because a team still has to understand the system and be willing to support what it ships.
Managing agent capabilities as dependencies
Drew Breunig’s drskill manages the skills and MCP servers available to an agent as a loadout.
I like that framing because these capabilities are dependencies. They can add tools, instructions, code, permissions, and context. Installing everything globally creates more security, maintenance, and context-management work, so it is worth keeping an inventory and selecting what each agent needs.
Running MCP servers without session state
The 2026-07-28 MCP update makes the protocol stateless, which should simplify how remote MCP servers are deployed and scaled.
For teams running MCP as a service, stateless requests fit existing infrastructure more easily. Replicas do not need to coordinate protocol sessions, and failed instances can be replaced without recovering that session state. This should reduce the operational complexity for MCP Server operators.
Agent runtimes as shared internal platforms
Y Combinator open-sourced QM, a customizable multi-agent runtime it uses across accounting, legal, events, and engineering.
Apparently the same runtime is being used outside core software development. Looks like a promising agent platform to deploy and also interesting to study for architectural insights. Would love to do a pilot with it.