Finding a tool before using it

Anthropic describes a pattern for discovering tools just in time. The agent searches a tool catalog instead of receiving every tool definition in its context.

This starts to matter once schemas take up enough context to interfere with selection. A large, undifferentiated list costs tokens and gives the model more nearly-correct choices. Search separates finding a capability from invoking it.

It also changes the failure mode. The discovery function has to retrieve the right tool, and the agent needs enough context to describe what it needs.

Gemini 3 arrives with broad product ambitions

Google announced Gemini 3 in mid-November and positioned it as the basis for new product experiences across Google. Jeff Dean’s announcement emphasizes that reach, but does not identify which engineering workloads improved most. That leaves the practical comparison open for now.

From a local coding tool to a remote service

The open-source Claude Agent Server packages the runtime behind Claude Code in a cloud sandbox and exposes a WebSocket control interface.

That bridges an interactive local tool and an asynchronous service. When agent runtime moves off the developer’s machine, it addresses isolation but then one has to figure out how to properly handle persistence. The runtime needs to handle event streaming, provide ways to track progress, ability to abort/cancel etc. Solves some issues but introduces a new set of considerations.

Some others to think through: authentication, tenancy, resource limits, and cleanup.

Productionizing AI Agents

A Google whitepaper on deploying, scaling, and productionizing AI agents picks up from there. The paper focuses on the route from prototype to production. For a product, monitoring, evaluation, failure handling, deployment, and cost all become major concerns. There’s lots of recommendations, not everything will transfer or land cleanly for everyone. In the messy real world, things always fail in different and novel ways. Nonetheless, there is some valuable information in there and it’s a great resource to come up with your own checklist.

Files as working memory

Harrison Chase wrote about agents using filesystems for context engineering.

Files are an old abstraction with convenient properties: they are hierarchical, persistent, selectively readable, diffable, and first-class citizens for programming tools and operating systems. An agent can put plans, intermediate results, retrieved documents, and handoff notes there without stuffing them all into the current prompt.

An agentic system can decide what to remember, where to store it and when to retrieve it. Filesystem gives that agentic system a legible place for state, which may be enough to avoid reaching immediately for a specialized memory service.

Models framed around agent work

Anthropic released Claude Opus 4.5 with explicit claims around coding, agents, and computer-use workloads.

Anthropic’s claims still need testing on representative work. Its positioning centers coding, agents, and computer use alongside chat. I think models cannot be compared or benchmarked in isolation, the surrounding system can easily matter as much as the model change.

Prime Intellect announced INTELLECT-3, a mixture-of-experts model that routes work among specialized components. It was trained with reinforcement learning using the company’s task environments, testing tools, training framework, and sandboxes.

The model itself is only one piece of that release. Training an agent this way also requires places to run tasks, infrastructure for running many attempts, scoring rules, and tests. Prime Intellect includes those pieces. The stack looks very interesting, would love to dig into it if I can find some spare cycles.