Context is starting to mean much more than prompt wording. A few projects this month made that more concrete.

Improving behavior without retraining the model

The Agentic Context Engineering paper proposes improving an AI system by changing the context it receives instead of retraining the model.

The post claims that this makes fine-tuning irrelevant, which is too broad. Updating context and retraining a model solve overlapping but different problems. Still, context is much cheaper and faster to change, so teams can test improvements there before committing to a fine-tuning cycle.

Letting a coding agent see the browser

A practical setup connects Claude Code to Chrome DevTools through MCP.

It closes a familiar gap in frontend work. A coding agent can edit CSS or application code, but without browser access it cannot inspect the rendered page, read console errors, or interact with the result. With a browser, it gets a feedback loop much closer to the one a developer uses.

There is plenty to scope carefully. Browser state can contain credentials, private data, and a lot of DOM information, only some of which is relevant and useful.

A small, readable model pipeline

Andrej Karpathy released nanochat, a compact repository that covers more of the model lifecycle than nanoGPT’s pretraining focus.

Production training stacks serve a different purpose. Repositories like this offer a readable route through training, post-training, inference, and chat serving without dropping the reader into a framework built for scale. That makes nanochat handy both as an educational reference and as a place to try small changes. It is valuable because it is small enough to understand.

Recursion for long inputs

The work on Recursive Language Models explores letting a model break a large input into smaller pieces, work through them, and combine the results.

Instead of putting every token into one huge context window, the model can split up a long input, inspect subproblems recursively, and combine intermediate results. That can save context, though it also introduces another route for errors to propagate. For debugging-sensitive work, I think we would want a trace of the decomposition and intermediate outputs before we can trust the final answer.

Loading procedures only when they are needed

Anthropic introduced Skills across Claude, Claude Code, and the API. Skills package instructions and supporting resources so they can be loaded as needed. A general agent can have access to many procedures without carrying every one in every request.

What triggers a skill and how does one test its behavior? How are changes to skills managed? Version control? Lots to unpack and understand.

Inspecting an agent run

Gang Rui Lim’s viewer for Claude Code traces tackles the other side of the problem. It displays subagents, tool calls, and tool results. A final diff cannot tell you why an agent picked a file, what it searched, which command failed, or how much context it burned before changing course. A viewer makes the run debuggable and can help compare agent-runtime changes by execution behavior rather than pass/fail alone.

Documents as visual context

The DeepSeek-OCR paper and implementation represent documents as images so more of their content can fit into the model’s working context. The paper calls this optical context compression, and the linked post notes an implementation running on vLLM.

Document images preserve layout, tables, and spacing that plain-text extraction may lose. The reported performance still needs comparison on the same documents and downstream tasks.