Longer runs need better checks
OpenAI updated its GPT-5.4 prompting guide with patterns for tool use, structured outputs, verification loops, and long-running agent workflows.
The announcement calls out verification loops and long-running workflows. There is guidance on the interaction patterns OpenAI recommends for GPT-5.4.
Andrej Karpathy’s autoresearch makes one such loop unusually easy to see. It is a self-contained repository built around roughly 630 lines of a single-GPU nanochat training core.
At roughly 630 lines, the training core lets an agent edit the code, run an experiment, read the result, and choose what to try next without navigating a large research platform. The repository leaves changes, metrics, failed experiments, and the next attempt in view. This looks like a very cool project and tool for trying a lot of small changes quickly.
Automatically changing the agent runtime
A DAIR.AI post describes research on automatically generating agent runtimes, including the glue code that connects a model to tools, code execution, filesystems, and APIs.
This lets the system change more than the prompt. The runtime decides which actions are available, how tool results are shown to the model, what state persists, and when the model receives feedback. Changing that code can change behavior without changing the model.
Generated control logic deserves careful review. An agent runtime can raise a task score while granting more permissions or making failures harder to see.
The OpenDev paper approaches terminal coding agents from the systems side, devoting 81 pages to runtime code, the main agent loop, context management, and how the components fit together.
That treatment is welcome because a terminal UI can make a coding agent look much simpler than it is. Behind the prompt are repository discovery, command execution, context selection, patch generation, testing, and error recovery.
Two different ways to rethink browser agents
Alibaba’s open-source Page Agent uses JavaScript to embed a GUI agent directly in a webpage, without an external headless-browser or screenshot stack.
Because it lives inside the application, the agent can work with application structure rather than screenshots. This will not fit arbitrary sites where the developer cannot add the agent, but for teams that control their own frontend, it removes the external screenshot and headless-browser layers.
Swapping out memory after compaction loses something
Peter Steinberger pointed OpenClaw users toward the QMD memory plugin when context compaction discarded useful state. The plugin boundary allows the memory behavior to be replaced without replacing the rest of OpenClaw.
A different browser approach surfaced in a March post about Lightpanda. It is a headless browser designed for automation, and the post claimed 11x higher speed and 9x lower memory use than Chrome. The summary does not identify the benchmark behind those figures.
A runtime built for automation can omit parts of the desktop-browser product and prioritize startup time, density, and protocol compatibility. Compatibility is what will decide its usefulness, though. Browser workloads are full of edge cases, and an efficient runtime still has to execute the pages agents actually encounter.
Guidance at the moment of action
Clare Liguori reported an evaluation of five agent-guidance approaches across 3,000 total runs. Strands steering hooks were the only approach to reach 100 percent accuracy. The 3,000 figure covers all five approaches, not just the steering-hook runs.
The result is specific to that evaluation, but the mechanism is practical. Static instructions arrive before the system knows what the agent is about to do. A hook can inject narrow guidance when it becomes relevant.
It also creates a place for action-specific checks. File writes, shell commands, network requests, and external messages can each receive different guidance instead of sharing one giant system prompt.