Keeping an eye on several agent runs is already tedious. One pane is waiting for approval, another has failed quietly, and a third finished ten minutes ago. A few recent updates and releases are aimed squarely at those challenges.

Memory that keeps observing

Mastra released Observational Memory and reported LongMemEval scores of 84.2 percent with GPT-4o and 94.9 percent with GPT-5-mini. Mastra says the system updates memory continuously instead of waiting for each query to search a separate memory store.

The system keeps a running record of what has happened instead of treating memory as a database searched only when a question arrives. The benchmark numbers are Mastra’s; a useful comparison would run the same model and LongMemEval setup against a system that retrieves stored memories on demand.

Memory quality obviously depends on first what the system chose to preserve and then eventually on what the final model can recall.

Watching how agent teams work

A command in claude-code-templates renders Claude Code Teams activity as a flow diagram, including lead-agent communication, tasks, messages, and tool use.

This feels better matched to multi-agent work than a flat transcript. Once a lead agent delegates across a team, the useful questions are about relationships: who received a task, which tools ran, and where a message went. A graph can put those in one view.

Ramp’s Inspect agent shows the same observability problem at a larger operational boundary. A post about the agent reported that it accounted for 57 percent of merged pull requests during one cited 24-hour period. It runs in sandboxed Modal VMs and can access Sentry, Datadog, and GitHub.

The 57 percent figure covers only that 24-hour period. The architecture pairs sandboxed VMs with real engineering context. That access needs careful permissions. Reading an error trace and opening a pull request are different authorities from changing production or merging code.

OpenAI also reported that a small team steering Codex opened and merged 1,500 pull requests for an internal product without manually writing the code.

OpenAI describes the work as steering Codex and engineering the runtime around it. The runtime is part of the result. Abundant generated code makes review capacity and tests the limits on how much can merge.

Sounds fantastic! Clearly very few organizations are operating at this frontier, there are many that have not even fully adopted coding agents in their developer workflows yet. Huge divergence of organization abilities.

Automatically improving skills

The gskill project uses GEPA to automatically rewrite and test agent skill files. Its authors reported that the generated skills solved nearly all of their repository tasks and ran 47 percent faster in their evaluation.

The reported result is tied to the authors’ repository tasks. The project treats skills as programs with measurable behavior rather than static documents improved through intuition alone.

There is a tradeoff, though. Automated rewriting can produce instructions that score well while becoming hard for a person to understand or maintain. A useful generated skill should improve the score and remain readable. One day I will have more time to run lots of experiments with DSPy & GEPA!

Small interface details for busy operators

The open-source cmux terminal launched with vertical tabs, a built-in browser, and visual indicators showing which coding-agent pane needs attention.

That indicator addresses a real issue if you have multiple sessions running simultaneously. Several asynchronous terminals can finish, block, or fail at different times, and manually polling each one gets painful very quickly.

Browser agents that keep running

Browser Use added continuously running agents with scoped-authentication profiles, file workspaces, and integrations such as Slack and Linear.

A persistent service needs tighter boundaries than a one-off browser session: which profile it may use, which files it may open, and how to stop it. The announced profiles and workspaces make some of that scope visible. A nice thing to add would be a record of what ran while nobody was watching.