Ralph and the appeal of a simple loop

Matt Pocock’s guide to Ralph describes a coding agent running the same prompt repeatedly, selecting tasks from a PRD, and committing after each feature.

The agent reads durable task state, does a bounded piece of work, records the result, and starts again. Repetition and external state provide autonomy without one enormous prompt.

That simplicity also makes the requirements obvious. Tasks need clear acceptance criteria, the repository needs tests, and commits or diffs need to be reviewable. Give the loop a vague PRD and weak tests and it will keep producing evidence of both.

Coding agents on remote VMs

Blackbox announced an Agents API for running several coding-agent CLIs on sandbox-backed remote VMs through one API. The announcement named Blackbox CLI, Claude Code, Codex CLI, and Gemini CLI, with the VMs powered by Vercel sandboxes.

The API packages environment provisioning. Applications still own tenant isolation, secret handling, network policy, quotas, and a record of what the agent changed.

Letting the agent look at localhost

The open-source Browser Use CLI and skill gives Claude Code and Codex a browser, including a headless mode that can inspect local applications.

This fills a gap in frontend work. An agent can change a page, open it, observe the result, and keep going. Without that feedback loop, it is effectively making visual changes blind.

A scoped development profile seems like the sensible boundary here, rather than a person’s everyday browser. Access to the local application and test credentials is usually enough.

Karpathy’s workflow shift

Andrej Karpathy wrote that his workflow had moved from roughly 20 percent agent coding in November to roughly 80 percent agent coding, with the remainder consisting of edits and touch-ups.

This is a single-workflow anecdote. In his description, most of the work consists of specifying, watching, reviewing, and correcting. How well that works depends heavily on the task and the repository. Suspect that codebases or repos that are already well designed, modular, include and use proper test frameworks will be a better fit for more hands-off or autonomous coding than the ones that don’t have these.

Testing repository instructions

Vercel compared ways to keep agents aligned with a project’s exact framework version and reported that its AGENTS.md approach scored 100 percent on its Next.js evaluations.

The result is specific to Vercel’s setup and tests, but testing the guidance is more useful than treating it as a style preference. Repository instructions, skills, and framework documentation can all be compared on actual tasks. In these tests, changing the repository guidance changed the result without changing the model.

Teaching a model to break work into smaller steps

The RLM team released RLM-Qwen3-8B, a small model trained to break larger tasks into smaller model calls. The team also tested it on task types that were not part of its training.

The model learns this behavior during training instead of relying on an external loop added at runtime. The release includes the model weights so others can run and test it.

Enterprise data agents need more than SQL

OpenAI published details about an internal data agent operating across more than 600 petabytes and 70,000 datasets.

The system relies on table-level knowledge plus product and organizational context. Natural-language-to-SQL is only a small part of enterprise data work. At 70,000 datasets, the agent also has to know which table is authoritative, how definitions differ, and which access rules apply. That catalog becomes critical.