Looking inside vLLM
Aleksa Gordić published an in-depth walkthrough of vLLM, covering the anatomy of a high-throughput LLM inference system.
vLLM often sits in architecture diagrams as a black box between a model and an API. Inside that box are request scheduling, memory management, batching, and the choices that shape throughput and latency under real traffic.
A code-oriented walkthrough shows why an inference service behaves one way during a single local test and another under concurrent load. It also makes it easier to separate model latency from serving-system latency.
Hallucination as an incentive problem
An OpenAI paper on why language models hallucinate argues that common evaluations can give more credit for a plausible guess than for saying “I don’t know.”
This frames hallucination as an incentive problem. If an evaluation gives more credit for a plausible answer than for admitting uncertainty, the measurement system permits confident guessing.
This carries straight into product work. Whether a model answers or admits uncertainty depends partly on the examples, scoring rules, and interfaces around it. Better models help, but a test still has to distinguish a useful answer from an unsupported one.
Optimizing a multi-agent RAG pipeline
Isaac Kargar published a worked example using DSPy and GEPA to optimize a multi-agent RAG system.
The example treats the surrounding program as something measurable and improvable. Rather than nudging a prompt until a handful of examples look better, it tests changes to instructions and strategies against a defined score.
Someone still has to choose representative examples and define what “better” means. The method offers a way to test changes across a multi-step AI system instead of treating the prompt as its only moving part.
Hooks in Claude Code
Anthropic documented custom tools and hooks in the Claude Code SDK, along with updated references and usage guides.
Hooks are named points where the system can validate an action, add context, record evidence, or stop a run. Pre-commit checks, tool-call logging, and blocks on selected shell commands are all natural fits.
Custom tools and hooks expose more of the agent runtime without requiring a separate orchestration layer for every extension. The question is how much behavior can be added cleanly before the hook system turns into another hidden control plane.
12 Factor for Agents
HumanLayer’s 12-factor agents framework takes on the same control problem from a different direction. Its principles treat agents as software systems with control flow, state, tools, and human interaction.
The checklist is useful mainly because it makes the difficult questions hard to skip. Where does state live? Who controls the loop? What happens when a tool fails? When does a person approve an action? Most agent demos one sees these days don’t address all of these open questions.
GPU costs are application concerns too
A dstack report on the state of cloud GPUs in 2025 covers costs, performance, and playbooks for choosing providers.
Infrastructure economics decide which experiments are affordable and which models can be served. A small difference in latency or utilization can matter once traffic is steady. The report is a useful reminder to benchmark the workload itself instead of making decisions primarily based on GPU hourly prices.