跳到正文
PyTorch:Blog· Ishan Dhanani, Jamie Li, Karen Chung·· 2 小時前AI 評分69

NVIDIA Dynamo 以 session ID 支援智能體推理排程與 KV cache 管理

Session-Aware Agentic Inference with NVIDIA Dynamo

AI 導讀

NVIDIA Dynamo 以統一的 session ID 識別智能體工作階段,將推理服務由逐請求處理擴展至支援工作階段感知的路由、排程及 KV cache 管理。

正文

TL;DR

Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, with KV cache remaining resident while tools execute between turns.

This technical blog post details how NVIDIA Dynamo uses a unified, session level identifier to transform request-level serving infrastructure into a program-aware system, unlocking session-aware routing, shared KV cache indexing, and programmatic cache movement across vLLM and SGLang.

Introduction

Agentic workloads have changed the shape of the traffic an inference server sees. A coding session opens with a large prefill, often tens of thousands of tokens of system prompt, tool definitions, and user-specific guidance, and every turn after that appends to a context the session resends in full. One task fans out into dozens of such calls, alongside short and long lived subagents running in parallel. Most of a session’s wall-clock time passes between tool calls, with its context sitting resident in the KV cache while nothing generates. Serving these well comes down to how much of that context a system keeps resident and reuses, and how many sessions it holds at once while still meeting latency targets This contrast between agentic workloads and that of a standard chatbot is shown in Figure one.

Figure 1: Compares the alternating user and model turns of a standard chatbot with an agentic workflow that includes tool calls and tool responses. In the agentic workflow, a single user request can trigger multiple inference calls, with subsequent steps depending on intermediate results.

Most open-source serving stacks are still built for the model on top of the figure. They route, admit, and cache each request on its own, with no notion of the harness or the session that ties it to previous turns.

Serving this shape optimally is something we have been building towards. In our previous post, Full-Stack Optimizations for Agentic Inference, we shared work across three layers of the NVIDIA Dynamo stack: the router, the inference engine, and the KV cache manager. Over the last few months we’ve unified this work around a single primitive: a common session-level identifier. Session-level ID (often called program or trajectory in papers and academia) turns request-aware infrastructure into program-aware infrastructure, and serves as the foundation for additional session-aware optimizations.

This post covers that progress in four parts:

  1. Session_ID: How Dynamo recognizes sessions and subagents from existing harness headers, and how custom harnesses can provide the same information.
  2. Tracing and replay: How session-linked traces connect harness activity to inference performance and support offline simulation and live benchmarking.
  3. Session-aware scheduling: How routing and admission control account for session working sets and apply backpressure at tool boundaries to reduce cache thrashing.
  4. Shared KV cache awareness: How the experimental shared-pool indexer lets the router account for reusable KV in external stores when placing requests.
  5. Programmatic KV cache: How the Session-Prefix Indexer and proposed KvHint interface support session-aware cache policies across memory tiers, with Dynamo supplying policy intent and inference engines controlling execution.

Identify Agent Sessions with Session ID

A Dynamo session ID is a unique stable identifier for a single agentic chain of reasoning and tool calls. Every LLM request in a given trajectory shares the same session_id. Usually, child sessions carry a parent_session_id, so traces and replay tools can establish a connection between agents and subagents or subtasks they spawn. Every optimization in this post is an opt-in consumer of the session information contained in session ID and parent session ID

There are a few different ways to share a stable session identifier with Dynamo, depending on what’s calling it.

Out-of-the-box support for coding agents

Dynamo recognizes the identity headers that today’s popular coding agents already emit and maps each onto the same internal agent_context: the session header becomes session_id, and the parent header, where present, becomes parent_session_id. Claude Code, Codex, and OpenCode work without configuration.

Coding agent Session header Parent header
Claude Code x-claude-code-session-id (root), x-claude-code-agent-id (child agents) x-claude-code-parent-agent-id
Codex thread-id x-codex-parent-thread-id
OpenCode x-opencode-session-id x-opencode-parent-session-id

We’re working with harness providers across the ecosystem to make this the default everywhere, so support expands without users changing anything. Dynamo’s CI tracks each first-class harness as it changes. You can find a full list of supported first class IDs here.

Plugins for harnesses that don’t send session identity natively

For harnesses that don’t emit a compatible header on their own, the agent-plugins repo provides small integrations that translate a harness’s native session identity into x-dynamo-session-id. Three are available today: a Pi provider for Dynamo’s OpenAI-compatible endpoint, a Hermes middleware that maps Hermes’s session_id, and an OpenClaw provider that maps OpenClaw’s sessionId. Each plugin ships its own install, configuration, and validation steps in its own folder in the repo.

Custom harnesses: opt in with one header

A custom harness opts in with one canonical header, X-Dynamo-Session-ID. Dynamo normalizes it into an internal agent_context struct that the rest of the stack reads, so nothing downstream needs to know which harness the request came from. If the harness supports subagents, you can also provide X-Dynamo-Parent-Session-ID. The following code block shows how any custom agent can pick up Dynamo’s optimizations without adopting a proprietary API.

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer sk-dummy' \
  -H 'x-dynamo-session-id: research-run-42:researcher' \
  -d '{"model":"my-model","messages":[{"role":"user","content":"..."}]}'

Harness ↔ Inference Codesign

Agent observability tools can already capture what happens at the harness level, but they don’t yet show a unified picture of how the inference stack performed underneath each of those steps. We think about this gap as harness ↔ inference codesign: giving both sides a shared, measurable trace of the same request that can be used for performance visualization, replay, and offline simulation, all keyed by a constant session ID. This trace is meant to be lightweight. It can be used as additional optional metadata alongside other trace collection formats, like Harbor’s ATIF or Letta’s new Trajectory format.

Capturing request traces

To enable trace collection, run your deployment with DYN_REQUEST_TRACE=1. Dynamo automatically records information on each request, without storing prompt, response, or tool-call content by default (these can be opt-in enabled). Instead, Dynamo emits a request_end record after each response stream finishes, carrying the session fields, the output token count, the finish reason, KV cache metrics, tool-call names, and a replay block with the request’s input length and per-block sequence hashes, similar to the Mooncake trace format. The sequence hashes can be used for replay with tools like the built-in dynamo.replay or our load-gen tool, AIPerf.

Figure 2: Shows a screen capture of the dynamo trace, with separate tracks for agents and their tool calls. Prefill, decode, and tool execution appear on the same timeline, showing how work is sequenced and overlaps across the session.

Because every record carries timestamps and finish reasons, we can see tool-call distributions and connect them to KV cache eviction behavior. Internally, we’ve used these techniques to optimize our model deployments end to end, using captured traces as backtesting data. Much of this data and these replay tools also power our internal autoresearch harnesses, which let AI generate novel routing strategies and research proactive KV cache offloading strategies, which we cover below.

Replaying agentic traces

A captured trace is a reusable benchmark. Because it records the workload the agent sent rather than the agent’s decisions, one capture replays the same request schedule as many times as you want without rerunning the model, the tools, or any external API. Replay stays content-free the same way the trace does: sequence hashes become synthetic block IDs that preserve prompt-prefix sharing across turns and sessions without exposing the original tokens.

One capture drives two paths. Offline, AISimulate runs the request graph against Dynamo’s simulated scheduler, router, and KV cache, letting you compare worker counts, serving topologies, routing policies, and cache capacity without spending GPU time. Live, AIPerf replays the same graph against a real Dynamo endpoint on real GPUs for final latency, throughput, and cache numbers.

Routing and admission control at the session layer

KV cache has always mattered for efficient LLM inference, but agentic inference adds real complexity: where the cache lives, how much of it can be reused at each memory tier, and whether the system is proactive about prioritization, eviction, and sharing. Throughput ends up being determined less by raw compute and more by how well the stack answers those questions.

Because the Dynamo router is flexible enough to act on more than just the next request, we’ve been able to iterate on session-aware routing and admission control strategies that target these characteristics directly, improving performance for serving agentic-shaped workloads (including RL post-training). The goal is to reuse as much KV cache as possible and avoid thrashing: evicting and re-prefilling for every request because HBM, and eventually DRAM, can’t hold every agent’s working set at once.

The problem with request-level routing

A request-level router, Dynamo’s stock KvRouter included, solves placement one request at a time: given a prompt, which worker has the most cache overlap. That’s the right unit for a single turn, but it’s the wrong unit for an agent, and at high concurrency it misses two things about the agent behind the request:

  • Cache-occupancy blowup. Between turns, an agent’s cache doesn’t get freed. It sits in HBM holding blocks while the agent runs a tool call outside the GPU. With N agents concurrently at step k, the aggregate working set is N × context_k, and a request-level router has no notion of that aggregate. It only ever sees one turn’s context at a time. Once the working set exceeds HBM, and then DRAM, the engine starts evicting live sessions’ cache, and every next turn pays a full re-prefill instead of a cache hit.
  • No tool-boundary backpressure. When a worker is over capacity, a request-level router’s only levers are to cancel or queue in-flight requests. Both are worse than the alternative: deferring the session at the pause point right before its next turn arrives, when it’s about to be idle anyway.

Session Aware admission control

Our first agent-aware routing strategy ports the scheduler from the ThunderAgent paper (Kang et al., 2026) on top of Dynamo’s KV-aware router. Written in Rust, it owns a KvRouter directly rather than sitting in front of one as an extra proxy hop, and it tracks real prompt_tokens + completion_tokens off each response instead of estimating token counts from raw bytes. Because this is written as a native Dynamo plugin, you can take its source code and adapt it to your workload shape without changing Dynamo’s underlying router code.

The scheduler groups requests by program_id (the session ID) and moves each program through two independent states: (REASONING | ACTING) × (ACTIVE | PAUSED). A program enters ACTING at a tool boundary. Under memory pressure, the scheduler pauses ACTING programs logically. There’s no decode preemption; the in-flight turn is allowed to finish, which frees the engine to evict their KV blocks. If a worker has no ACTING programs left to pause, it instead marks the smallest REASONING program for pause at its next tool boundary. When a worker has headroom again, the scheduler resumes the smallest-token paused programs first, bin-packing them back under threshold.

The control loop

Pause and resume are both driven off one quantity per worker: utilization, the fraction of a worker’s KV pool occupied by its assigned programs’ working sets:

U_worker = (sum of token weights for ACTING programs on worker) / (worker's KV pool capacity)

Three thresholds (defaults shown) turn that number into a control loop:

  • U ≥ 0.95 (pause-threshold): the worker is over-subscribed. The tick pauses the smallest ACTING programs first until U falls back to 0.80 (pause-target).
  • 0.80 ≤ U < 0.95 (soft-demote band): programs aren’t paused yet, but take a -2.0 priority penalty as early backpressure before a hard pause is needed.
  • Resume only fires once U ≤ 0.85: that’s pause-threshold minus a 0.10 resume-hysteresis, a buffer below the pause line that keeps the loop from flapping between pausing and resuming right at the threshold.

Resume itself is a bin-packing pass: paused programs are sorted by token count ascending, and the scheduler admits them smallest-first until the next one would push utilization back over threshold, accounting for a fixed per-program buffer. Resumed requests get a one-second priority boost so they don’t get stuck behind fresh admissions, and a 30-minute forced-resume cap guarantees no program starves indefinitely.

Results

At moderate to high concurrency, we see a large increase in throughput. the difference shows up directly in throughput. On SWE-bench, running two TP4 MiniMax-M2 replicas on a single 8xH100 node, program-aware scheduling improved throughput roughly 12-16% over KV-routing alone, with the gap coming from the re-prefill work that request-level routing can’t avoid.

The same strategy shows a similarly large boost on agentic RL rollouts. Running Uni-Agent SWE-Bench on 8xH20-3e with Qwen3-Coder-30B-A3B-Instruct against VERL’s default Global LB, our strategy stayed roughly even at low concurrency (CC32-128), where nothing needed to pause. At medium concurrency (CC192-256) it pulled ahead, with 11.0-14.6% higher model-token throughput than Global LB, and at high concurrency (CC384-512) Global LB throughput dropped sharply while our strategy kept scaling. Across the full sweep it held prefix-cache hit rate above 94.5%.

Figure 3: Compares model-token throughput across agent admission concurrency levels for VERL Global LB, Dynamo Python TA, and Dynamo’s native session-aware scheduler. At high concurrency, the native scheduler maintains the highest throughput of the three, while Global LB throughput drops sharply.

Shared KV cache awareness in the router

KV cache placement gets harder once our working set outgrows HBM. The cache doesn’t disappear, it moves into CPU memory and then into shared stores like Mooncake or NIXL backed storage. So far the Dynamo router was only aware of GPU and native framework CPU offloaded memory. When KV moved into an external store, Dynamo was unable to factor it into the cost function itself. To solve this, we worked with the Mooncake team to enable KV event emission natively from the Mooncake Store and implemented a new experimental shared-pool indexer that works in conjunction with Dynamo’s blazing fast FlashIndexer. When KV is moved from HiCache (SGLang’s native CPU offloading mechanism) into Mooncake, a removal event is emitted from SGLang, and a store event is emitted from Mooncake. A page only counts as a hit when every physical object for the worker’s parallelism layout is present, and credit applies beyond its device-local prefix.

At request time, Dynamo hashes the prompt into SGLang KV pages and expands each logical page into the physical Mooncake objects implied by the worker’s TP/PP layout, K/V tensors, and optional backend tag. Mooncake publishes ordered Store and Remove events and Dynamo consumes them into a global shared-pool index, conservatively treating a page as reusable only when every required physical object is present. The index returns shared hits as contiguous page ranges rather than probing Mooncake on the routing hot path. FlashIndexer independently computes each candidate worker’s native device, CPU, and disk prefix overlap. For a worker with device-resident prefix depth D, the router credits only shared hits at positions >= D: shared_beyond_device = |shared_hit_ranges ∩ [D, request_pages)|. That prevents a GPU-local hit from being counted twice, while allowing a worker with a shorter local prefix to receive credit for the suffix already available in Mooncake. The final placement score combines weighted native-tier overlap, configurable shared-pool credit, and live worker load.

Towards Programmatic KV Cache

Ultimately, performant inference on agentic workloads requires session-aware KV block management. To predict the priority value of a given KV cache block in such workloads, we need complementary information from the router and the engine. The router is aware of session lifecycle and reuse patterns across workers; the engine is individually aware of the KV cache status it holds. Hence, for engines (vLLM, SGLang) to execute a maximally informed cache policy, (1) Dynamo should express policy intent, and (2) the engine should resolve that intent against exact cache state and execute it.

This is the central idea of programmatic KV cache. Essentially, the Dynamo router provides the “brain” that calculates smart, session/load/cache-aware KV movement policies, while the engine “executes” the movements that Dynamo mandates. As such, we have been developing features for both Dynamo and vLLM/SGLang: the machinery for Dynamo to calculate+emit KV policies, and the first-class abstractions in the engines required to lower Dynamo’s session-based policies into concrete, in-engine KV movements.

Router-Driven KV Hints Policy

Dynamo’s programmatic KV cache movements are driven by a narrow, router-initiated hint surface (”KvHint”) which enables Dynamo to pass KV cache intent to vLLM/SGLang. This avoids invasive changes to the engine’s native scheduler and cache manager. (SGLang design, vLLM design).

The orchestrator holds what a request-local policy cannot see. Which sessions are still live, which token ranges are a shared prefix versus a unique tail, whether a tool-call gap will last milliseconds or minutes, when a subagent opens and closes. That knowledge belongs in the router layer, where the harness contract and the session lifecycle already live, and it should stay there. The engine’s job is to expose the mechanism and let the router supply the intent.

The proposal is a narrow set of soft hints the router can emit to the engine. Here is a partial taxonomy of potential KvHints and their functions:

KvHint name Functionality Example use
Share Moves a cached prefix from one worker to another via p2p transfer A new engine worker warms its cache from the existing workers’ cache.
Prefetch Fetches KV from a higher tier (e.g. DRAM) to lower tier, ahead of need Reloading a main agent’s context into HBM as a subagent approaches completion.
Demote Pushes KV from a lower tier to a higher tier. Long tool call, or paused trajectory.
Pin Pins a high-value KV cache onto its current tier for a bounded TTL Pinning a system cache shared by all requests.
Retain Biases KV blocks to be more (or less) likely to be retained under cache pressure. Retaining a subagent’s KV blocks during a tool call.

Each hint carries the session ID with it, so the tier receiving the KV knows which session it belongs to and can group, place, and reclaim it as a unit rather than as anonymous blocks. That matters most in the colder tiers, where a storage backend holding host memory or disk otherwise sees only opaque hashes and has no way to tell a live trajectory’s prefix from a finished one. The engine can clip, defer, or ignore any hint, and workloads that send no hints behave exactly as they do today.

Share already has a working implementation layered into HiCache with minimal scheduler hooks, and Prefetch and Demote are designed to build on that machinery once Share lands.

Enabling KV Hint Policy Computations with Session-Prefix Indexer

For Dynamo to compose and execute these KvHint policies on the session level, it must own a continuously updated structure which bidirectionally maps between KV blocks, their session identities, and their in-session positional lineage. For this, we implemented the Session-Prefix Indexer, a radix tree structure for which each branch represents a unique session (or subagent session which branched from another).

The Session-Prefix Indexer is continuously hydrated from both inference engine KV events (Stored / Removed) and the prefix cache hits from Dynamo router’s main indexer (“FlashIndexer”). Consequently, through the FlashIndexer and Session-Prefix Indexer, the Dynamo router has a cumulative view of (1) inference engine workers’ block ownership as well as (2) those same blocks’ session membership. This information-rich view can be consumed by various KV hint policies that require signals on KV blocks’ worker location, active session lifecycles, and/or engine workloads.

With these indexers, we are actively experimenting with KvHint policies with Dynamo vLLM/SGLang on agentic workload datasets.

Figure 4: The Session-Prefix Indexer tracks shared prefixes, session lineage, and KV block locations within the Dynamo router. KV events from inference workers update this view, supporting routing decisions and policies that send optional KV hints to vLLM or SGLang.

Looking forward

Efficient agentic inference depends on managing context across the full session lifecycle. Dynamo’s session ID connects harness behavior to routing, admission control, and cache management, giving the serving stack a shared view of an agent’s evolving working set.

Session-linked traces and replay make policies measurable, while session-aware admission control reduces cache thrashing by applying backpressure at tool boundaries. In the SWE-bench configuration tested here, program-aware scheduling improved throughput by roughly 12–16% over KV-aware routing alone. The experimental shared-pool indexer extends routing visibility to external KV stores.

The Session-Prefix Indexer and proposed KvHint interface build toward proactive cache movement: Dynamo supplies policy intent, and engines retain control over execution. The key takeaway is that cache reuse and session concurrency must be optimized together. Keeping useful context available across turns and avoiding repeated prefill as concurrency grows are central to serving more agent sessions efficiently.

The post Session-Aware Agentic Inference with NVIDIA Dynamo appeared first on PyTorch.

來源:PyTorch:Blog · pytorch.org