feat(recipe): add ThunderAgent to the Dynamo rollout backend - #126
Merged
Merged
Conversation
Integrate ThunderAgent program-aware routing with the PR verl-project#110 Dynamo backend, wire agent lifecycle and finalization, pin the tested verl revision, and add documentation and targeted tests.
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
Contributor
wuxibin89
approved these changes
Jul 28, 2026
sophiayyya
added a commit
to sophiayyya/verl-recipe
that referenced
this pull request
Sep 9, 2026
…oject#126) ## Summary Add recipe-side [ThunderAgent](https://arxiv.org/abs/2602.13692) integration on top of the Dynamo rollout backend introduced in verl-project#110(verl-project#110). This change adds program-aware routing for multi-turn agent trajectories while preserving the existing verl agent-loop and colocated training architecture. All integration remains inside the recipe; it does not modify core verl or Dynamo source code. ## Changes - Enable ThunderAgent through dynamo_trainer.yaml while retaining: - thunderagent.enabled=true for the Dynamo with ThunderAgent. - thunderagent.enabled=false for the native Dynamo KV-router baseline. - VARIANT=global for the native verl vLLM baseline. - Add concise launch examples: - dynamo/run_uniagent_variant.sh - dynamo/smoke_dynamo_v1.sh - dynamo/train_30b_rl_dynamo_kv_metrics.sh - Add focused unit tests for program lifecycle, routing headers, finalization, process ordering, configuration, registration, and verl compatibility. - Document KV-aware routing and end-to-end ThunderAgent benchmark results, including the corresponding plots. ## Required versions - Dynamo: 59d614641837e593f0567b79d75394aae5f864e0 - Includes the ThunderAgent lifecycle support merged through ai-dynamo/dynamo#11185 (ai-dynamo/dynamo#11185). ## Benchmark evidence For the matched Uni-Agent × verl synchronous GRPO experiment: - Concurrency 384: 1.94× rollout throughput and 1.39× observed full-step throughput. - Concurrency 512 <img width="416" height="224" alt="thunder agent rollout" src="https://github.com/user-attachments/assets/21f3abbd-b868-42d2-8065-6fe055fffbfb" /> <img width="416" height="226" alt="thunder_agent_globalstep" src="https://github.com/user-attachments/assets/2b95f045-e95a-4c7c-a435-7b3108f5e3b6" /> These benchmark results come from the matched runs documented in dynamo/README.md; no additional GPU training job was launched while preparing this PR. --------- Co-authored-by: Jinyan Chen <jinyanc@nvidia.com> Co-authored-by: Sophia Yang <sopyang@nvidia.com> Co-authored-by: OpenAI Codex <codex@openai.com>
sophiayyya
pushed a commit
to sophiayyya/verl-recipe
that referenced
this pull request
Sep 9, 2026
With ThunderAgent enabled the ONLY handler advertising the public model name must be the TA router; the vLLM path has renamed its workers to <model>--verl-thunderagent-backend since PR verl-project#126, but the sglang path launched workers under the public name — the frontend could then route generation (though not session-finals, which the validated run proved were all router-terminated) directly to a worker, silently bypassing program affinity. Mirror the rename in a _build_sglang_cmd override, and fail fast when an explicit sglang.page_size contradicts thunderagent.router_block_size (the base builder already aligns page size to the router block size when unset). The earlier TA x sglang smoke is hereby downgraded to partial evidence pending a rerun with positive routing assertions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Summary
Add recipe-side ThunderAgent integration on top of the Dynamo rollout backend introduced in #110(#110).
This change adds program-aware routing for multi-turn agent trajectories while preserving the existing verl agent-loop and colocated training architecture. All integration remains inside the recipe; it does not modify core verl or Dynamo source code.
Changes
Enable ThunderAgent through dynamo_trainer.yaml while retaining:
Add concise launch examples:
Add focused unit tests for program lifecycle, routing headers, finalization, process ordering, configuration, registration, and verl compatibility.
Document KV-aware routing and end-to-end ThunderAgent benchmark results, including the corresponding plots.
Required versions
(feat(thunderagent): include HiCache retention capacity ai-dynamo/dynamo#11185).
Benchmark evidence
For the matched Uni-Agent × verl synchronous GRPO experiment: