Repository navigation
fix(grpc): forward scheduler load info for DP-aware load balancing #1115
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We鈥檒l occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -28,6 +28,7 @@ | |
| HealthCheckOutput, | ||
| TokenizedEmbeddingReqInput, | ||
| TokenizedGenerateReqInput, | ||
| WatchLoadUpdateReq, | ||
| ) | ||
| from sglang.srt.observability.req_time_stats import ( | ||
| APIServerReqTimeStats, | ||
|
|
@@ -679,6 +680,15 @@ async def cleanup(request_id): | |
|
|
||
| cleanup_tasks.append(asyncio.create_task(cleanup(rid))) | ||
|
|
||
| # Forward load info to DataParallelController for token-aware balancing. | ||
| # Mirrors TokenizerManager._handle_batch_output logic: when dp_size > 1, | ||
| # each scheduler piggybacks its load (num_reqs, num_tokens) on batch output. | ||
| # Without this, DPBudget stays at zero and total_tokens/total_requests | ||
| # policies degenerate to always picking rank 0. | ||
| if self.server_args.dp_size > 1 and batch_out.load is not None: | ||
| load_update = WatchLoadUpdateReq(loads=[batch_out.load]) | ||
| self.send_to_scheduler.send_pyobj(load_update) | ||
|
Comment on lines
+688
to
+690
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Make DP load forwarding best-effort to avoid dropping batch outputs. At Line 690, an unguarded send can raise and abort Suggested fix- if self.server_args.dp_size > 1 and batch_out.load is not None:
- load_update = WatchLoadUpdateReq(loads=[batch_out.load])
- self.send_to_scheduler.send_pyobj(load_update)
+ if self.server_args.dp_size > 1 and batch_out.load is not None:
+ try:
+ await self._send_to_scheduler(WatchLoadUpdateReq(loads=[batch_out.load]))
+ except Exception as e:
+ logger.warning(f"Failed to forward DP load update: {e}")馃 Prompt for AI Agents |
||
|
|
||
| # Execute all queue.put() operations in parallel | ||
| if put_tasks: | ||
| await asyncio.gather(*put_tasks, return_exceptions=True) | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
To maintain consistency with the rest of the
GrpcRequestManagerclass (see lines 368, 427, and 471), please use the_send_to_schedulerhelper method instead of callingself.send_to_scheduler.send_pyobjdirectly. This ensures that the send operation is covered by the helper's error logging. Since_handle_batch_outputis anasyncfunction, the call should beawaited to follow the established pattern in this class.References