Repository navigation
fix(session): disable training replay outputs during evaluation - #3323
Conversation
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
| try: | ||
| params = CreateSessionRequest.model_validate_json(await request.body() or b"{}") | ||
| except ValidationError as exc: | ||
| return JSONResponse(status_code=400, content={"error": str(exc)}) |
|
@claude review always |
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
|
@claude review always |
10518c5 to
e8e1b63
Compare
| @@ -50,7 +50,7 @@ async def generate(input: GenerateFnInput) -> GenerateFnOutput: | |||
| ) | |||
| use_v2 = getattr(input.args, "use_session_server", None) == "v2" | |||
| collect_spec_metrics = use_v2 and input.args.sglang_speculative_algorithm is not None | |||
| tracer = await OpenAIEndpointTracer.create(input.args) | |||
| tracer = await OpenAIEndpointTracer.create(input.args, evaluation=input.evaluation) | |||
There was a problem hiding this comment.
instead of passing evaluation flag, how about explicitly passing "is_replay"? In case, we need to turn it off for non-evaluation requests too. What do you think?
There was a problem hiding this comment.
I also thought of this... My idea is the evaluation semantic is easy to be understood by user, and also we might skip more TITO components for evaluation to ensure perf. So I think keep this param name is ok.
Agentic evaluation inherited training replay settings and R3-only request constraints because session creation did not record its purpose. Propagate evaluation through the tracer and v1/v2 session state, then force replay outputs off after model request resolution. Preserve sampling defaults, TITO behavior, and the training path. Add HTTP coverage for creation, continuation, retries, concurrent sessions, and sample collection without replay data, plus user documentation.
Validate the creation body with CreateSessionRequest at the HTTP route and pass only the evaluation boolean into SessionCore. Keep parse_chat_request specific to chat completions instead of sharing its input contract with session creation. Preserve empty-body defaults, strict booleans, error response shape and allocation ordering. Validation errors now use schema diagnostics. Verification: 481 focused CPU tests passed on an unchanged rerun; the first run had an intermittent upstream empty-404 failure in the parity test. All pre-commit hooks passed. No GPU backend code changed.
e8e1b63 to
eeddbe1
Compare
Session creation previously ignored misspelled or unsupported fields, which could silently select the training policy. Use StrictBaseModel for CreateSessionRequest and cover unknown keys at the HTTP boundary before session allocation.
There was a problem hiding this comment.
The CreateSessionRequest gap I flagged earlier (plain BaseModel silently dropping unknown/misspelled keys like evalution) is now fixed: it subclasses StrictBaseModel and POST /sessions validates the body via model_validate_json, returning HTTP 400 on ValidationError before the session is allocated.
Extended reasoning...
Diff vs base commit 418be99 is the PR's three commits (6b6378c, eeddbe1, 6bad30d), touching session creation/evaluation-flag plumbing across core.py, sessions.py, linear_trajectory.py, request_args.py, session_state.py, types.py, and openai_endpoint_utils.py. Verified the previously-flagged StrictBaseModel/extra-forbid gap in types.py is now closed and the /sessions route rejects invalid bodies with 400 before allocation, matching the PR description. No new issues found in the evaluation-flag threading; a human reviewer already approved.
Summary
Keep evaluation sessions free of training replay data and reject unknown creation fields.
Symptom & Reproduction
{"evalution": true}was accepted and silently created a training session.{"evalution": true}or an unrelated field toPOST /sessions; each request must return HTTP 400 before allocating a session.Root Cause
OpenAIEndpointTracer.createand registries dropped evaluation purpose.resolve_request_args_by_configapplied training replay policy to every session.CreateSessionRequestinherited permissiveBaseModel.evaluationtofalse.Fix
Validate
POST /sessionswithCreateSessionRequest, now based onStrictBaseModel, and pass only the resolvedevaluation: boolintoSessionCore.create_session. Unknown, misspelled, or invalid creation fields return HTTP 400 before allocation; empty bodies still default to training. Chat completion parsing remains separate.Propagate
GenerateFnInput.evaluationthrough the tracer into v1/v2 session state so the purpose survives rollback and branching. After model rules resolve the request, evaluation forcesreturn_sampling_mask,return_routed_experts, andreturn_indexer_topktofalseand removesrouted_experts_start_len.Training behavior, sampling default/override precedence, TITO rendering, and compatibility checks stay unchanged. Evaluation may vary temperature between turns, and sample collection remains supported without replay data.
Verification
/opt/sglang/bin/python -m pytest tests/fast/router/test_session_evaluation.py -q: 37 passed at6bad30d5327c92533eda7457bb61c2f98c875546, covering valid creation, typo/unknown-field rejection before allocation, continuation/retry, concurrent train/eval sessions, sampling resolution, and collection without replay outputs.ruff,autoflake,isort, andblackchecks on the changed files passed;git diff --checkpassed.Review Focus
CreateSessionRequest/SessionCore.create_session: reject unknown creation fields at the HTTP boundary while passing only the resolved boolean into core state.SessionStateV2/LinearTrajectory: keep session purpose across continuation, retries, and concurrent training/evaluation.prepare_chat_request: enforce replay suppression without changing the training path; disabling response fields does not guarantee zero internal engine capture.