Repository navigation
Serving-load PR3/3: V4.1 on the shared runner; staircase feeds v41_capacity_measurement; fleet mode spark_v41_serving_load - #12875
Merged
Conversation
…bracket (PR3 groundwork) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ds v41_capacity_measurement gunbc.spark.v41_serving_load: V4.1's subject (Group A, head srv6 :30000, deepseek-v4.1-flash, gunbc-v41-tp4), one bench step per v41_staircase_concurrencies level, and a host-pressure reading on every Group A rank around each step (runner's new per-step bracket). Each step's records are projected into V41StaircaseStepReading taking the WORST rank per stop rule, and run through v41_run_staircase under the signed policy. Operator ruling 2026-10-01: pressure on all four ranks, worst rank; protocol is a copy of GLM's 256/128/0 shape declared as V4.1's own row; the 262k TTFT gate is not exercised by it and is a stated frontier. Runner: per-step records + a model-owned host-reading bracket; GLM passes NoHostReading and its receipt stays byte-equal (witness re-run). Fleet-converge mode spark_v41_serving_load: a row in every FleetConvergeWorkflowMode match, a step, and a ci_spec target. No new job. Executed locally: 5 V4.1 witnesses + 2 GLM runner witnesses rc=0; worst-rank comparison inverted rc=1 worst_rank. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…24-pr3 # Conflicts: # dag/gunbc/fleet/fleet_converge_workflow.dag
… generated_artifact_gate) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oncurrencyUnbound) instead of being judged at level 0 Addresses review 73600. Executed locally: 8 witnesses rc=0; restoring the fabricated-0 arm reds a_request_without_a_concurrency_refuses_instead_of_judging_level_zero (rc=1). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Author
|
Review 73600 (fabricated default concurrency): fixed in d57fae0. The level is now read only by |
github-merge-queue
Bot
removed this pull request from the merge queue due to a conflict with the base branch
Oct 1, 2026
…24-pr3 # Conflicts: # .github/workflows/fleet-converge.yml # dag/gunbc/fleet/fleet_converge_workflow.dag
…ed_artifact_gate) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR3 of 3 for the shared serving-load runner (#12838, #12842 merged). It binds DeepSeek V4.1 to the same runner GLM uses. Per-model numbers stay per model; only the engineering is shared.
V4.1's binding:
gunbc.spark.v41_serving_load127.0.0.1:<spark_pair_serve_port>(30000); served aliasv41_group_a_served_alias(deepseek-v4.1-flash); containerv41_group_a_container_name(gunbc-v41-tp4).v41_production_produced_image). The standing is believed, not established: this probe does not read the head's image back, and thewould_settleobligation names that read.vllm bench serveperv41_staircase_concurrencieslevel (1, 2, 4, 6, 8, 12, 16), withnum_prompts = 2candmax_concurrency = c.v41_serving_load_datasetis 256 in / 128 out / prefix 0. It is a copy of GLM's shape, declared as V4.1's own row: it does not import GLM's.v41_ttft_gate_frontier, printed in every receipt), not a pass.findmnton that directory rather than named by guess;memory.stat.v41_capacity_measurement: each step becomes aV41StaircaseStepReading. For each stop rule (refault growth, PSI full stall, NVMe read await), the step takes the rank that rule judges worst. An unmeasurable interval ranks worst, so it surfaces as a missing reading. Preemptions come from the head's/metrics.v41_run_staircasethen reports a missing reading at that concurrency instead of judging three ranks as four.v41_capacity_policy. The receipt carries every step's observation, each projection, andsla_sustainable.Stated bet. The cgroup path assumes dockerd's default systemd cgroup driver:
system.slice/docker-<id>.scope, perextdeps.docker.cgroup. The driver is not modeled. A cgroupfs daemon makes the read refuse with the path it tried; it is never misread.Runner change.
serving_load_records<H>returns per-step records (request, observation, bodies, bench,host_beforeandhost_after).serving_load_runis now that function plus a receipt. GLM passesno_host_reading, and its receipt is unchanged: the byte-equality witness was re-run.Fleet-converge mode
spark_v41_serving_load. A mode is a row, not a job:FleetConvergeWorkflowModematch: wire, modes list, live-deploy, key, mutation domain, job id;if;ci_spectarget and invoke;The mutation domain is Group A's arm, so it serializes with the launch. The target must be the head (srv6); any other host refuses. No job is added.
.github/workflows/fleet-converge.ymlis generated from these rows; if the generated lane reports drift, I'll regenerate it here.Evidence (local
gunbc run). These modules are outside every required lane, so CI does not execute them.each_stop_rule_reads_its_own_worst_rank,an_unread_rank_unprojects_the_step_and_names_it,every_staircase_step_addresses_the_v41_subject,the_protocol_row_is_glms_shape_declared_as_v41s_own,findmnt_source_names_the_diskstats_device, plus GLM's byte-equality and subject-admission witnessesworst_rankgunbc compile --entry fleet_converge_workflow.dagtypechecks the workflow,ci_specand the new modules with no diagnostic in them. Its 3 blocking errors are emission-time refusals inextdeps.gunbc(WitnessBin.Run's shell output channels), a module this PR does not touch.Frontiers (named, not passed)
deploy/ds41-dev-launch, unmerged). No wet run was performed.v41_capacity_reportstill needs the startup receipts and probe cells. This PR feeds only the staircase (thesla_sustainableaxis).v41_wet_runner_frontieris partially discharged: the staircase readings now have a producer, while the startup, residency and probe readings do not.🤖 Generated with Claude Code