Skip to content

Serving-load PR3/3: V4.1 on the shared runner; staircase feeds v41_capacity_measurement; fleet mode spark_v41_serving_load - #12875

Merged
gunbai-bot[bot] merged 7 commits into
mainfrom
session/fierce-carp-324-pr3
Oct 1, 2026
Merged

gunbai-bot[bot] merged 7 commits into
mainfrom
session/fierce-carp-324-pr3

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

PR3 of 3 for the shared serving-load runner (#12838, #12842 merged). It binds DeepSeek V4.1 to the same runner GLM uses. Per-model numbers stay per model; only the engineering is shared.

V4.1's binding: gunbc.spark.v41_serving_load

  • Subject: Group A; head srv6 at 127.0.0.1:<spark_pair_serve_port> (30000); served alias v41_group_a_served_alias (deepseek-v4.1-flash); container v41_group_a_container_name (gunbc-v41-tp4).
    • Runtime is the production image (v41_production_produced_image). The standing is believed, not established: this probe does not read the head's image back, and the would_settle obligation names that read.
    • V4.1 has its own admission and qualification frontier rows.
  • Steps: one vllm bench serve per v41_staircase_concurrencies level (1, 2, 4, 6, 8, 12, 16), with num_prompts = 2c and max_concurrency = c.
  • Protocol (operator ruling 2026-10-01):
    • v41_serving_load_dataset is 256 in / 128 out / prefix 0. It is a copy of GLM's shape, declared as V4.1's own row: it does not import GLM's.
    • This protocol does not exercise the 262k-context TTFT gate, so that gate's reading is a stated frontier (v41_ttft_gate_frontier, printed in every receipt), not a pass.
  • Host pressure (operator ruling 2026-10-01): every rank, the worst rank wins. The runner gained a per-step host-reading bracket, generic over the model's reading type. V4.1's reader runs on all four Group A ranks (srv5–8), before each bench and after each after-scrape, and reads:
    • host PSI memory;
    • the diskstats row of the device holding the Engram row stores, found per host by findmnt on that directory rather than named by guess;
    • the rank container's memory.stat.
  • Projection into v41_capacity_measurement: each step becomes a V41StaircaseStepReading. For each stop rule (refault growth, PSI full stall, NVMe read await), the step takes the rank that rule judges worst. An unmeasurable interval ranks worst, so it surfaces as a missing reading. Preemptions come from the head's /metrics.
    • A rank unread on either side leaves the step unprojected, with the rank and cause named. v41_run_staircase then reports a missing reading at that concurrency instead of judging three ranks as four.
    • The staircase runs under the signed v41_capacity_policy. The receipt carries every step's observation, each projection, and sla_sustainable.

Stated bet. The cgroup path assumes dockerd's default systemd cgroup driver: system.slice/docker-<id>.scope, per extdeps.docker.cgroup. The driver is not modeled. A cgroupfs daemon makes the read refuse with the path it tried; it is never misread.

Runner change. serving_load_records<H> returns per-step records (request, observation, bodies, bench, host_before and host_after). serving_load_run is now that function plus a receipt. GLM passes no_host_reading, and its receipt is unchanged: the byte-equality witness was re-run.

Fleet-converge mode spark_v41_serving_load. A mode is a row, not a job:

  • one arm in every FleetConvergeWorkflowMode match: wire, modes list, live-deploy, key, mutation domain, job id;
  • a step and its if;
  • a ci_spec target and invoke;
  • a sentence in the mode description.

The mutation domain is Group A's arm, so it serializes with the launch. The target must be the head (srv6); any other host refuses. No job is added. .github/workflows/fleet-converge.yml is generated from these rows; if the generated lane reports drift, I'll regenerate it here.

Evidence (local gunbc run). These modules are outside every required lane, so CI does not execute them.

Run Result
5 V4.1 witnesses: each_stop_rule_reads_its_own_worst_rank, an_unread_rank_unprojects_the_step_and_names_it, every_staircase_step_addresses_the_v41_subject, the_protocol_row_is_glms_shape_declared_as_v41s_own, findmnt_source_names_the_diskstats_device, plus GLM's byte-equality and subject-admission witnesses rc=0
Mutation: worst-rank comparison inverted rc=1, worst_rank

gunbc compile --entry fleet_converge_workflow.dag typechecks the workflow, ci_spec and the new modules with no diagnostic in them. Its 3 blocking errors are emission-time refusals in extdeps.gunbc (WitnessBin.Run's shell output channels), a module this PR does not touch.

Frontiers (named, not passed)

  • Wet run: waits until V4.1 serves on main's path (live bring-up is on deploy/ds41-dev-launch, unmerged). No wet run was performed.
  • TTFT gate: not exercised by the GLM-shaped protocol until a V4.1-specific long-context protocol is ruled.
  • Capacity report: v41_capacity_report still needs the startup receipts and probe cells. This PR feeds only the staircase (the sla_sustainable axis). v41_wet_runner_frontier is partially discharged: the staircase readings now have a producer, while the startup, residency and probe readings do not.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 5 commits October 1, 2026 04:05
…bracket (PR3 groundwork)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ds v41_capacity_measurement

gunbc.spark.v41_serving_load: V4.1's subject (Group A, head srv6 :30000, deepseek-v4.1-flash,
gunbc-v41-tp4), one bench step per v41_staircase_concurrencies level, and a host-pressure reading
on every Group A rank around each step (runner's new per-step bracket). Each step's records are
projected into V41StaircaseStepReading taking the WORST rank per stop rule, and run through
v41_run_staircase under the signed policy. Operator ruling 2026-10-01: pressure on all four ranks,
worst rank; protocol is a copy of GLM's 256/128/0 shape declared as V4.1's own row; the 262k TTFT
gate is not exercised by it and is a stated frontier.

Runner: per-step records + a model-owned host-reading bracket; GLM passes NoHostReading and its
receipt stays byte-equal (witness re-run).

Fleet-converge mode spark_v41_serving_load: a row in every FleetConvergeWorkflowMode match, a step,
and a ci_spec target. No new job.

Executed locally: 5 V4.1 witnesses + 2 GLM runner witnesses rc=0; worst-rank comparison
inverted rc=1 worst_rank.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…24-pr3

# Conflicts:
#	dag/gunbc/fleet/fleet_converge_workflow.dag
… generated_artifact_gate)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oncurrencyUnbound) instead of being judged at level 0

Addresses review 73600. Executed locally: 8 witnesses rc=0; restoring the fabricated-0 arm reds
a_request_without_a_concurrency_refuses_instead_of_judging_level_zero (rc=1).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Review 73600 (fabricated default concurrency): fixed in d57fae0. The level is now read only by v41_project_for_request. A request with no max_concurrency becomes V41StepConcurrencyUnbound { cause }: it contributes no step reading, so v41_run_staircase reports the level as missing, and the receipt prints the cause. New witness a_request_without_a_concurrency_refuses_instead_of_judging_level_zero checks both arms: absent refuses, and a real step projects at its own level. Local runs: all 8 witnesses pass (rc=0); putting the old Absent => 0 arm back reds the new witness (rc=1). — sent from fierce-carp-324

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 1, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Oct 1, 2026
gunbc-ci-auto-heal and others added 2 commits October 1, 2026 10:23
…24-pr3

# Conflicts:
#	.github/workflows/fleet-converge.yml
#	dag/gunbc/fleet/fleet_converge_workflow.dag
…ed_artifact_gate)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit a280b65 Oct 1, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/fierce-carp-324-pr3 branch October 1, 2026 11:57
@briansrls
briansrls restored the session/fierce-carp-324-pr3 branch October 1, 2026 12:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants