Skip to content

Promote shape-aware r14 EXL3 prefill - #16

Merged
malaiwah merged 2 commits into
mainfrom
codex/r14-capacity-shape-default
Aug 1, 2026
Merged

malaiwah merged 2 commits into
mainfrom
codex/r14-capacity-shape-default

Conversation

@malaiwah

@malaiwah malaiwah commented Aug 1, 2026

Copy link
Copy Markdown
Owner

What changed

  • Compose vLLM #210's bounded/sliced EXL3 prefill arena before vLLM #219's
    shape-aware mixed-K executor in the fail-closed field-review manifest.
  • Keep the cooperative one-grid/block-8 route for decode-sized rows and use
    serial homogeneous K3/K4 block-64 plans for large prefills.
  • Select a 1,024-row reusable arena for the 3.25-bpw profile and document its
    measured memory/performance trade.
  • Update the cache ABI namespace, patch checksums, profile tests, README,
    changelog, self-service guidance, and full AIBeast qualification record.

Why

GG v20-r14's one-grid mixed-K route regressed 3K/32K/128K prefill by about 23%
versus r13. The executor fix recovers 28.4-29.3% over pristine r14 while
retaining r14's hardening and decode route. The 1,024-row #210 arena then
returns about 665 MiB/GPU. A 1,536-row arm returned only about 332 MiB/GPU but
paid the same roughly 10-11% PP cost, so 1,024 is the useful safety/performance
point for the full 524,288-token profile.

Live qualification

Source-exact turnkey image on AIBeast: 4x RTX PRO 6000 Blackwell 96 GB at
280 W, driver 595.71.05, CUDA 13.2; willfalco 3.25-bpw snapshot
61d2b6b757f6a4ac7098a78d861f2033497532dc; TP4/DCP4; 2,048 GPU blocks;
125 GiB DRAM plus bounded 512 GiB NVMe LMCache.

  • cold unique-prefix PP: 2,121 / 1,932 / 1,837 tok/s at 3K / 32K / 128K
  • aggregate MTP5 TG: 93 / 126 / 162 / 240 tok/s at C1 / C2 / C4 / C8
  • MTP5 MAL: 3.31 / 3.17 / 3.11 / 3.70
  • actual 509,022-token prompt: 3/3 needles, no degeneration
  • 524,288 active GPU-KV tokens; no request failure, preemption, exception, or
    CUDA OOM
  • both GLM-5.2 and local-primary aliases passed

The turnkey profile itself retains MTP3 because the matched field-review
workload found better acceptance and tail latency at depth 3; MTP5 was used in
the AIBeast control so only the executor/capacity stack changed.

Upstream lineage

Validation

  • exact repository CI suite from AGENTS.md
  • all three configuration-smoke profiles
  • field-review manifest/application tests
  • EXL3 mixed-K and parity-ABI patch tests
  • structured-output patch and field-review log-audit tests
  • JSON and whitespace checks
  • source-exact GPU image boot, cold performance matrix, alias checks, LMCache
    activity, and near-maximum retrieval gate

@malaiwah
malaiwah merged commit 6ab4b09 into main Aug 1, 2026
2 checks passed
@malaiwah
malaiwah deleted the codex/r14-capacity-shape-default branch August 1, 2026 00:51
@malaiwah

malaiwah commented Aug 1, 2026

Copy link
Copy Markdown
Owner Author

Post-merge production handoff on AIBeast is green:

  • glm52-turnkey-r14-cap1024-shape-prod is healthy on port 8000; port 18000
    is closed and no qualification container owns GPU memory.
  • Warm compile-cache reuse took 0.60 s for the backbone; graph capture took
    8 s and 0.20 GiB/GPU. The runtime reported the qualified 1,024-row target
    and draft plans and exactly 524,288 GPU-KV tokens.
  • The built-in authenticated verifier passed short prompts, strict JSON with
    thinking, and 3/3 needles at 32,853 tokens. A separate local-primary
    verifier passed arithmetic, factual, instruction, and structured-output
    gates.
  • Under immediate agent traffic, generation windows reached 120-125 tok/s at
    C1 with MTP5 MAL 4.27-4.45; external prefix-cache hit rate was 18.6%.
  • Full boot and soak logs contain no ERROR, traceback, CUDA OOM, request
    failure, or preemption.
  • Checkpoint remains read-only; LMCache remains 125 GiB DRAM + bounded 512 GiB
    local NVMe.

The host's Podman release cannot add a restart policy to an existing container
in place, so I deliberately did not replace the healthy service solely for
that non-runtime setting. PID 1 still supervises/restarts the vLLM process;
after a host reboot the standard launch command is required.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant