Skip to content

compass(kv): refuse a parallel width written as integer text - #340

Merged
jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-258-text-widths
Sep 23, 2026
Merged

jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-258-text-widths

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Closes #258

What changed

whole_number in atom/compass/kv/handoff.py no longer reads text. At 196fd3711 it ran int() on any str first, so tp_size="8" and " 8 " went out in the blob as 8. The special case is deleted: int(value) is now compared with value itself, so "8" is refused because 8 != "8". It raises the same ValueError with the same message as every other width that is not a whole number. int and a float with no fractional part are still taken, and bool is still refused.

No refusal text changed. The message template is untouched; only which inputs reach it changed.

Measured before changing it (node 18, xiaobizh_n18_cpu, tree 196fd3711)

probe result
whole_number("tp_size", "8") returns 8
whole_number("tp_size", " 8 ") returns 8
Config(model=..., tensor_parallel_size="8") TypeError: '<=' not supported between instances of 'int' and 'str'
ParallelConfig(data_parallel_rank="2") TypeError: '<' not supported between instances of 'str' and 'int' (with ATOM_DP_RANK unset)

A grep of atom/, scripts/ and tests/ for a width passed as a string literal or str(...) finds only the three tests in tests/compass/test_kv_remote_prefill.py that this PR changes. The Compass spec schema refuses tensor_parallel_size as a key (DEPLOYMENT_OWNED), so no spec document carries one. -tp and --data-parallel-rank are type=int in atom/model_engine/arg_utils.py, and ATOM_DP_RANK is parsed with int() in atom/utils/envs.py.

At the head, the same probe refuses "8", " 8 ", "3" and "9007199254740993" with tp_size is '8', which is not a whole number; .... It still takes 8 and 8.0 as 8.

Tests (tests/compass/test_kv_remote_prefill.py)

  • test_a_width_that_is_not_a_whole_number_is_refused_by_name gains [tp_size-integer-text] ("8"), [dp_rank-integer-text] ("3") and [tp_size-padded-text] (" 8 "). Its docstring's last sentence now says text is refused whatever it spells.
  • test_the_ranks_the_router_reads_are_numbers drives the connector with 8.0 / 3.0 instead of "8" / "3". The sentence saying a width filled from an environment variable or JSON file arrives as a string is gone; the docstring now says a whole float is converted and text is refused.
  • test_a_width_in_text_keeps_its_exact_value is removed. It asserted that "9007199254740993" went out exactly, and text no longer goes out at all.
  • test_a_whole_width_is_taken_whatever_it_is_spelled_as drops its text and padded cases, and is renamed test_a_whole_width_is_taken_as_an_int_or_a_whole_float to match the two spellings it still takes.

Named result (node 18, xiaobizh_n18_cpu)

tests/compass/test_kv_remote_prefill.py, whole file, with atom.__file__ checked to be under each staged root:

tree result
head 8bab7230b 44 passed
head with 196fd3711's handoff.py restored (112 lines, byte-identical to the tip; commit fd31f767d via git commit-tree) 3 failed, 41 passed

The three that fail, each with Failed: DID NOT RAISE <class 'ValueError'>:

  • test_a_width_that_is_not_a_whole_number_is_refused_by_name[tp_size-integer-text]
  • test_a_width_that_is_not_a_whole_number_is_refused_by_name[dp_rank-integer-text]
  • test_a_width_that_is_not_a_whole_number_is_refused_by_name[tp_size-padded-text]

Null controls, which pass on both trees: test_a_whole_width_is_taken_as_an_int_or_a_whole_float[int] and [float], test_the_ranks_the_router_reads_are_numbers, and the 11 refusal cases that were already there.

Gate 1: scripts/compass/gate_cpu.sh, the tree's own copy, as a delta

Staged with git archive plus both stamps, piped into xiaobizh_n18_cpu under /tmp/i258gates/, never the shared mount. PYTHONPATH was set to each root, and atom.__file__ was asserted to be under it and printed before each run. Each gate printed gpu: not required (.compass-changed stamp).

run commit stamp result GATE_CPU_RC
control 1 196fd3711 5244 passed, 155 skipped, 3 xfailed 0
branch 8bab7230b 5244 passed, 155 skipped, 3 xfailed 0
control 2 (--junitxml) 196fd3711 5244 passed, 155 skipped, 3 xfailed 0

Node-id delta, branch vs control 2 (junit, 5402 cases each side). No shared node id changed outcome. The only differences are in tests/compass/test_kv_remote_prefill.py:

  • only at control: test_a_width_in_text_keeps_its_exact_value, test_a_whole_width_is_taken_whatever_it_is_spelled_as[int], [text], [float], [padded], all passed;
  • only at branch: test_a_whole_width_is_taken_as_an_int_or_a_whole_float[int], [float], and test_a_width_that_is_not_a_whole_number_is_refused_by_name[tp_size-integer-text], [dp_rank-integer-text], [tp_size-padded-text], all passed.

The timing classes in tests/entrypoints/test_stream_marker_properties.py and tests/test_gc_utils.py had the same outcome in the two runs with a per-test record (control 2 and branch). Control 1 has no per-test record; its totals are the same.

git merge-tree --write-tree 196fd3711 8bab7230b gives 10c58af3df157f99cc9c85bb25cb48a424b27e96, which is 8bab7230b^{tree}.

ruff check and black --check on both changed files: clean at the tip and at the head.

Lines

added removed
production (atom/compass/kv/handoff.py) 10 11
tests (tests/compass/test_kv_remote_prefill.py) 18 38

Estimate was small (10–25). Of the production lines, 3 code lines are removed and 2 added; the rest of the production diff is the docstring. Net: 28 added, 49 removed.

Left alone

#258 records two things as out of scope, and this PR does not touch either: numpy.bool_(True) accepted as 1, and whole_number checking integrality and not range.

🤖 Generated with Claude Code

whole_number read text through int() when int() took it as an integer
literal, so a config carrying tp_size="8" or dp_rank=" 3 " went out in
the blob as 8 and 3. Nothing in ATOM delivers a width as text: the -tp
and --data-parallel-rank flags are type=int, ATOM_DP_RANK is parsed with
int(), and Config raises TypeError on a string width. Reading text was
a conversion with no launch behind it.

The special case is gone. int(value) is compared with value itself, so
"8" is refused because 8 != "8", by the same ValueError and the same
message as any other width that is not a whole number. int and a float
with no fractional part are still taken; bool is still refused.

Tests:
- the refusal test gains "8", "3" and " 8 " by name;
- test_the_ranks_the_router_reads_are_numbers drives the connector with
  whole floats instead of text, and its docstring no longer says widths
  arrive as strings from env vars or JSON;
- test_a_width_in_text_keeps_its_exact_value is removed: it asserted
  that "9007199254740993" went out exactly, and text no longer goes out;
- test_a_whole_width_is_taken_whatever_it_is_spelled_as drops its two
  text cases and is renamed for the two spellings it still takes.

Closes #258

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
counts as an `int` and which would go out as a width of 1 or 0; text;
and anything else `int` does not take exactly. No value that is accepted
changes on the way to the `int` returned, and that check is what refuses
integer text: `int("8")` is 8, which is not equal to "8".

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking. ponytail shrink:

L88/L95/L97: shrink: "text is refused" is said three times in this docstring. Keep L88's "Text is refused, "8" included." and the L97-98 sentence on why the equality check refuses integer text, and drop "text;" from the Refused list. -1 line.

Rule: AI_DEV_RULES gate 4 says the reviewer "runs the ponytail-review skill over the diff to catch over-engineering; its findings are posted like any others." This one is optional.

)
def test_a_whole_width_is_taken_whatever_it_is_spelled_as(geometry, value):
"""The refusal above is not of integer text or of a float with no fraction."""
@pytest.mark.parametrize("value", [8, 8.0], ids=["int", "float"])

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking. ponytail delete:

L313: delete: the [float] case repeats test_the_ranks_the_router_reads_are_numbers (L262-273), which already drives whole floats (8.0 / 3.0) through the same connector and asserts an int goes out. Drop 8.0, and the parametrize with it. -1 line.

I measured it on node 18 (xiaobizh_n18_cpu) with the mutation return whole -> return value at L111 (111 lines kept). The result was 2 failed, 42 passed, and both tests failed the same way (assert (<class 'float'> is int) / assert (False)):

  • test_the_ranks_the_router_reads_are_numbers
  • test_a_whole_width_is_taken_as_an_int_or_a_whole_float[float]

So neither case catches anything the other misses on this defect. The ranks test is the one to keep, because it uses two different values (8 and 3) and so would also catch tp/dp being swapped. The duplication was already there at 196fd3711, as [text] beside the text-driven ranks test. This PR carried it over rather than adding it.

Principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more."

integer text: `int("8")` is 8, which is not equal to "8".
"""
try:
if isinstance(value, bool):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking. The numpy.bool_ gap needs an issue before #258 closes.

This isinstance(value, bool) refuses a Python bool. It does not refuse numpy.bool_. On node 18 at this head, whole_number("tp_size", numpy.bool_(True)) returns 1 (int), and it did the same at 196fd3711.

#258 records this as "not in scope", and the PR's "Left alone" section repeats that. That is fine for this PR. The trouble is where the note lives: it is only in #258's body, and #258 closes when this PR lands, so the one record of the gap ends up in a closed issue.

AI_DEV_RULES: "A finding not fixed in the PR that found it gets an issue: PR bodies are squashed away on landing."

Please file it as its own issue, with #258's other out-of-scope note (integrality is checked but range is not: 0 and -1 are accepted) if you want them together. Do this before or at landing. It does not hold this PR.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review cycle 1: PR #340 (issue #258), head 8bab7230b203289aea216e2fdb62507f7d491e05

Written by a reviewer agent (Claude). I read the eight design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md, both at tip abb70ef48, before reviewing.

Verdict: APPROVE at 8bab7230b203289aea216e2fdb62507f7d491e05. No finding blocks. There are three non-blocking inline comments: two ponytail items and a request to file the numpy.bool_ gap as its own issue before #258 closes.


1. Named result, reproduced on node 18

Setup:

  • Container: xiaobizh_n18_cpu.
  • Staging: each tree was a git archive piped into /tmp/pr340r1/<tree> with docker exec -i … tar -x, with md5 matched on both ends. Nothing was written into the shared mount.
  • Python path: PYTHONPATH was set to each root, and atom.__file__ was printed and asserted to be under that root before every run.
  • Command: pytest tests/compass/test_kv_remote_prefill.py.
tree handoff.py lines result
head 8bab7230b 111 44 passed, rc 0
head with 196fd3711's handoff.py put back (git show, nothing else changed) 112 3 failed, 41 passed, rc 1

The 3 failures are the ones the PR names, each with Failed: DID NOT RAISE <class 'ValueError'>:

  • tests/compass/test_kv_remote_prefill.py::test_a_width_that_is_not_a_whole_number_is_refused_by_name[tp_size-integer-text]
  • ...[dp_rank-integer-text]
  • ...[tp_size-padded-text]

handoff.py is byte-identical at 196fd3711 and at tip abb70ef48, so reinstating the old file is the pre-fix code at the current tip as well.

Principle 8: "Every claim carries its measurement." The PR's named result is confirmed.

2. Caller audit: does any real path send text?

No. Measured on the merged tree.

whole_number has exactly one caller, atom/compass/kv/connector.py:264-265. It reads config.tensor_parallel_size and config.parallel_config.data_parallel_rank. Here is every writer of those two fields:

source what it produces
CLI -tp / --data-parallel-rank (atom/model_engine/arg_utils.py:116-118, 156-157) type=int
ATOM_DP_RANK (atom/utils/envs.py:27) int(os.getenv(...))
arg_utils.py:712 ParallelConfig(**parallel_config_kwargs) built from the argparse values above
engine_core_mgr.py:295, 361, 1428 (DP-attention rewrite, per-rank deepcopy) integer arithmetic over iter_dp_rank_assignments
vLLM / sglang plugin (atom/plugin/config.py:344, 495, 596, 680) the host framework's int fields, or literals
Compass spec tensor_parallel_size is DEPLOYMENT_OWNED (atom/compass/spec/rules.py:81), so a spec is refused if it carries one

JSON from the wire. The one place a rank arrives as JSON is the atomesh request field data_parallel_rank, read by _get_data_parallel_rank in atom/entrypoints/atomesh/atom_standalone_service.py:1596. That field is a per-request routing hint: it goes onto the Sequence and never reaches parallel_config, so it never reaches whole_number. None of the json.loads-typed CLI flags (--online_quant_config, --hf-overrides, --dspark-config, --dcp-config, --eplb-config) sets a width. The connector never reads a width out of kv_transfer_config or out of an incoming blob.

A real Config cannot carry text anyway. I measured this on node 18 at the head:

  • Config(model=<Qwen3-0.6B>, tensor_parallel_size="8") raises TypeError: '<=' not supported between instances of 'int' and 'str'.
  • ParallelConfig(data_parallel_rank="2") raises TypeError: '<' not supported between instances of 'str' and 'int'.

So the removed branch was reachable only from a hand-built config double.

No text widths in the tree. A git grep at 196fd3711 over atom/, scripts/ and tests/ finds text widths only in the three tests this PR changes. At the head there are none outside the refusal parametrization.

Principle 6: "Refuse rather than fall back." Refusing text is not a behaviour change for any real path.

3. Did the removed or renamed tests hide coverage?

No. Principle 8. Here is where each of the 5 removed node ids went:

  • whatever_it_is_spelled_as[int] and [float] became as_an_int_or_a_whole_float[int] and [float]. The body is unchanged; only the decorator, name and docstring changed.
  • [text] ("8") and [padded] (" 8 ") were inverted into the refusal cases [tp_size-integer-text] and [tp_size-padded-text], and [dp_rank-integer-text] was added.
  • test_a_width_in_text_keeps_its_exact_value pinned that text was not routed through float. That path no longer exists. The defect it guarded against, text read through float, is still pinned by [tp_size-exponent-text] ("1e1") and [tp_size-decimal-text] ("8.0"), which both refuse at the head. An int input goes through int() as the identity, so large-int exactness has nothing left to lose.

test_the_ranks_the_router_reads_are_numbers still holds the conversion it exists for. I checked with a mutation on node 18: return whole changed to return value at L111, 111 lines kept. The result was 2 failed, 42 passed:

  • test_the_ranks_the_router_reads_are_numbers, failing assert (False);
  • test_a_whole_width_is_taken_as_an_int_or_a_whole_float[float], failing assert (<class 'float'> is int).

Both failures come from one mutation, which is the basis for the ponytail delete: inline comment.

4. Edge cases

Measured on node 18 by calling whole_number("tp_size", v) directly. The "before" column is 196fd3711's handoff.py.

input head before
8.5 refused, named ValueError refused
float("inf"), float("-inf") refused, named refused
float("nan") refused, named refused
-1, 0 accepted (-1, 0) accepted. Range is out of scope per #258.
True, False refused, named refused
numpy.int64(8) accepted, 8 (int) accepted
Decimal("8"), Decimal("8.0") accepted, 8 (int) accepted
Decimal("8.5"), Decimal("NaN"), Decimal("sNaN"), Decimal("Infinity") refused, named refused
"8.0" refused, named refused
"8", " 8 " refused, named accepted as 8
"8_0" refused, named accepted as 80
"٨" (Arabic-Indic 8) refused, named accepted as 8
numpy.str_("8") refused, named accepted as 8
b"8", bytearray(b"8") refused, named refused
numpy.float64(8.0), Fraction(8, 1) accepted, 8 (int) accepted
numpy.bool_(True) accepted as 1 accepted. Out of scope per #258; see the inline comment.
None, [8], complex(8, 0), numpy.array([8]) refused, named refused

What the table shows:

5. Docstring truth checks

Principle 8. Every changed sentence is true.

  • handoff.py, L87-91. "Text is refused, "8" included" is measured (table above). "CLI flags and ATOM_DP_RANK are parsed with int" is true: arg_utils.py:118,157 and envs.py:27. "Its Config raises on a string width" is measured: TypeError from both Config and ParallelConfig.
  • handoff.py, L93-98.
    • "8.5 would go out as 8" is true of int().
    • "A bool … would go out as 1 or 0" is true of a Python bool.
    • "No value that is accepted changes on the way to the int returned" holds in the == sense for every accepted row above.
    • "that check is what refuses integer text: int("8") is 8, which is not equal to "8"" is true. For "8" and " 8 ", int() succeeds and whole != value refuses.
  • Test, L251-257. "A whole float is converted to an int on the way in" is shown by the mutation. "Text is refused, which the test below pins" is true: the refusal test follows at L277.
  • Test, L302-304. "Text is refused whatever it spells" is measured, including numpy.str_, "8_0" and non-ASCII digits.
  • Test, L315 is true.

PR body claims. Every claim is confirmed except the gate count, which does not carry over and is re-gated below:

  • The before-probes are reproduced.
  • "Only three tests": git grep at 196fd3711.
  • "Refusal text unchanged": the template is untouched in the diff.
  • The line counts match git diff --stat (+10/−11 and +18/−38).

Other checks:

  • No design-doc references. Grepped at the head over the PR's whole file set (D\d+, P\d.\d, T\d+, "principle", "Gate N"): nothing. AI_DEV_RULES: "No design-doc references in code or runtime output."
  • Lint. ruff check and black --check on both files at the head: RUFF_RC=0, BLACK_RC=0.

Observation, not a finding. The PR keeps whole floats (8.0 becomes 8). No ATOM launch delivers a float either. But unlike text, a real Config does not refuse one: Config(tensor_parallel_size=8.0) gets past the width assert at config.py:1738 and only fails later, on rocminfo at L1750, in the CPU container. So a whole float is the one non-int a real Config can hand this function. The conversion is lossless and inside #258's brief ("keep int accepted"). Nothing is asked of this PR.

6. ponytail-review

Over 196fd3711..8bab7230b:

atom/compass/kv/handoff.py L88/L95/L97: shrink: "text is refused" said three times in one docstring. Drop "text;" from the Refused list.
tests/compass/test_kv_remote_prefill.py L313: delete: [float] repeats test_the_ranks_the_router_reads_are_numbers (same mutation reds both). Drop 8.0 and the parametrize.
net: -2 lines possible.

Both lines are non-blocking. The production code change itself (whole = int(value) and whole != value) is already minimal. AI_DEV_RULES gate 4: the reviewer "runs the ponytail-review skill over the diff… its findings are posted like any others."

7. Gate on the merged tree

Tree:

Run: the tree's own scripts/compass/gate_cpu.sh on xiaobizh_n18_cpu, bounded by timeout -k 10 3000, unpiped.

run result GATE_CPU_RC
merged fe9b53e7b 5253 passed, 155 skipped, 3 xfailed, 185 s 0

Decomposition against the developer's 5244 at 8bab7230b, per principle 7 ("Never report an aggregate without its decomposition"):

Timing classes, from this run's junit: all passed:

  • every node in TestTheRegionIsNotCopiedPerChunk;
  • TestNoSizeAtWhichACallStopsBeingOne::test_a_long_call_costs_what_it_is_and_not_its_square[minimax];
  • tests/test_gc_utils.py::test_freezing_twice_is_additive_and_harmless.

No re-run was needed.


Blocking findings: none.
Non-blocking: 3, all inline:

  1. shrink: docstring repetition at handoff.py L98.
  2. delete: the duplicate [float] case at test L313.
  3. File the numpy.bool_ gap, plus the range note if you want them together, as its own issue before compass(kv): widths written as integer text are still accepted, and a test docstring says they arrive that way #258 closes (handoff.py L101). AI_DEV_RULES: "A finding not fixed in the PR that found it gets an issue."

Next action: undraft and land per the landing rule. The approved tree is 545c49d12 on tip abb70ef48. If the tip moves again, recompute the merge tree.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant