Skip to content

Record the failed hosted trial of unittest-parallel on the two required test jobs - #652

Merged
seathatflowsinourveins merged 16 commits into
mainfrom
foundation/suite-parallelism-trial-outcome-20261003
Oct 4, 2026
Merged

seathatflowsinourveins merged 16 commits into
mainfrom
foundation/suite-parallelism-trial-outcome-20261003

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes, in one or two sentences: records the failed preregistered trial of unittest-parallel 1.8.6 on the two required test jobs (hosted run 37109532421 of draft Trial: unittest-parallel for the two required test jobs (draft, preregistered, never merged) #646, head 1e4bb5ab): a validated observed convergence record (decision reject, 36 observations), a dated decision record with the unit's completeness critique (not adopted; PR-A and PR-A2 are not opened), a durable sanitized receipt, supporting artifacts with the sha256 of the originals that expire on 2026-11-02, byte copies of the 15 files the preregistration froze (the observed record's frozen inputs), and four anti-pattern rows. Docs and evidence only: no workflow, test, ruleset or setting changes.
  • Base commit: 4ced2923063db6a6dcafa9f25af5ee05a4153c75
  • Lane: lane:foundation
  • Owned paths touched: docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md, evidence/receipts/suite-parallelism-trial-20261003.json, evidence/artifacts/suite-parallelism-trial-20261003/ (13 files, among them the observed convergence record experiment.json), evidence/artifacts/suite-parallelism-trial-20261003/frozen/ (15 byte copies, stored as .txt, except the two lock copies, which are named *-pins.frozen so no scanner pattern matches them), docs/harness-defaults.md (five rows at the top of the anti-pattern table), manifests/evidence.json (registration and the convergence_records entry only, the branch's last commit, per the hot-file protocol in docs/lanes.md)

SOTA sources

  • craigahobbs/unittest-parallel at release commit bda5d77dc1a2fa2df90f5f5a7de297ea375e345c: src/unittest_parallel/main.py lines 150-160 (spawned pool, pool.map at 160), 164-211 (report only after the map), 240-242 (exit status), 62-63 (--thread); README.md lines 43-71; issue 27; commits after the release; PyPI JSON (1.8.6 latest on 2026-10-03).
  • python/cpython at v3.12.3 (f6650f9ad73359051f3e558c2431a109bc016664): Lib/unittest/async_case.py lines 35-38, Lib/unittest/case.py lines 159-161, 176-182 and 612-617 (a class-level skip only marks the class and applies when a test runs), Lib/unittest/loader.py line 94 (one instance per test name), Lib/multiprocessing/pool.py lines 367, 528-546, 774 and 809-831, Modules/_pickle.c lines 3646-3704; at v3.13.16 (cbc944f4bc59639a444dd971c737788ba2283a91): Lib/unittest/async_case.py line 42.
  • github/docs at 068546469ae7f079368d12f969991a045121c4a0: data/reusables/actions/workflows/triggering-a-workflow-paths5.md (a pull request's paths filter uses the three-dot diff). GitHub REST: workflow jobs, artifacts.
  • Completeness critique, read on 2026-10-03 after the trial: repository files with GET requests at these pinned commits (each tag resolved through the GitHub commits API): pytest-dev/pytest 9.1.1 at cf470ec0bf7eb89cd97dd56df4859eae5db46447; pytest-dev/pytest-xdist v3.8.0 at 1e3e4dc16523c8a8f6c67d95a950166420718c99; mtreinish/stestr 4.2.1 at 2802f1425f45b9ad0857ff4685ab469dcd2cf2e0 (stestr/output.py lines 154-161); testing-cabal/testtools 2.9.1 at 088c98e24961ecf6d94ea5204457f2dcffe2f1c6; nose-devs/nose2 0.16.0 at c93ca65ac6f62475209aa8944fda14c671615549; cgoldberg/concurrencytest 0.1.11 at 266e27c833f9b1ecb47142bdddbac5199bebab72 (concurrencytest.py lines 89-90, 158-160 and 179-187); testing-cabal/subunit 1.4.6 at c85605280cb1d975b3076b9ad7ac38f17e858945; CleanCut/green 4.0.2 at 4de285b05b8e6b161159be2241b219dc5176f0e5; zopefoundation/zope.testrunner 8.3 at 1061ccc4b6b5a824870f2142bc948863b7ba1f70; python/cpython v3.13.16 at cbc944f4bc59639a444dd971c737788ba2283a91 (Lib/unittest/loader.py, Lib/test/libregrtest/); craigahobbs/unittest-parallel at bda5d77d (main.py:123-128, :142-160, README.md:59-70); github/docs at 2bd66de8cea336061c9ea060c9b37385136e6ab3 (workflow-syntax.md: background, wait, wait-all, cancel, parallel; limits.md lines 64-69, concurrent macOS jobs per plan). With GET requests that no commit pins: the GitHub changelog of 2026-06-25; the GNU findutils 4.10.0 announcement; PyPI JSON per release and the repository records; unittest-parallel's issue tracker (issue 22, issue 27 and a search for "pickle") and CPython's (python/cpython#83438, bpo-39257, and searches for IsolatedAsyncioTestCase with pickle, pickling or multiprocessing). Locally: the installed GNU xargs 4.10.0 manual page (the gnu.org manual did not answer).
  • In-repository, at the base commit: docs/acceptance-evidence-policy.md:55-63 (preserve failed attempts, declare every sanitization), docs/evidence.md:18 (compatibility_attempt), docs/lanes.md:94-141 (hot-file protocol), blueprints/convergence-practice/contract-reference.md:27-28 (distinct source, input and evaluation artifacts) and :84-92 and scripts/validate_convergence.py:99-123 and :136-146 (usage fields, failure coverage, run status, scope of qualifying runs), tools/sota-convergence/landscape-sweep/build_inputs.py:21-22 and :225-234 (the seeds format); for matrix sharding, .github/main-ruleset.json:35 and :41, docs/github-automation.md:26-28 and tests/test_workflow_hardening.py:1249-1257. At the trial head 1e4bb5ab: the 15 frozen files (byte-copied under frozen/) and the preregistration's decision record, "Alternatives".

Evidence-class table

Claim Evidence class Command / receipt
The observed convergence record: 36 observations, of which 12 are timed ubuntu-24.04 runs (S passed 3 of 3; L4, L4F and L4C failed 9 of 9 with exit 1), 4 are ubuntu-24.04 control jobs (passed; each observation's exit code is its control step's own exit, and the commands' own exits and the controls file's step seconds, 0, 0, 1 and 0 s under S, L4, L4F and L4C, are in its scope text) and 20 are macos-15 jobs that never received a runner (skipped, no exit code, no quality); decision reject, no qualification, token usage null (unknown), native_retries 0 for the 16 executed jobs, runner time in each executed observation's scope text (11,461 s over the 16 observed jobs); frozen_inputs cite byte copies of the 15 files the preregistration froze (3 sources, 3 inputs, 9 evaluation, in its order, each with its preregistered sha256) and the preregistration copy at the end of sources; every departure from the plan stated in limitations native_cli_execution observations generated from the raw run directories, the jobs API (GET, steps included) and result.json; scripts/validate_convergence.py checks declared consistency and artifact hashes, not truth evidence/artifacts/suite-parallelism-trial-20261003/experiment.json
The 15 byte copies under frozen/ equal their originals at 1e4bb5ab byte for byte and carry the preregistered sha256 values (15 of 15); all pass scripts/validate.py --scan-file; 264,054 bytes reproducible artifact check (git show at 1e4bb5ab against each copy, and sha256 against the preregistration's frozen_inputs) evidence/artifacts/suite-parallelism-trial-20261003/frozen/; receipt data.retained_files
Completeness critique: the missed modality (no whole-suite local run under the candidate runner; the hosted behaviour of background steps never examined), the sources missed before the trial (not recorded as read by the preregistration or its decision record: unittest-parallel's issue tracker, whose issue 27 is titled with the trial's exact exception; CPython's tracker, issue 83438; the suite's own async and type()-built test modules; GitHub's changelog), nine candidate classes the trial did not evaluate plus matrix sharding with its blockers, each relayed claim marked VERIFIED or LEAD-ONLY, and seeds for the next ci-supply-chain sweep source review, by source class: repository files with GET requests at the pinned commits above; repository records, PyPI JSON, the GitHub changelog, the GNU findutils announcement and the two issue trackers with GET requests that no commit pins; the GNU xargs exit statuses from the installed GNU xargs 4.10.0 manual page; the unittest.mock count (106 of 234 test modules, 100 at module level), the suite's modules and the matrix-sharding blockers locally with git grep and reads at 1e4bb5ab and the base. The candidate list is a lead from a read-only single-family Codex lane decision record, "Completeness critique"
On ubuntu-24.04 (CPython 3.12.3, image 20260927.320.1) S ran 9,823 tests, OK (skipped=967), exit 0, in 3 of 3 runs (1,688, 1,659 and 1,186 s); all 9 unittest-parallel runs (L4 649/438/459 s, L4F 539/692/710 s, L4C 667/659/704 s) exited 1 with TypeError: cannot pickle '_contextvars.Context' object from pool.map (main.py:160); 199 ids never ran at module level and 82 at class level (in S, 90 of the 199 and 8 of the 82 were ok and the rest skipped, all 66 IsolatedAsyncioTestCase tests among the skipped); oracle: ubuntu-24.04 reject, macos-15 no verdict; id mapping ok native_proven (hosted run 37109532421, attempt 1; receipt retained with hashes and sanitized excerpts) evidence/receipts/suite-parallelism-trial-20261003.json; evidence/artifacts/suite-parallelism-trial-20261003/
Every control and crash control passed: the controls file reported Ran 8 and FAILED (failures=2, errors=2, skipped=1, expected failures=1, unexpected successes=1), exit 1 under S and 5 under unittest-parallel; the crash control exited 3 under S and 124 (the 300 s bound) under every parallel arm synthetic: synthetic fixtures executed on the hosted runner (run 37109532421) receipt data.controls; control-logs.txt
K4 timing FAIL in 3 of 3 L4C runs; test_cli_refuses_live_base_url_from_env_file_without_network ERROR in 3 of 3 L4F runs native_proven (observed); causes not established, the L4F one unknown receipt data.non_missing_mismatches
Jobs-API test-step seconds equal each run's own; the 20 macOS jobs never received a runner; 5 of this repository's macOS jobs held runners in 46 of 50 minutes of the queued window independent observation of platform records (read-only gh api --method GET); nearest template class source_review receipt data.jobs, data.cancellation
Root causes: IsolatedAsyncioTestCase stores a contextvars.Context in __init__, so its suites cannot be pickled even when every test in them is skipped, as all 66 were on ubuntu-24.04; four type()-built classes in tests/test_adoption_version_probes.py are not importable under their __qualname__ local_integration (toy modules on CPython 3.12.3 and 3.13.16; the three real modules on 3.12.3 only; the coordinator also saw the version-probes PicklingError on 3.13.16, an output not in the retained artifacts) plus source_review (async_case.py, case.py, loader.py, pool.py, _pickle.c at pinned tags) local-reproductions.txt; receipt data.root_causes
Upstream issue 27 (closed by its reporter, no maintainer comment), README silent on asyncio and pickling, 2 commits after the release touching only AGENTS.md, PyPI latest 1.8.6 source_review (GET reads on 2026-10-03) receipt data.upstream
Speed ratios (L4 0.277 and 0.547; L4F 0.417 and 0.599; L4C 0.402 and 0.594) computed from measured seconds; not an adoption claim (no arm eligible, runs incomplete, three runs per arm, and each slowest-over-fastest ratio divides by the single fastest S run, 1,186 s, inside a serial spread of 1,186 to 1,688 s) receipt data.arms_ubuntu_24_04
Usage: the hosted jobs ran deterministic commands and no model; runner time is the per-job seconds of the 18 ubuntu-24.04 jobs (11,524 s in total); the 20 macOS jobs were never assigned a runner; the model usage of the planning, build and review sessions is unknown and not counted as zero independent observation of platform records (jobs API) for the job seconds; unknown for model usage receipt data.usage, data.jobs
The run's cancellation at 09:12:40Z coordinator decision, not part of the preregistered rule; recorded with the run and job records decision record, "The cancellation"
Repository checks on this branch local_integration commands below

Local commands run

$ python3 scripts/validate.py
passed (69 components, 187 receipts, 9,450 hashed files); exit 0
$ python3 scripts/validate_convergence.py evidence/artifacts/suite-parallelism-trial-20261003/experiment.json --root . --json
valid: true, 36 observations; exit 0
$ python3 scripts/validate_convergence.py --all-recorded --root . --json
valid: true, 27 records (this record among them, 36 observations); exit 0
$ python3 scripts/evidence_manifest.py --check
{"files": 9450, "status": "passed"}; exit 0
$ python3 scripts/component_matrix.py --check
{"rows": 32, "status": "checked"}; exit 0
$ python3 scripts/new_host_grand_list.py --check
{"status": "passed", "layers": 32, "winners": 66}; exit 0
$ python3 -m unittest tests.test_osv_lockfile_coverage.LockfileInventoryTests.test_every_tracked_lockfile_and_manifest_is_listed tests.test_blind_checkout.RepositoryClassificationTests.test_every_blueprint_value_under_a_label_key_is_classified tests.test_workflow_security_coverage.NewWorkflowSecurityCoverageTests.test_all_published_workflows_are_listed_and_covered
Ran 3 tests, OK (zizmor 1.30.1 and actionlint 1.17.0 on PATH); exit 0
$ python3 -m unittest discover -s tests -p test_adoption_docs_consistency.py
Ran 39 tests, OK (skipped=1); exit 0
$ python3 -m unittest tests.test_blind_checkout tests.test_ecosystem_manifest tests.test_jcodemunch_config tests.test_verdict_review_gate tests.test_gitleaks_config tests.test_validate tests.test_claude_repository_evidence tests.test_workflow_hardening tests.test_osv_lockfile_coverage tests.test_workflow_security_coverage tests.test_pre_push_gate tests.test_new_host_grand_list tests.test_catalog_freshness_propose tests.test_convergence_contract tests.test_evidence_manifest tests.test_lane_packets tests.test_claude_recovery tests.test_gpt6_family_tiering_20260926 tests.test_local_inference_latest_20260926
Ran 927 tests, OK (skipped=30); exit 0 (at the head 405014de)
$ python3 scripts/validate.py --scan-file <each of the 31 added or changed files other than manifests/evidence.json>
passed, 31 of 31; exit 0
$ for each frozen copy: git show HEAD:<copy> | sha256sum, against git show 1e4bb5ab:<original> | sha256sum and the preregistered sha256
15 of 15 equal (264,054 bytes); exit 0
$ rg -n -i -f <pattern file kept outside the repository> <every added line of git diff origin/main HEAD, 13,371 lines>   # host path prefixes, the user and host names, the email identifier, the account's private repository names (word-bounded), dollar figures (\$[0-9])
4 matches; exit 0: the shell positional parameters kind="$1" and control_file="$2" at lines 252-253 and 452-453 of frozen/github-workflows-macos-suite-parallel-trial.yml.txt, a byte copy that cannot change without breaking its preregistered hash; they are not dollar figures. The same patterns over the other 13,367 added lines: no match; exit 1
observed-record generator kept outside the repository: re-derives every per-run value from the raw run directories (exit-code.txt, step-seconds.txt, command.txt, meta.json, runtime.json, git-status.txt, log.txt), the jobs API read with GET (steps included) and result.json, independently of the receipt, re-hashes the 15 frozen files of the preregistration at 1e4bb5ab and checks each byte copy under frozen/ against them
194 of 194 checks (exit codes, test-step seconds equal to the jobs API's, job seconds, missing ids per run equal to compare.py's, crash lines, control exits, macOS jobs without a runner, exact speed ratios, 15 of 15 frozen hashes, 15 of 15 byte copies); exit 0
consistency pass kept outside the repository over the decision record, the receipt, the observed record, the artifacts README, the draft thread replies and this description: the observed record against the re-derivation, every number repeated across them, each missing id's serial outcome parsed from the three S logs (module level 90 ok / 109 skipped, class level 8 ok / 74 skipped, identical in all three S runs), the controls files' raw step seconds (0, 0, 1 and 0 s), each frozen copy against git show at 1e4bb5ab and the preregistration's role lists, the receipt's retained-file hashes (28 files, recursively), the new sections and the superseded phrasings absent
655 of 655 checks; exit 0
$ git fetch origin main; git merge-tree --write-tree --name-only origin/main HEAD
main has not moved (4ced2923); clean, exit 0

Local reproductions (unittest-parallel 1.8.6, coverage 7.16.2):
$ python -m unittest_parallel -j 2 --level module -v        # toy module: one plain and one IsolatedAsyncioTestCase class
TypeError: cannot pickle '_contextvars.Context' object; exit 1 (CPython 3.12.3 and 3.13.16)
$ python -m unittest_parallel -j 2 --level class -v         # same toy module
the plain test ok, then the same TypeError; exit 1 (both)
$ python -m unittest_parallel -j 2 --level module -v        # plain-only toy module
Ran 2 tests, OK; exit 0 (both)
$ python -m unittest_parallel -j 2 --level module -s tests -t . -p test_adoption_version_probes.py
_pickle.PicklingError: Can't pickle <class 'tests.test_adoption_version_probes.LinuxReportUnderbash32Tests'>: attribute lookup LinuxReportUnderbash32Tests on tests.test_adoption_version_probes failed; exit 1 (3.12.3)
$ python -m unittest_parallel -j 2 --level module -s tests -t . -p test_adaptive_paper_fees.py   # and test_adaptive_paper_transport.py
TypeError: cannot pickle '_contextvars.Context' object; exit 1 each (3.12.3)
$ python -m unittest_parallel -j 4 --level class -v -s tests -t . -p <each of the three modules>
6 of 22, 54 of 61 and 57 of 116 results, then the PicklingError or the TypeError; exit 1 each (3.12.3)
$ python -m unittest tests.test_order_throughput
$ python -m unittest_parallel -j 2 --level module --disable-process-pooling -s tests -t . -p test_order_throughput.py
Ran 76 tests, OK (skipped=1); exit 0 each (3.12.3 and 3.13.16)

Read-only GETs on 2026-10-03: gh api --method GET for the run, its 38 jobs (again with steps for the observed record), its 18 artifacts, PR #646 and the review comments of this PR, the macOS jobs of the repository's runs around the window, craigahobbs/unittest-parallel (repository, tags, issue 27 and its comments, compare, contents at bda5d77), python/cpython (including Lib/unittest/case.py and loader.py at v3.12.3, and loader.py and Lib/test/libregrtest at v3.13.16), github/docs contents at the pins, and the critique's sources at the pinned commits above (contents, tag-to-commit lookups and repository records); PyPI JSON, the GitHub changelog page and the GNU findutils announcement with curl (the gnu.org xargs manual did not answer; the installed GNU xargs 4.10.0 manual page was read instead). For the delta review, also with gh api --method GET (queries in the URL, no -f or -F): stestr/output.py and concurrencytest.py at their full release commits and the two tag-to-commit lookups, unittest-parallel's issue list, its issue search for "pickle" and issue 27 with its comment, python/cpython issue 83438 with its comments and eight issue searches, one code search of github/docs, the github/docs commit lookup (2bd66de8) and limits.md at 2bd66de8, the state of PRs #646 and #652 and the repository's owner type; and the GitHub changelog page again with curl.

Decision record

docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md: the decision (not adopted), the observed convergence record it cites, what was tried with the preregistered rule and a pointer to #646 at 1e4bb5ab, what happened, the cancellation and why, root causes, what the attempt does not show, departures from the preregistration (among them where the observed record lives and the seven departures of its shape), the completeness critique (missed modality, missed sources, candidate classes including matrix sharding) with its seeds, overturn conditions and next-trial preconditions.

Host evidence

Not applicable: no files under evidence/hosts/ change.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no workflow changes in this PR).
  • New/changed workflows declare top-level permissions: contents: read (no workflow changes in this PR).
  • No secrets are printed, logged or committed; no new required secret was added.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

Notes for the merger: manifests/evidence.json is the shared hot file; if main moves, take main's copy, register again docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md, docs/harness-defaults.md, evidence/receipts/suite-parallelism-trial-20261003.json and the 28 files under evidence/artifacts/suite-parallelism-trial-20261003/ (15 of them under frozen/), append evidence/artifacts/suite-parallelism-trial-20261003/experiment.json to convergence_records again, then run the --write commands, scripts/validate.py and scripts/validate_convergence.py --all-recorded --root . --json (docs/lanes.md). #632 inserts five rows at the same top of the anti-pattern table, so whichever of the two merges second has a trivial table conflict (keep both sets of 2026-10-03 rows). This branch was rebuilt on main at 4ced2923 (after #634 and #636), with this PR's five rows above the 2026-10-02 rows; when main moves again, git merge-tree is likely to report conflicts in those same two files (keep both sets of rows; take main's manifests/evidence.json and register again as above). #646 was closed unmerged at 11:05:52Z on 2026-10-03, ten seconds after this PR opened; do not reopen it or open another pull request that adds its workflow, because either re-runs the whole trial. The record names no private repository and no dollar figure.

🤖 Generated with Claude Code

seathatflowsinourveins and others added 5 commits October 3, 2026 07:04
…tized artifacts

Hosted run 37109532421 of draft PR #646 (head 1e4bb5a) failed its preregistered
rule. On ubuntu-24.04 the serial arm passed 3 of 3 runs (1,688, 1,659 and
1,186 s); all nine unittest-parallel 1.8.6 runs (L4, L4F, L4C) ended with
TypeError: cannot pickle '_contextvars.Context' object from pool.map, with 199
(module level) or 82 (class level) ids never run. The oracle gave
ubuntu-24.04 reject and macos-15 no verdict; the coordinator cancelled the run
after the Linux jobs finished and before any macOS job received a runner.

Root causes, reproduced locally on CPython 3.12.3 and 3.13.16: the two
IsolatedAsyncioTestCase classes hold a contextvars.Context, and four classes
built by type() in tests/test_adoption_version_probes.py are not importable
under their qualified names. Nothing is adopted.

The receipt keeps the run and job records, per-run data, missing ids per class,
the non-missing mismatches, the crash signature, the local reproductions and
the upstream reads with pins; the artifacts keep sanitized excerpts, a trimmed
oracle output and the hashes of the originals, which expire on 2026-11-02.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two dated 2026-10-03 rows at the top of the anti-pattern log: a hosted
whole-suite trial preregistered after a preflight of one 29-test module (the
whole suite never ran locally under the candidate runner), and a trial
workflow whose pull_request paths filter names its own file, so the record
path has to be decided before the first push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e review

Independent review of the records pull request (one medium, seven low
findings) checked against the raw run artifacts; the decision is unchanged.

- Serial outcome of the ids that never ran (F1), re-derived from the three
  S logs and equal in all three: module level 90 ok and 109 skipped, class
  level 8 ok and 74 skipped. All 66 IsolatedAsyncioTestCase tests and the
  8 bash32 tests were skipped on ubuntu-24.04, so the async classes break
  unittest-parallel at pickling although none of their tests runs there
  (a class-level skip applies only in TestCase.run; CPython v3.12.3
  case.py:159-161 and :612-617, loader.py:94). Per-class counts are in the
  receipt's data.missing_ids and the trimmed result; the serial seconds
  sentence now names the two missing ids in S-r3's durations table.
- Ids with a result record versus records (F2): L4C-r3 has 9,743 records
  for 9,741 ids (three K4 subtest FAIL records).
- n=3 limitation (F3), controls labelled as synthetic fixtures executed on
  the hosted runner and reproduction scope per interpreter and module (F4),
  usage statement with 11,524 ubuntu-24.04 job-seconds and unknown session
  model usage (F6), runner image recorded only by the 16 arm and control
  jobs (F7), and the two macOS phrasings (F8).
- retained_files hashes updated for README.md and result-trimmed.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whole suite was not run locally under the candidate runner, so the
row now says a local whole-suite run would very likely have shown the
crash, inferred from the per-module reproductions. It attributes the
TypeError to the two IsolatedAsyncioTestCase modules and the masked
PicklingError to the dynamic-class module (first-exception rule,
Lib/multiprocessing/pool.py:822 at CPython v3.12.3). The unmeasured "in
seconds" is dropped. Row format and position unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the trial environment, not because of the operating system

Delta review finding on the outcome record: the skip is skipUnless(HAS_SDK), true only when both alpaca and requests import; the trial arms installed neither (requirements-ci.txt installs neither), so the 66 tests were skipped in every S run in that environment. Reworded the two sentences that tied the skip to ubuntu-24.04.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Oct 3, 2026
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T11:13:40.715580Z 571520b PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 571520b084

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md
Comment thread docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md Outdated
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Exact-head SOURCE ACCEPT at 571520b084c2710d4619c4ecdd8ef9b0a9e1df62, base/current main read 4ced2923063db6a6dcafa9f25af5ee05a4153c75. Scope: publication and classification of the failed suite-parallelism attempt. This supplies no adoption, macOS execution, SDK qualification or fresh test run.

I checked the actual pinned publication, original hosted compare artifact, frozen preregistration and GitHub run/job records. The decision follows the recorded eligibility rule: nine Linux candidate runs exit 1 and remain ineligible; three serial runs exit 0 and are eligible; Linux is rejected. Twenty macOS jobs have no assigned runner and zero steps, so macOS has no verdict. Candidate timing below the speed thresholds does not qualify an incomplete, crashing suite.

The frozen experiment equals the retained byte copy: 43,208 bytes, SHA256 305fc2fa2ddd2c0a0eba4d85d748807db87302bde59a6b6489de866bee416c08. The original run is attempt 1 at full head 1e4bb5ab7c79faa83e4cd8d04cf8dc6d33c7b0de. Original compare artifact 11270430044 has ZIP SHA256 bc0136254601b3ece5ecfec7c2a0525bcbc60c0f606de1b03c05ff7c729d6f2c; its result is 815,073 bytes/SHA256 f26cb0c1c460029ffd481a844bc524ae5e83c84b73fda52cc3efe3496475eb9a. The six declared unchanged fields, all retained fields of twelve runs, reconstructed missing-id sets and their serial outcome counts match the original. Summary and checkout-status byte copies match too.

Original job records reconcile all twelve arm-step durations and the 11,524 seconds across eighteen Linux jobs. Serial arm-step durations are 1,688/1,659/1,186 seconds. Step success alone is not an arm exit-code proof: this trial intentionally records failing arm exits inside a successful collection step. The original result retains the actual arm exits. The publication explicitly records cancellation outside the preregistered rule and the changed record location. Those departures stay visible. The proposed cross-platform failure is explicitly an inference; it supplies no macOS run.

The source mechanism is supported by unittest-parallel at bda5d77dc1a2fa2df90f5f5a7de297ea375e345c, which sends suites through a spawned pool before reporting, and CPython at f6650f9ad73359051f3e558c2431a109bc016664, which stores a context while constructing an async test instance. The L4F error's cause and K4 contention explanation remain unknown; no suppressed traceback is invented.

All 27 actual Git source identities verify. The registry preserves all 9,420 base paths in relative order, adds fourteen, replaces only the owned harness-defaults binding, and retains all 187 receipt records and 26 convergence records unchanged. All fifteen content bindings and twelve retained-file bindings match their actual bytes. Preserve current-main foreign rows when composing later merges.

The evidence classes and usage correctly distinguish repository execution, synthetic controls, local reproductions, source inspection, independent platform observations and unknown planning/build/review model usage. These are not upstream acceptance. I did not redownload the twenty full raw logs or rerun the trial.

Verification custody: 34 native GH commands exit 0, retained stdout/stderr manifest SHA256 9b3734aad72471ddfc2e9e0c5143d38a962ff0715daa0cc33d23690668e76297; original artifact and source bytes were independently checked by root. Source ACCEPT is limited to this failed-attempt publication and its declared boundaries. Required hosted checks and fresh-main composition remain merge conditions.

seathatflowsinourveins and others added 7 commits October 3, 2026 08:07
…itique of the suite-parallelism trial

The trial's convergence record is now evidence/artifacts/suite-parallelism-trial-20261003/experiment.json
(status observed, decision reject), generated from the raw run directories, the jobs API and result.json.
It has 36 observations: the 12 timed ubuntu-24.04 runs, the 4 ubuntu-24.04 control jobs and the 20 macos-15
jobs that never received a runner (skipped). It carries the preregistered metrics, commands, lane, roles and
base revision unchanged; its frozen_inputs cite the retained preregistration copy, whose 15 frozen hashes were
re-checked at 1e4bb5a. Token usage is null (unknown), not the preregistered observed zeros, and runner time
is in each observation's scope text; both departures are in its limitations and the decision record.

The decision record, the receipt and the artifacts README cite it as the convergence record; the receipt
keeps the detailed data and the claim text, and no longer says it replaces the observed record. The decision
record's new "Completeness critique" section names the missed modality, nine candidate classes with each
relayed claim marked VERIFIED or LEAD-ONLY against its pinned upstream source, and seeds for the next
ci-supply-chain landscape sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A free-form receipt in place of the preregistered observed convergence record, and a substantive trial closed
without its completeness critique. Both rows sit above this pull request's two earlier 2026-10-03 rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The table now has four rows dated 2026-10-03, so the two references to "the anti-pattern row of 2026-10-03"
name the row by its title.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…en inputs

The observed record's frozen_inputs pointed only at the preregistration copy and listed result-trimmed.json, the
oracle's output, as an evaluation input, so scripts/validate_convergence.py checked no byte of the preregistered
oracle, controls, B1 fix or lock files (delta review of #652, D6).

- frozen/ holds byte copies, stored as .txt, of the 15 files the preregistration froze, written with git show at the
  trial head 1e4bb5a (draft #646, closed unmerged): 15 of 15 byte-equal to their originals and to their preregistered
  sha256, 264,054 bytes, all passing scripts/validate.py --scan-file. No lock copy's name starts with requirements, so
  tests/test_osv_lockfile_coverage.py does not treat it as a lockfile.
- experiment.json, regenerated from the raw run data: frozen_inputs cite the copies in the preregistration's role
  lists and order (3 sources, 3 inputs, 9 evaluation) plus the preregistration copy at the end of sources;
  result-trimmed.json stays an observation artifact only. A new limitation lists the other departures from the plan
  (command 11, one scope text per run, the job and step timestamps kept in the receipt's data.jobs), and each control
  observation's scope states its controls file's step seconds (0, 0, 1 and 0 s under S, L4, L4F and L4C; verifier
  note). rollback records #646 as closed.
- README: a frozen/ section says what each copy is, that the originals live on the closed trial branch and that the
  hash equality with the preregistration is the point.
- Receipt: data.retained_files lists all 28 files under the artifacts directory with their new hashes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ow each claim was read

Delta review of #652 (D1 to D5) and two corrections found while verifying it.

- D1: a "Sources missed before the trial" subsection names four sources that the preregistration and its decision
  record do not record reading (plan W1 and its Codex review, kept outside the repository, were not re-read):
  unittest-parallel's issue tracker (25 issues; issue 27, filed 2025-08-24, is titled with this trial's exact
  exception, and issue 22 reports a pickling failure of the family of root cause (b)), CPython's tracker
  (python/cpython#83438, whose only reply calls the pickling error expected: "contextvars are not compatible with
  multiprocessing"; no issue on pickling IsolatedAsyncioTestCase was found), this suite's own async and type()-built
  test modules, and GitHub's changelog. Matrix sharding, which plan W1 compared and the preregistration rejected, is
  listed with its blockers re-checked at the base (.github/main-ruleset.json:35 and :41, docs/github-automation.md:26-28,
  tests/test_workflow_hardening.py:1249-1257, GitHub's five-job macOS limit at github/docs 2bd66de8 limits.md:64-69
  and the trial's measured macOS queue) and noted as outside the one-job task; a seed and a critic question follow.
- D2: stestr's reader hang is VERIFIED at stestr/output.py:160-161 (4.2.1), and concurrencytest's repeated class
  fixtures under its default round-robin partitioning at concurrencytest.py:158-160 (0.1.11); only the lead's
  module-fixture wording stays LEAD-ONLY.
- D3: the critique, the Evidence class and the SOTA sources say how each class of claim was read: repository files
  with GET requests at the release tag or commit; repository records, PyPI JSON, the changelog, the GNU announcement
  and the issue trackers with unpinned GET requests; the GNU xargs statuses from the installed manual page; the
  unittest.mock count and the matrix-sharding blockers locally.
- D4: the second trial's local preflight is attributed to the coordinator and marked not verified here.
- D5: the departures of the observed record's shape list all seven (command 11, the control observations, the skipped
  macOS observations, one scope text per run, the frozen inputs at other paths, the timestamps in the receipt, null
  token usage), with runner time in each executed observation's scope text.
- Corrections: 106 of the 234 test modules import unittest.mock (100 at module level, 6 inside a test), not 100; and
  #646 was closed unmerged at 11:05:52Z on 2026-10-03, after this record's pull request opened.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ply the last delta-review wording fixes

The two requirements-style lock copies (byte copies, same sha256) are renamed trial-ci-suite-<os>-pins.frozen: a .txt name containing requirements is what Python dependency scanners and the OSV inventory test match, so the .txt suffix alone did not keep them out. README and experiment.json now say what each name rule does and that the dependency-graph effect is expected, not verified. Also: the sources-missed list now includes the pytest and coverage reads the preregistration recorded, matrix sharding's attribution names the concurrent Adoption runs, the receipt separates pinned source reads from unpinned GETs, the experiment record hedges the absence claim, and the anti-pattern log gains the frozen-inputs row.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…(hot-file protocol; last commit)

One registration on main's manifest at 4ced292: the decision record, the receipt, the artifacts including the 15 frozen byte copies (the two lock copies renamed so no scanner pattern matches them), docs/harness-defaults.md (five rows), and the observed record in convergence_records.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Scout and others added 4 commits October 4, 2026 15:44
…ries to the bytes after a986c4c (review 652 P1)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry plus the owned rows)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore the final hot-file commit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (hot-file protocol: every hot-file edit in the last commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-5f: landing at head e87dd3b9be8a8476b160a862aba2d1fae3824947. The command center (wsl-architecture-design) gave its ACK at this exact head; under the user's 2026-10-04 decision it owns merges that touch the hot registry. The read-only cross-family review (GPT-6.1 Sol max, the packaged Codex SDK worker) returned FINDINGS at 13e084c93aab; this head is reached from it through recorded carry edges (equal owned patch-ids, or a cross-family delta read returning ACCEPT at the edge's target), with its 1 P1 item(s) resolved in one repair round as recorded.

Observed main 14048b840425c2569e0df60a6596e94e601da15b. coordination/merge_tree_landing_check.py (sha256 a11649fcf2e5844c…) exit 0:

main 14048b840425c2569e0df60a6596e94e601da15b head e87dd3b9be8a8476b160a862aba2d1fae3824947 base 14048b840425c2569e0df60a6596e94e601da15b merged-tree 3c54d436ffbc01183bce44f7e328d4cd46dcd890 merge-tree-exit 0
ok   1: clean three-way merge
ok   2: merged-vs-main paths 32, outside PR-owned 0 []
ok   3: main drift 0 paths, overlap with PR-owned inputs (registry excluded) []
ok   4: registry foreign rows equal True, order preserved True, PR-owned rows 31, unowned top-level keys differing []
ok   5: merged files[] sorted by path with no duplicates: True (9926 rows)
LANDABLE
condition 5: merged files[] sorted and unique (tree 3c54d436ffbc01183bce44f7e328d4cd46dcd890)

Required checks at this head: 7 pass 1 skipping . Unresolved review threads: 0.

@seathatflowsinourveins
seathatflowsinourveins merged commit 4cfac6c into main Oct 4, 2026
25 checks passed
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Claude session native-agent-stack-5f: post-merge observation. Landed as 4cfac6c4add2446c315882be72477f7dff65a60f (parent 14048b840425c2569e0df60a6596e94e601da15b), tree 3c54d436ffbc01183bce44f7e328d4cd46dcd890. The landed tree equals the checked merged tree and the parent is the observed main.

@seathatflowsinourveins
seathatflowsinourveins deleted the foundation/suite-parallelism-trial-outcome-20261003 branch October 4, 2026 20:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant