Repository navigation
Record the failed hosted trial of unittest-parallel on the two required test jobs - #652
Conversation
…tized artifacts Hosted run 37109532421 of draft PR #646 (head 1e4bb5a) failed its preregistered rule. On ubuntu-24.04 the serial arm passed 3 of 3 runs (1,688, 1,659 and 1,186 s); all nine unittest-parallel 1.8.6 runs (L4, L4F, L4C) ended with TypeError: cannot pickle '_contextvars.Context' object from pool.map, with 199 (module level) or 82 (class level) ids never run. The oracle gave ubuntu-24.04 reject and macos-15 no verdict; the coordinator cancelled the run after the Linux jobs finished and before any macOS job received a runner. Root causes, reproduced locally on CPython 3.12.3 and 3.13.16: the two IsolatedAsyncioTestCase classes hold a contextvars.Context, and four classes built by type() in tests/test_adoption_version_probes.py are not importable under their qualified names. Nothing is adopted. The receipt keeps the run and job records, per-run data, missing ids per class, the non-missing mismatches, the crash signature, the local reproductions and the upstream reads with pins; the artifacts keep sanitized excerpts, a trimmed oracle output and the hashes of the originals, which expire on 2026-11-02. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two dated 2026-10-03 rows at the top of the anti-pattern log: a hosted whole-suite trial preregistered after a preflight of one 29-test module (the whole suite never ran locally under the candidate runner), and a trial workflow whose pull_request paths filter names its own file, so the record path has to be decided before the first push. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e review Independent review of the records pull request (one medium, seven low findings) checked against the raw run artifacts; the decision is unchanged. - Serial outcome of the ids that never ran (F1), re-derived from the three S logs and equal in all three: module level 90 ok and 109 skipped, class level 8 ok and 74 skipped. All 66 IsolatedAsyncioTestCase tests and the 8 bash32 tests were skipped on ubuntu-24.04, so the async classes break unittest-parallel at pickling although none of their tests runs there (a class-level skip applies only in TestCase.run; CPython v3.12.3 case.py:159-161 and :612-617, loader.py:94). Per-class counts are in the receipt's data.missing_ids and the trimmed result; the serial seconds sentence now names the two missing ids in S-r3's durations table. - Ids with a result record versus records (F2): L4C-r3 has 9,743 records for 9,741 ids (three K4 subtest FAIL records). - n=3 limitation (F3), controls labelled as synthetic fixtures executed on the hosted runner and reproduction scope per interpreter and module (F4), usage statement with 11,524 ubuntu-24.04 job-seconds and unknown session model usage (F6), runner image recorded only by the 16 arm and control jobs (F7), and the two macOS phrasings (F8). - retained_files hashes updated for README.md and result-trimmed.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whole suite was not run locally under the candidate runner, so the row now says a local whole-suite run would very likely have shown the crash, inferred from the per-module reproductions. It attributes the TypeError to the two IsolatedAsyncioTestCase modules and the masked PicklingError to the dynamic-class module (first-exception rule, Lib/multiprocessing/pool.py:822 at CPython v3.12.3). The unmeasured "in seconds" is dropped. Row format and position unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the trial environment, not because of the operating system Delta review finding on the outcome record: the skip is skipUnless(HAS_SDK), true only when both alpaca and requests import; the trial arms installed neither (requirements-ci.txt installs neither), so the 66 tests were skipped in every S run in that environment. Reworded the two sentences that tied the skip to ubuntu-24.04. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 571520b084
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Exact-head SOURCE ACCEPT at I checked the actual pinned publication, original hosted compare artifact, frozen preregistration and GitHub run/job records. The decision follows the recorded eligibility rule: nine Linux candidate runs exit 1 and remain ineligible; three serial runs exit 0 and are eligible; Linux is rejected. Twenty macOS jobs have no assigned runner and zero steps, so macOS has no verdict. Candidate timing below the speed thresholds does not qualify an incomplete, crashing suite. The frozen experiment equals the retained byte copy: 43,208 bytes, SHA256 Original job records reconcile all twelve arm-step durations and the 11,524 seconds across eighteen Linux jobs. Serial arm-step durations are 1,688/1,659/1,186 seconds. Step success alone is not an arm exit-code proof: this trial intentionally records failing arm exits inside a successful collection step. The original result retains the actual arm exits. The publication explicitly records cancellation outside the preregistered rule and the changed record location. Those departures stay visible. The proposed cross-platform failure is explicitly an inference; it supplies no macOS run. The source mechanism is supported by unittest-parallel at All 27 actual Git source identities verify. The registry preserves all 9,420 base paths in relative order, adds fourteen, replaces only the owned harness-defaults binding, and retains all 187 receipt records and 26 convergence records unchanged. All fifteen content bindings and twelve retained-file bindings match their actual bytes. Preserve current-main foreign rows when composing later merges. The evidence classes and usage correctly distinguish repository execution, synthetic controls, local reproductions, source inspection, independent platform observations and unknown planning/build/review model usage. These are not upstream acceptance. I did not redownload the twenty full raw logs or rerun the trial. Verification custody: 34 native GH commands exit 0, retained stdout/stderr manifest SHA256 |
…itique of the suite-parallelism trial The trial's convergence record is now evidence/artifacts/suite-parallelism-trial-20261003/experiment.json (status observed, decision reject), generated from the raw run directories, the jobs API and result.json. It has 36 observations: the 12 timed ubuntu-24.04 runs, the 4 ubuntu-24.04 control jobs and the 20 macos-15 jobs that never received a runner (skipped). It carries the preregistered metrics, commands, lane, roles and base revision unchanged; its frozen_inputs cite the retained preregistration copy, whose 15 frozen hashes were re-checked at 1e4bb5a. Token usage is null (unknown), not the preregistered observed zeros, and runner time is in each observation's scope text; both departures are in its limitations and the decision record. The decision record, the receipt and the artifacts README cite it as the convergence record; the receipt keeps the detailed data and the claim text, and no longer says it replaces the observed record. The decision record's new "Completeness critique" section names the missed modality, nine candidate classes with each relayed claim marked VERIFIED or LEAD-ONLY against its pinned upstream source, and seeds for the next ci-supply-chain landscape sweep. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A free-form receipt in place of the preregistered observed convergence record, and a substantive trial closed without its completeness critique. Both rows sit above this pull request's two earlier 2026-10-03 rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The table now has four rows dated 2026-10-03, so the two references to "the anti-pattern row of 2026-10-03" name the row by its title. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…en inputs The observed record's frozen_inputs pointed only at the preregistration copy and listed result-trimmed.json, the oracle's output, as an evaluation input, so scripts/validate_convergence.py checked no byte of the preregistered oracle, controls, B1 fix or lock files (delta review of #652, D6). - frozen/ holds byte copies, stored as .txt, of the 15 files the preregistration froze, written with git show at the trial head 1e4bb5a (draft #646, closed unmerged): 15 of 15 byte-equal to their originals and to their preregistered sha256, 264,054 bytes, all passing scripts/validate.py --scan-file. No lock copy's name starts with requirements, so tests/test_osv_lockfile_coverage.py does not treat it as a lockfile. - experiment.json, regenerated from the raw run data: frozen_inputs cite the copies in the preregistration's role lists and order (3 sources, 3 inputs, 9 evaluation) plus the preregistration copy at the end of sources; result-trimmed.json stays an observation artifact only. A new limitation lists the other departures from the plan (command 11, one scope text per run, the job and step timestamps kept in the receipt's data.jobs), and each control observation's scope states its controls file's step seconds (0, 0, 1 and 0 s under S, L4, L4F and L4C; verifier note). rollback records #646 as closed. - README: a frozen/ section says what each copy is, that the originals live on the closed trial branch and that the hash equality with the preregistration is the point. - Receipt: data.retained_files lists all 28 files under the artifacts directory with their new hashes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ow each claim was read Delta review of #652 (D1 to D5) and two corrections found while verifying it. - D1: a "Sources missed before the trial" subsection names four sources that the preregistration and its decision record do not record reading (plan W1 and its Codex review, kept outside the repository, were not re-read): unittest-parallel's issue tracker (25 issues; issue 27, filed 2025-08-24, is titled with this trial's exact exception, and issue 22 reports a pickling failure of the family of root cause (b)), CPython's tracker (python/cpython#83438, whose only reply calls the pickling error expected: "contextvars are not compatible with multiprocessing"; no issue on pickling IsolatedAsyncioTestCase was found), this suite's own async and type()-built test modules, and GitHub's changelog. Matrix sharding, which plan W1 compared and the preregistration rejected, is listed with its blockers re-checked at the base (.github/main-ruleset.json:35 and :41, docs/github-automation.md:26-28, tests/test_workflow_hardening.py:1249-1257, GitHub's five-job macOS limit at github/docs 2bd66de8 limits.md:64-69 and the trial's measured macOS queue) and noted as outside the one-job task; a seed and a critic question follow. - D2: stestr's reader hang is VERIFIED at stestr/output.py:160-161 (4.2.1), and concurrencytest's repeated class fixtures under its default round-robin partitioning at concurrencytest.py:158-160 (0.1.11); only the lead's module-fixture wording stays LEAD-ONLY. - D3: the critique, the Evidence class and the SOTA sources say how each class of claim was read: repository files with GET requests at the release tag or commit; repository records, PyPI JSON, the changelog, the GNU announcement and the issue trackers with unpinned GET requests; the GNU xargs statuses from the installed manual page; the unittest.mock count and the matrix-sharding blockers locally. - D4: the second trial's local preflight is attributed to the coordinator and marked not verified here. - D5: the departures of the observed record's shape list all seven (command 11, the control observations, the skipped macOS observations, one scope text per run, the frozen inputs at other paths, the timestamps in the receipt, null token usage), with runner time in each executed observation's scope text. - Corrections: 106 of the 234 test modules import unittest.mock (100 at module level, 6 inside a test), not 100; and #646 was closed unmerged at 11:05:52Z on 2026-10-03, after this record's pull request opened. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ply the last delta-review wording fixes The two requirements-style lock copies (byte copies, same sha256) are renamed trial-ci-suite-<os>-pins.frozen: a .txt name containing requirements is what Python dependency scanners and the OSV inventory test match, so the .txt suffix alone did not keep them out. README and experiment.json now say what each name rule does and that the dependency-graph effect is expected, not verified. Also: the sources-missed list now includes the pytest and coverage reads the preregistration recorded, matrix sharding's attribution names the concurrent Adoption runs, the receipt separates pinned source reads from unpinned GETs, the experiment record hedges the absence claim, and the anti-pattern log gains the frozen-inputs row. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…(hot-file protocol; last commit) One registration on main's manifest at 4ced292: the decision record, the receipt, the artifacts including the 15 frozen byte copies (the two lock copies renamed so no scanner pattern matches them), docs/harness-defaults.md (five rows), and the observed record in convergence_records. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
571520b to
13e084c
Compare
…ries to the bytes after a986c4c (review 652 P1) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry plus the owned rows) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fore the final hot-file commit Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (hot-file protocol: every hot-file edit in the last commit) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude session native-agent-stack-5f: landing at head Observed main Required checks at this head: 7 pass 1 skipping . Unresolved review threads: 0. |
|
Claude session native-agent-stack-5f: post-merge observation. Landed as |
Scope
1e4bb5ab): a validated observed convergence record (decisionreject, 36 observations), a dated decision record with the unit's completeness critique (not adopted; PR-A and PR-A2 are not opened), a durable sanitized receipt, supporting artifacts with the sha256 of the originals that expire on 2026-11-02, byte copies of the 15 files the preregistration froze (the observed record's frozen inputs), and four anti-pattern rows. Docs and evidence only: no workflow, test, ruleset or setting changes.4ced2923063db6a6dcafa9f25af5ee05a4153c75lane:foundationdocs/decisions/2026-10-03-suite-parallelism-trial-outcome.md,evidence/receipts/suite-parallelism-trial-20261003.json,evidence/artifacts/suite-parallelism-trial-20261003/(13 files, among them the observed convergence recordexperiment.json),evidence/artifacts/suite-parallelism-trial-20261003/frozen/(15 byte copies, stored as.txt, except the two lock copies, which are named*-pins.frozenso no scanner pattern matches them),docs/harness-defaults.md(five rows at the top of the anti-pattern table),manifests/evidence.json(registration and theconvergence_recordsentry only, the branch's last commit, per the hot-file protocol indocs/lanes.md)SOTA sources
bda5d77dc1a2fa2df90f5f5a7de297ea375e345c:src/unittest_parallel/main.pylines 150-160 (spawned pool,pool.mapat 160), 164-211 (report only after the map), 240-242 (exit status), 62-63 (--thread);README.mdlines 43-71; issue 27; commits after the release; PyPI JSON (1.8.6 latest on 2026-10-03).f6650f9ad73359051f3e558c2431a109bc016664):Lib/unittest/async_case.pylines 35-38,Lib/unittest/case.pylines 159-161, 176-182 and 612-617 (a class-level skip only marks the class and applies when a test runs),Lib/unittest/loader.pyline 94 (one instance per test name),Lib/multiprocessing/pool.pylines 367, 528-546, 774 and 809-831,Modules/_pickle.clines 3646-3704; at v3.13.16 (cbc944f4bc59639a444dd971c737788ba2283a91):Lib/unittest/async_case.pyline 42.068546469ae7f079368d12f969991a045121c4a0:data/reusables/actions/workflows/triggering-a-workflow-paths5.md(a pull request's paths filter uses the three-dot diff). GitHub REST: workflow jobs, artifacts.cf470ec0bf7eb89cd97dd56df4859eae5db46447; pytest-dev/pytest-xdist v3.8.0 at1e3e4dc16523c8a8f6c67d95a950166420718c99; mtreinish/stestr 4.2.1 at2802f1425f45b9ad0857ff4685ab469dcd2cf2e0(stestr/output.pylines 154-161); testing-cabal/testtools 2.9.1 at088c98e24961ecf6d94ea5204457f2dcffe2f1c6; nose-devs/nose2 0.16.0 atc93ca65ac6f62475209aa8944fda14c671615549; cgoldberg/concurrencytest 0.1.11 at266e27c833f9b1ecb47142bdddbac5199bebab72(concurrencytest.pylines 89-90, 158-160 and 179-187); testing-cabal/subunit 1.4.6 atc85605280cb1d975b3076b9ad7ac38f17e858945; CleanCut/green 4.0.2 at4de285b05b8e6b161159be2241b219dc5176f0e5; zopefoundation/zope.testrunner 8.3 at1061ccc4b6b5a824870f2142bc948863b7ba1f70; python/cpython v3.13.16 atcbc944f4bc59639a444dd971c737788ba2283a91(Lib/unittest/loader.py,Lib/test/libregrtest/); craigahobbs/unittest-parallel atbda5d77d(main.py:123-128,:142-160,README.md:59-70); github/docs at2bd66de8cea336061c9ea060c9b37385136e6ab3(workflow-syntax.md:background,wait,wait-all,cancel,parallel;limits.mdlines 64-69, concurrent macOS jobs per plan). With GET requests that no commit pins: the GitHub changelog of 2026-06-25; the GNU findutils 4.10.0 announcement; PyPI JSON per release and the repository records; unittest-parallel's issue tracker (issue 22, issue 27 and a search for "pickle") and CPython's (python/cpython#83438, bpo-39257, and searches forIsolatedAsyncioTestCasewithpickle,picklingormultiprocessing). Locally: the installed GNU xargs 4.10.0 manual page (the gnu.org manual did not answer).docs/acceptance-evidence-policy.md:55-63(preserve failed attempts, declare every sanitization),docs/evidence.md:18(compatibility_attempt),docs/lanes.md:94-141(hot-file protocol),blueprints/convergence-practice/contract-reference.md:27-28(distinct source, input and evaluation artifacts) and:84-92andscripts/validate_convergence.py:99-123and:136-146(usage fields, failure coverage, run status, scope of qualifying runs),tools/sota-convergence/landscape-sweep/build_inputs.py:21-22and:225-234(the seeds format); for matrix sharding,.github/main-ruleset.json:35and:41,docs/github-automation.md:26-28andtests/test_workflow_hardening.py:1249-1257. At the trial head1e4bb5ab: the 15 frozen files (byte-copied underfrozen/) and the preregistration's decision record, "Alternatives".Evidence-class table
reject, no qualification, token usage null (unknown),native_retries0 for the 16 executed jobs, runner time in each executed observation's scope text (11,461 s over the 16 observed jobs);frozen_inputscite byte copies of the 15 files the preregistration froze (3 sources, 3 inputs, 9 evaluation, in its order, each with its preregistered sha256) and the preregistration copy at the end ofsources; every departure from the plan stated inlimitationsnative_cli_executionobservations generated from the raw run directories, the jobs API (GET, steps included) andresult.json;scripts/validate_convergence.pychecks declared consistency and artifact hashes, not truthevidence/artifacts/suite-parallelism-trial-20261003/experiment.jsonfrozen/equal their originals at1e4bb5abbyte for byte and carry the preregistered sha256 values (15 of 15); all passscripts/validate.py --scan-file; 264,054 bytesgit showat1e4bb5abagainst each copy, and sha256 against the preregistration'sfrozen_inputs)evidence/artifacts/suite-parallelism-trial-20261003/frozen/; receiptdata.retained_filestype()-built test modules; GitHub's changelog), nine candidate classes the trial did not evaluate plus matrix sharding with its blockers, each relayed claim marked VERIFIED or LEAD-ONLY, and seeds for the nextci-supply-chainsweepunittest.mockcount (106 of 234 test modules, 100 at module level), the suite's modules and the matrix-sharding blockers locally withgit grepand reads at1e4bb5aband the base. The candidate list is a lead from a read-only single-family Codex laneTypeError: cannot pickle '_contextvars.Context' objectfrompool.map(main.py:160); 199 ids never ran at module level and 82 at class level (in S, 90 of the 199 and 8 of the 82 were ok and the rest skipped, all 66IsolatedAsyncioTestCasetests among the skipped); oracle: ubuntu-24.04 reject, macos-15 no verdict; id mapping oknative_proven(hosted run 37109532421, attempt 1; receipt retained with hashes and sanitized excerpts)evidence/receipts/suite-parallelism-trial-20261003.json;evidence/artifacts/suite-parallelism-trial-20261003/Ran 8andFAILED (failures=2, errors=2, skipped=1, expected failures=1, unexpected successes=1), exit 1 under S and 5 under unittest-parallel; the crash control exited 3 under S and 124 (the 300 s bound) under every parallel armsynthetic: synthetic fixtures executed on the hosted runner (run 37109532421)data.controls;control-logs.txttest_cli_refuses_live_base_url_from_env_file_without_networkERROR in 3 of 3 L4F runsnative_proven(observed); causes not established, the L4F one unknowndata.non_missing_mismatchesgh api --method GET); nearest template classsource_reviewdata.jobs,data.cancellationIsolatedAsyncioTestCasestores acontextvars.Contextin__init__, so its suites cannot be pickled even when every test in them is skipped, as all 66 were on ubuntu-24.04; fourtype()-built classes intests/test_adoption_version_probes.pyare not importable under their__qualname__local_integration(toy modules on CPython 3.12.3 and 3.13.16; the three real modules on 3.12.3 only; the coordinator also saw the version-probesPicklingErroron 3.13.16, an output not in the retained artifacts) plussource_review(async_case.py,case.py,loader.py,pool.py,_pickle.cat pinned tags)local-reproductions.txt; receiptdata.root_causesAGENTS.md, PyPI latest 1.8.6source_review(GET reads on 2026-10-03)data.upstreamdata.arms_ubuntu_24_04data.usage,data.jobslocal_integrationLocal commands run
Decision record
docs/decisions/2026-10-03-suite-parallelism-trial-outcome.md: the decision (not adopted), the observed convergence record it cites, what was tried with the preregistered rule and a pointer to #646 at1e4bb5ab, what happened, the cancellation and why, root causes, what the attempt does not show, departures from the preregistration (among them where the observed record lives and the seven departures of its shape), the completeness critique (missed modality, missed sources, candidate classes including matrix sharding) with its seeds, overturn conditions and next-trial preconditions.Host evidence
Not applicable: no files under
evidence/hosts/change.Checklist
permissions: contents: read(no workflow changes in this PR).Notes for the merger:
manifests/evidence.jsonis the shared hot file; ifmainmoves, takemain's copy, register againdocs/decisions/2026-10-03-suite-parallelism-trial-outcome.md,docs/harness-defaults.md,evidence/receipts/suite-parallelism-trial-20261003.jsonand the 28 files underevidence/artifacts/suite-parallelism-trial-20261003/(15 of them underfrozen/), appendevidence/artifacts/suite-parallelism-trial-20261003/experiment.jsontoconvergence_recordsagain, then run the--writecommands,scripts/validate.pyandscripts/validate_convergence.py --all-recorded --root . --json(docs/lanes.md). #632 inserts five rows at the same top of the anti-pattern table, so whichever of the two merges second has a trivial table conflict (keep both sets of 2026-10-03 rows). This branch was rebuilt onmainat4ced2923(after #634 and #636), with this PR's five rows above the 2026-10-02 rows; whenmainmoves again,git merge-treeis likely to report conflicts in those same two files (keep both sets of rows; takemain'smanifests/evidence.jsonand register again as above). #646 was closed unmerged at 11:05:52Z on 2026-10-03, ten seconds after this PR opened; do not reopen it or open another pull request that adds its workflow, because either re-runs the whole trial. The record names no private repository and no dollar figure.🤖 Generated with Claude Code