Repository navigation
Pin socraticode 1.15.0 and ccusage 20.0.26 with upstream-suite qualification receipts - #446
Conversation
…ws, receipt wording
R1: the Codex user template ran tools/socraticode-1.15.0 on every platform,
while the macOS pin and bootstrap install 1.14.0. render_config.py now derives
${SOCRATICODE_VERSION} from the selected platform's pin, as it derives
${AI_MEMORY_BIN}; --set SOCRATICODE_VERSION names another install. New
SocratiCodeVersionTests render both platforms (Linux 1.15.0, macOS 1.14.0) and
fail against the old literal template.
R1b: the examples and recipes name the per-platform version.
R2: docs/token-efficiency-stack.json installs socraticode 1.15.0 (with
--before and the macOS 1.14.0 exception) and ccusage 20.0.26.
R3: the socraticode receipt gives the stale-graph cause its own times show,
records the post-cutover scratch Codex home that still names 1.14.0, and
relabels the notifications/message zero-matches as untested.
R4: the ccusage receipt drops the token-guide deferral, discloses the daily
comparison's unretained argument vector, and names the dated blueprint parser.
R5: docs/stack.md keeps its September 19 snapshot value.
R6, R7: provenance sentences in pins-macos-arm64.json and foundation-stack.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
3f93d0e to
19e5f29
Compare
|
Other-lane acknowledgement requested (lane:shared), head 19e5f29.
🤖 Generated with Claude Code |
|
Merge-time requirement: evidence manifest size bound. This PR adds 14,248 bytes to The fix is in 499a661 on
|
… receipts SocratiCode 1.15.0 moves the Linux pin, launcher templates and examples after the user waived the 7-day cooldown. The unchanged upstream suite passed at f6191f07, pre-cutover native checks ran against production services, and the workstation cut over at 2026-09-27T21:42:49Z. The macOS pin stays 1.14.0 (MAC_PIN_LAGS_LINUX). The receipt records the partial all-sessions state: two 1.14.0 MCP children remained at 02:03:14Z, and the stored graph was still built by v1.15.0 at about 02:07Z. ccusage 20.0.26 fires the 2026-09-26 publication trigger. At d9821088 the upstream Rust suite passed (944/0/3), the Node tests passed 34/34, the native_token_ci fixture is byte-identical across versions, and daily, weekly and monthly native token totals are unchanged. 20.0.26 prices claude-opus-5-5 and gpt-6-luna, which 20.0.24 left unpriced. The workstation launcher switched at 2026-09-27T22:01:01Z. Linux and macOS pins move together; macOS rests on registry evidence only. Also updates the saturation-audit rows, the upstream-snapshot and landscape release identities, the recipes row, the decision records and the generated new-host grand list. Landscape winner pins stay unchanged, so host receipts at the new versions need --allow-unbound-version. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ws, receipt wording
R1: the Codex user template ran tools/socraticode-1.15.0 on every platform,
while the macOS pin and bootstrap install 1.14.0. render_config.py now derives
${SOCRATICODE_VERSION} from the selected platform's pin, as it derives
${AI_MEMORY_BIN}; --set SOCRATICODE_VERSION names another install. New
SocratiCodeVersionTests render both platforms (Linux 1.15.0, macOS 1.14.0) and
fail against the old literal template.
R1b: the examples and recipes name the per-platform version.
R2: docs/token-efficiency-stack.json installs socraticode 1.15.0 (with
--before and the macOS 1.14.0 exception) and ccusage 20.0.26.
R3: the socraticode receipt gives the stale-graph cause its own times show,
records the post-cutover scratch Codex home that still names 1.14.0, and
relabels the notifications/message zero-matches as untested.
R4: the ccusage receipt drops the token-guide deferral, discloses the daily
comparison's unretained argument vector, and names the dated blueprint parser.
R5: docs/stack.md keeps its September 19 snapshot value.
R6, R7: provenance sentences in pins-macos-arm64.json and foundation-stack.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Carry only the observability/grand-dashboard/progress.py and tests/test_grand_dashboard.py hunks of 499a661 on origin/claude/prompt-audit-instructions-20260927 (two-file `git patch-id --stable` 8f693abd4c60119019b7a419b4cca4344ac4deed). That branch's decision and control files are not taken. With this PR's two receipts registered, manifests/evidence.json passes the 2,000,000-byte MAX_BYTES bound that progress.py applies to every source. The manifest gets its own SOURCE_MAX_BYTES bound of 8,000,000; every other source keeps MAX_BYTES. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… hashes Rebased onto main dd99677. On the first rebase, onto 9f8db58, the manifests/evidence.json conflict was resolved by taking main's version; main's next commit (#459) then applied without conflict. This commit re-registers the branch on top of main's manifest: - receipts[]: the two entries from this PR's original registration commit 19e5f29, appended verbatim after main's last entry; each equals its receipt file's id, kind, component_ids, claim and limitations. - files[]: scripts/host_receipts.py register_file for all 31 paths in `git diff --name-only origin/main...HEAD`: the 29 PR files (the two receipt files inserted) and the two grand-dashboard files last. Re-running register_file for those paths on dd99677 leaves the manifest byte-identical. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
19e5f29 to
184d1f2
Compare
|
Lane acknowledgement (docs/lanes.md): on 2026-09-28 at about 13:40Z the repository owner, who owns both lanes, answered "Yes, merge now" to merging this PR without a live trading-lane session. This comment records that as the acknowledgement. Rebased head 184d1f2 (base main dd99677):
Merge follows green CI on this head. |
…by subject After the rebase onto 3058b23, the dashboard commit keeps only its record subsection and the two receipts: #446 landed the same progress.py and test hunks (two-file patch-id 8f693abd4c601190). The review README now names the reviewed commits by subject, since each rebase changes their SHAs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…by subject After the rebase onto 3058b23, the dashboard commit keeps only its record subsection and the two receipts: #446 landed the same progress.py and test hunks (two-file patch-id 8f693abd4c601190). The review README now names the reviewed commits by subject, since each rebase changes their SHAs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…CLAUDE.md, 4 of 4 blind judgments; dashboard manifest bound (#444) * AGENTS.md: scope role dispatch to new or ad-hoc stages; current-state holdout wording Two items of the 2026-09-27 prompt audit, resolved by cross-family convergence (docs/decisions/2026-09-27-prompt-audit-resolution.md, lane:foundation PR): - Role dispatch applies to each new or ad-hoc workflow agent() stage; the saved scripts vendored in examples/claude-native/workflows keep their reviewed routing, byte-identical to agent-lab, as that README states in its opening paragraph and the Workflow contract's Dispatch by role bullet (F1). - The inspected 2021 control segment: "has been inspected, so no experiment may present it as a fresh untouched holdout" (F4). o200k (gpt-tokenizer 3.4.0): AGENTS.md 2,374 -> 2,401 tokens (+27). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * X9 second round: blind two-family re-judgment splits again; CLAUDE.md unchanged After round 1 split 2 and 2 on X9 (CLAUDE.md:3-4), a second round added: - the installed client's own refusals of /context and /mcp through the Skill tool; - an executed comparison of the three texts, all headless claude -p runs through promptfoo 0.123.1: K1-K3 as 27 preregistered runs on Opus 5.5, and K4 as 9 runs on Sonnet 5; - 38 dated sources, plus 5 recorded searches that found nothing. Attempt 1 was void, because its judges could read this record and the coordinator's work directory. Attempt 2 was final. Its judges read a plain export of ba1700a and judged in both orders: blind-adjudicator agents and the packaged GPT-6 runner. A void audit, hashed before any return was read, found no voiding hit. Both GPT-6 judgments chose the GPT-6 lane's text, and both Claude judgments chose the Claude lane's text, so under the unanimity rule nothing is applied. All eight judgments across both rounds reject the current line. The addendum records both positions and the comparison that would decide it. Review: - One round, by GPT-6 and Claude reviewers, then one repair round: - disclose that K4 ran on Sonnet 5 while the packets said Opus 5.5; - make check_orders.py exit 1 on a difference, and add retally.py; - correct the counts; - add receipts (run windows, settings file names per run, the K4 skill-list sentences, the isolation receipt, a source supplement); - fix the wording. - A separate re-check found all 13 items fixed and two new defects: an overstated "keeps only" claim and an overstated account of K4's origin. Both are fixed. The recovered first check_orders.py is kept to show its flaw. - A pre-handoff review found that the published packets, prompts and K4 files quoted answer heads containing client configuration (an enabled-plugin entry, status-line wiring, plugin install records). Every published copy now withholds them. attempt2/sent-sha256.json keeps the hashes of the files as sent. This commit squashes the earlier ones, so no commit on the branch carries that text. docs/harness-defaults.md gains three dated anti-pattern rows: - a model scope taken from the preregistration instead of each run's recorded model; - a check that reported a mismatch and exited 0; - judge inputs published with answer text that quotes client configuration. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prompt audit round 3: X9 applied to CLAUDE.md; lane A's S1 and X5b in AGENTS.md; AN-13 receipt - CLAUDE.md lines 3-8 (X9): round 3 compared the two texts on the measure the second round named (native /mcp and /context refusal in print mode, with a promptfoo-graded, preregistered comparison) and the blind adjudication chose the GPT-6 lane's text in all four judgments. - AGENTS.md:3 (X5b): the operator's 2026-09-28 top-rule paragraph, byte for byte; AGENTS.md:37 (S1): build a PR description from the pull request template. Both converged in lane A (agree / agree). X5a and X5c land in #458. - AN-13: the headless /doctor prompt-audit run the 2026-09-28 community sweep left open. Facts of both runs in evidence/artifacts/prompt-audit-20260927/an13/ (the report text is not published); no edit made from it; each finding's disposition is in its README and in the decision record. - Decision record and x9-round3/ evidence (preregistration, isolation receipt, fixtures, judges, usage). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prompt audit: print-mode anti-pattern row (AN-13), line-7 co-change with X5b, round-3 token receipts - docs/harness-defaults.md: a dated row for running a headless skill that starts background agents under print mode's default 600 s wait (AN-13's first run: four Explore agents stopped by the client, an interim note instead of the report); the rerun with CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 returned the report (code.claude.com/docs/en/env-vars). - docs/harness-defaults.md:7 restates AGENTS.md's new top-rule heading (X5b), disclosed in the decision record. - controls/token-counts-r3-{f508ffb,base,after}.txt: the o200k counts the record cites for CLAUDE.md (30 -> 96 at f508ffb; 96 -> 162 at fb14ded) and AGENTS.md (2,374 -> 2,484 at fb14ded). - The record cites validate.yml's heading rule at both revisions and notes the Codex-lane host step done on this workstation after #458. - x9-round3/README.md: the scan note names every kind of match (names, not values), including one judgment's quote of the untested cases. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Grand dashboard: the evidence manifest gets its own read bound This pull request's registrations take manifests/evidence.json past the 2,000,000-byte bound progress.py applied to every source (1,987,653 bytes on main at fb14ded; 2,017,827 with this branch), so the snapshot refused it and eight dashboard tests failed. The snapshot reads only the manifest's receipt count: SOURCE_MAX_BYTES gives the manifest 8,000,000 bytes and every other source keeps MAX_BYTES. read(root, relative) keeps its signature (the tests patch it). The new test fails on the unchanged reader (controls/dashboard-bound-before.txt) and passes with 16 others (dashboard-bound-after.txt). Overturn: when the manifest nears the new bound, count receipts from a smaller source instead of raising it again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prompt audit round 3: review repairs, the audit's control, the Codex lane host receipt One review round on 0779e91a (GPT-6 codex exec and a Claude evidence-reviewer, both changes-needed); every finding was in the record or the evidence text: - the record: --max-turns 10 with one 11-turn result; every builder change after the freeze; the applied X9 text ties with the line it replaced on the preregistered metrics; the runs used ba1700a's files; X5b's wording and the tests that read AGENTS.md; the advisor calls; M4 is project-side; a review subsection; - x9-round3: audit_control_r3.py and judges/audit-control.json (clean copies reproduce the published verdicts; one planted access per judgment voids all four); the README lists every builder change, the judge-action flag and the tested configuration; - codex-worker-lane-host-20260928: the 12:14-12:15Z dry run, apply and proof, sanitized as #406's were, and a recheck with exit codes and times (dry run: block in place; proof 7 of 7); - review-444: both prompts, returns, usage and each finding's disposition. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prompt audit: map the reviewed commits to the rebased ones The branch was rebased onto 9f8db58 (#400) with every patch unchanged (git range-diff shows each pair as =). review-444/README.md says 0779e91a, the reviewed head, was never pushed and was 7e5eddcd plus a registration commit, and gives the rebased SHAs of 9e036e7d and 7e5eddcd; the decision record points there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Prompt audit: #446 carried the dashboard bound; map reviewed commits by subject After the rebase onto 3058b23, the dashboard commit keeps only its record subsection and the two receipts: #446 landed the same progress.py and test hunks (two-file patch-id 8f693abd4c601190). The review README now names the reviewed commits by subject, since each rebase changes their SHAs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * AN-13: point M4 and M5 to their 2026-09-28 second-family lanes M4: both lanes reject the rewrite (no edit; the preload stays with the 2026-10-25 comparison). M5: the lanes amended with different texts and the blind adjudication split (unchanged). Both are recorded in docs/decisions/2026-09-28-an13-m4-m5.md, a separate lane:foundation PR. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * AN-13 M4 pointer: the trial's removal rule is open (#460) #460's record now says the M4 lanes did not test the skills trial's rule that removes a trial skill whose instructions conflict with CLAUDE.md or AGENTS.md, and leaves that question open. Say so where this record points to it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Register the instruction files, the decision record, the anti-pattern rows and the prompt-audit evidence Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…16.3.8 pin The snapshot's Next.js entry is refreshed with its own method (gh api on the repository, the latest release and the pin and tag commits; one call each, no retry), as #446 did for ccusage; the latest stable is v16.3.8, so the selected version matches it. The lifecycle audit row carries the stack version, which its test compares. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…anifest rows (hot-file protocol, last commit) manifests/stack.json: nextjs 16.3.6 -> 16.3.8 at tag commit b0fad0d4, upstream sources, the freshness note, and the new receipt id in evidence_ids beside the two 2026-09-20 receipts, which qualify 16.3.5 only. manifests/evidence.json: the receipt nextjs-1638-qualification-20261001 in receipts[] (kind native_cli_e2e, as the #446 qualification receipts are); three new file rows (the two retained 16.3.6 files and the receipt) and the re-registered recipe, history, snapshot, audit and stack files. Files 8,906 -> 8,909; receipts 183 -> 184; convergence records unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alification receipt and dated record (#587) * U2: Next.js 16.3.6 -> 16.3.8 in the application-delivery recipe, with its qualification receipt and the dated U2 record The v16.3.8 release notes list seven security advisories, the most severe GHSA-cjq9-62q9-8jv4 (high, SSRF in Image Optimization). The qualified two-file patch changes next, @next/env and the eight @next/swc-* lock entries. pnpm's release-age rule refused the first frozen install at 12:13Z; no exclusion was added, and after the window the unchanged candidate passed the frozen install, peer check, typecheck and production build (exit 0 each). The 16.3.6 files are retained under history/, the receipt keeps the failed attempt, and make verify is not claimed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * U2: the upstream snapshot and the lifecycle audit follow the Next.js 16.3.8 pin The snapshot's Next.js entry is refreshed with its own method (gh api on the repository, the latest release and the pin and tag commits; one call each, no retry), as #446 did for ccusage; the latest stable is v16.3.8, so the selected version matches it. The lifecycle audit row carries the stack version, which its test compares. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * U2: stack pin Next.js 16.3.8, its receipt indexed, and the evidence manifest rows (hot-file protocol, last commit) manifests/stack.json: nextjs 16.3.6 -> 16.3.8 at tag commit b0fad0d4, upstream sources, the freshness note, and the new receipt id in evidence_ids beside the two 2026-09-20 receipts, which qualify 16.3.5 only. manifests/evidence.json: the receipt nextjs-1638-qualification-20261001 in receipts[] (kind native_cli_e2e, as the #446 qualification receipts are); three new file rows (the two retained 16.3.6 files and the receipt) and the re-registered recipe, history, snapshot, audit and stack files. Files 8,906 -> 8,909; receipts 183 -> 184; convergence records unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Scope
ed3cd96c(origin/main)lane:shared. It changes whatmanifests/stack.jsonsays (versions), so it needs the trading lane's acknowledgement before merge (docs/lanes.md:145-149).blueprints/token-native-focus/saturation-audit.json,catalogs/landscape/upstream-snapshot.json, docs (stack, foundation-stack, memory-RAG, the two dated decisions), examples,recipes/README.md,scripts/native_token_ci.py,tests/test_adoption_bootstrap_macos.py, two new receipts, and the hot filesmanifests/stack.json,manifests/landscape.jsonandmanifests/evidence.json(the last commit).SOTA sources
f6191f076a42405f0d5508139f3a8b505cfef93a. Upstream'snpm testunit suite and its Qdrant regression tests at that commit.d9821088b98aa536c7a385aa1a4579d6fa02269b. Upstream's Node tests andcargo test --workspaceat that commit; npm tarball integrity checked against the registry (sha256 prefixb8d59c19).Evidence-class table
evidence/receipts/socraticode-1150-qualification-20260927.json: 2475/2475evidence/receipts/ccusage-20026-qualification-20260927.json: 944 passed, 0 failed, 3 ignoredLocal commands run
Decision record
docs/decisions/2026-09-26-token-practice-f1-f9.md(dated update: the currency trigger fired, and the next trigger is anything newer than 20.0.26) anddocs/decisions/2026-09-25-workstation-sota-refresh.md. The user waived the socraticode 7-day cooldown on 2026-09-27.Residuals
docs/token-efficiency-stack.json(owned by the cards session) still says socraticode 1.14.0 and ccusage 20.0.24.The landscape verdict pins stay at 1.14.0 and 20.0.24 until a verdict re-run.
Host receipts (
host_receipts.py record --allow-unbound-version) are not recorded.native_token_ci.py --installis left to PR CI.ccusage SLSA attestation not fetched, and the installed platform binary's integrity not compared with the registry: both recorded as limitations.
On this host,
native_token_ci.py(installed executables, no--install) exits 1 at 19e5f29. The only failure isserena-tools-list-byte-parity-with-the-pinned-install, and the same check also fails at origin/main ed3cd96. That main run has a second failure, "Installed ccusage differs from its manifest pin", which this branch fixes. The serena failure is a host-specific residual on main; this branch does not touch serena.Review round 1 and repair (2026-09-28)
socraticode-1.15.0, but the macOS pin is 1.14.0.render_config.pynow derivesSOCRATICODE_VERSIONfrom the selected platform's pins file, the way it already derivesAI_MEMORY_BIN.--platform linux-x86_64renderssocraticode-1.15.0and--platform macos-arm64renderssocraticode-1.14.0, each matching its pins file.docs/token-efficiency-stack.json;docs/stack.mdrestored;pins-macos-arm64.jsonandfoundation-stack.md.🤖 Generated with Claude Code