Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
243 changes: 243 additions & 0 deletions docs/decisions/2026-10-03-macos-full-suite-coverage.md

Large diffs are not rendered by default.

5 changes: 5 additions & 0 deletions docs/harness-defaults.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,11 @@ When a claim or action proves wrong, correct it where it was relayed and record

| Date | Anti-pattern | What happened | Rule or check that prevents it | Where enforced |
| --- | --- | --- | --- | --- |
| 2026-10-03 | Passing `-f` or `-F` to `gh api` for a read | Adding a parameter turned the read into a POST: `gh api --help` at the installed gh 2.102.0 (lines 20-27) says the method is `GET` normally and `POST` if any parameters were added. A read-only audit's listing sent `POST user/repos` and got HTTP 422; nothing was created. | Put read parameters in the URL query string, or pass `--method GET` with them. | This log; verification path: `gh api --help` at the installed version, lines 20-27 |
| 2026-10-03 | Projecting allowance exhaustion from a partial snapshot | A "gone in 8-9 days" projection for the monthly Actions allowance that private repositories draw on rested on two days of a bursty series (daily job minutes ranged from 0 to 763). | Project from the full daily series, per repository and runner OS, and give a range of scenarios rather than one date. | This log; verification path: daily job minutes recomputed from the jobs API (`gh api --method GET`) per repository and runner OS |
| 2026-10-03 | Reporting non-success runs as failures | A summary said about half of pull-request runs failed. Grouped by conclusion over 2026-09-25 to 2026-10-02, 6.5% and 7.3% of the pull-request runs of the two workflows that hold the required test jobs failed and 25-33% were cancelled (36-46% on 2026-10-01 and 10-02); 672 of all 712 cancelled runs in the window (701 of them pull-request runs of these two workflows) were superseded by a newer run. The day-wide list for 2026-10-01 had also reached the 1,000-result cap: the complete counts were 176 and 134 pull-request runs, not 160 and 120. | Report success, failure and cancelled separately (`gh run list --json conclusion`, grouped), and partition every query so that no slice's `total_count` reaches 1,000. | This log; [measurement receipt](../evidence/receipts/github-ci-measurements-20261003.json) |
| 2026-10-03 | Researching through raw web search before using the repository's lanes | The first external lookups of a GitHub CI review ran on WebSearch before the user directed research through the Codex runtime lanes, whose dispatch contract is the [routing record](decisions/2026-09-30-sol-primary-quality-defaults.md). | Route external research through those lanes, and relay each claim with the upstream citation that settled it. | This log; [routing record](decisions/2026-09-30-sol-primary-quality-defaults.md) |
| 2026-10-03 | Repeating a known guard false positive in a read-only lane | An Explore agent's command with inline Python `set(...)` and the word "environment" was blocked as `environment_dump` by the PreToolUse secret guard, the case of the 2026-09-25 row; that row's remedy, a script file, is unavailable to an agent that may not write files. | In a read-only lane, aggregate with `jq` or `gh --jq` instead of inline Python; send the reproduction to the guard owner for a failing-first case in `tests/test_secret_path_guard.py`; never bypass or edit the guard | `scripts/hooks/secret_path_guard.py` (`is_environment_dump`); this log until the guard owner's failing-first case lands |
| 2026-10-03 | Pointing an observed record's frozen inputs at a copy of the plan instead of at the files the plan froze | The first observed convergence record of the suite-parallelism trial listed the preregistration copy as its source, input and evaluation artifact, and the oracle's trimmed output as an evaluation input, so `scripts/validate_convergence.py` re-checked none of the preregistered bytes; the delta review of #652 found it | Keep byte copies, with the same sha256, of every file the preregistration froze and list them under the preregistration's own role lists; never list an oracle output as a frozen input; name a lock copy so that no dependency scanner pattern matches it | this log; the trial outcome artifacts under `evidence/artifacts/suite-parallelism-trial-20261003/frozen/` and their README |
| 2026-10-03 | Putting a free-form receipt in place of a preregistered observed convergence record | The suite-parallelism trial's outcome pull request (#652) first kept the failed trial's per-run data only in a receipt and stored the preregistered record as `.json.txt`, so `scripts/validate_convergence.py --all-recorded` checked nothing about the trial: no artifact hashes, run and failure coverage or usage fields. The automated review of #652 found it | Write the observed record that the preregistration names, in a path that does not re-trigger the trial, validate it with `scripts/validate_convergence.py` and list it in `convergence_records`; a receipt adds detail but does not replace it | AGENTS.md:18; this log; [trial outcome record](decisions/2026-10-03-suite-parallelism-trial-outcome.md) and its observed record |
| 2026-10-03 | Closing a substantive trial without its completeness critique | The same outcome record listed the runner classes the trial did not evaluate "for a pending survey" and recorded no critique of the missed modality, sources or candidate classes. The review of #652 found it | End every substantive research or adoption unit with the completeness critic, mark each relayed candidate claim as verified against its upstream source or as a lead, and record the seeds for the layer's next landscape sweep in the unit's record | AGENTS.md:18; this log; the trial outcome record's "Completeness critique" |
Expand Down
Loading
Loading