Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
59 commits
Select commit Hold shift + click to select a range
fe17df9
test(noema): add observed defect false-negative corpus
seonghobae Sep 1, 2026
95b66f9
ci(temp): apply and verify PR1641 review-corpus repair
seonghobae Sep 1, 2026
e27d5c8
fix(ci): repair PR1641 temporary writer workflow
seonghobae Sep 1, 2026
eae62d1
ci(temp): trigger repaired PR1641 writer
seonghobae Sep 1, 2026
f702789
ci(temp): gate PR1641 writer while fixing review findings
seonghobae Sep 1, 2026
fe9b6da
test(noema): add generic class-evidence relabel regression
seonghobae Sep 1, 2026
468c450
fix(noema): stage concrete class-observation repair
seonghobae Sep 1, 2026
6e766c0
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
0dcfa8f
test(noema): reject vacuous class evidence
seonghobae Sep 1, 2026
d8eb254
test(noema): preserve workflow-trigger review regression
seonghobae Sep 1, 2026
10f6d11
fix(noema): bind class evidence to exact source
seonghobae Sep 1, 2026
d0c8b1f
fix(noema): harden nonvacuous review evidence
seonghobae Sep 1, 2026
8c1c82a
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
e1a4406
test(noema): avoid accidental source-token match
seonghobae Sep 1, 2026
b501d26
test(noema): keep duplicate observation regression precise
seonghobae Sep 1, 2026
b046568
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
579d5a3
fix(noema): add exact-source follow-up repair
seonghobae Sep 1, 2026
397d9d2
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
cc3c980
fix(noema): replace lexical causal admission with structural roles
seonghobae Sep 1, 2026
33b6966
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
f3ae6d2
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
8da5c3f
fix(noema): repair exact-source evidence fixtures
seonghobae Sep 2, 2026
1c3346e
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
33ec0a7
test(noema): isolate repair deadline from external DNS
seonghobae Sep 2, 2026
f33a22e
ci(temp): stage repaired PR1641 retrigger
seonghobae Sep 2, 2026
39ae210
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
e8e08cc
test(noema): cover observed-evidence validator branches
seonghobae Sep 2, 2026
02724d3
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
6a4ff1d
test(noema): cover final class-evidence failure branches
seonghobae Sep 2, 2026
f5cf383
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
0918cb3
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
249f28c
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
7c7e2fc
fix(noema): enforce structural observed-defect evidence
Sep 2, 2026
b196d3f
docs(noema): record successor-check proof
seonghobae Sep 2, 2026
7b145bd
ci(temp): repair Noema source-evidence edge cases
seonghobae Sep 2, 2026
0443565
ci: move PR 1641 edge repair off saturated runner pool
seonghobae Sep 2, 2026
812853e
ci: add robust PR 1641 edge repair
seonghobae Sep 2, 2026
f445324
ci: add fail-closed PR 1641 exact-source edge repair
seonghobae Sep 2, 2026
168f053
ci: add executable PR 1641 exact-source edge repair
seonghobae Sep 2, 2026
cdcd220
ci: make PR 1641 edge repair executable
seonghobae Sep 2, 2026
361d9fb
ci(noema): execute and self-clean exact-source edge repair
seonghobae Sep 2, 2026
2136826
chore(noema): remove superseded PR1641 repair workflow
seonghobae Sep 2, 2026
a78131a
chore(noema): remove superseded PR1641 edge workflow
seonghobae Sep 2, 2026
ec53e9c
chore(noema): remove superseded PR1641 repair helper
seonghobae Sep 2, 2026
70bc213
ci(noema): retrigger exact-source edge repair
seonghobae Sep 2, 2026
468d0ed
ci(noema): supersede flawed PR1641 source writer
seonghobae Sep 2, 2026
e4eeaa5
ci(noema): replace indentation-fragile PR1641 writer
seonghobae Sep 2, 2026
2224c45
ci(noema): run robust PR1641 exact-source GREEN repair
seonghobae Sep 2, 2026
d73afbd
ci(noema): keep broad source-evidence regression suite causal
seonghobae Sep 2, 2026
ba2ee0a
fix(ci): make PR1641 GREEN writer executable
seonghobae Sep 2, 2026
c759f23
fix(noema): bind claimed verification to trusted receipts
seonghobae Sep 5, 2026
fbe2602
fix(noema): unify exact diff evidence parsing
seonghobae Sep 5, 2026
2946017
fix(noema): close whitespace evidence provenance gaps
seonghobae Sep 5, 2026
43a1fdc
fix(noema): align structured probe contracts
seonghobae Sep 5, 2026
9df1ea4
test(noema): cover malformed provenance containers
seonghobae Sep 5, 2026
409638a
Merge protected main into Noema evidence root
seonghobae Sep 5, 2026
80fc255
fix(noema): align decision status schema and close HTTP errors
seonghobae Sep 6, 2026
ad48dd6
fix(noema): 응답 정리 오류와 원래 통신 실패를 분리
seonghobae Sep 6, 2026
b8c986e
test(noema): reproduce unreceipted Concept35 claim
seonghobae Sep 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,9 @@
- Raised `hourly-review-repair.yml`'s discovery ceiling from 50 to 200 while rotating deterministic 50-PR deep-inspection windows by hourly run number. The scheduler hydrates only the selected window and stops immediately after its single dispatch, preserving access to newer PRs without quadrupling expensive review/check/comment work. See `docs/doctoring/hourly-review-repair-single-file-consolidation.md`'s 2026-09-03 follow-up.

## [Unreleased]
- Failed Noema reviews retain the original network failure even when response
cleanup also fails, while process cancellation still stops the review. Direct
redirect-rejection tests now release their responses explicitly.
- Include merge-scheduler entrypoint, core, and regression-test changes in
the existing runtime-quality workflow's trigger and suite selector. Scheduler
workflow edits retain queue checks and also select the full review-repair
Expand Down Expand Up @@ -151,6 +154,9 @@ this file. The format follows Keep a Changelog, and versioned releases follow
Semantic Versioning where the repository publishes a release.

## [Unreleased]
- **Keep Noema's strict output schema and deterministic probe validator identical (#1641).** Each structured probe now declares its closed `probe_kind` together with the exact required `class_evidence` witness roles and source receipt fields. Nested `anyOf` variants preserve strict OpenAI-compatible required/additional-property semantics, so a realistic verdict cannot be rejected merely because the outbound schema and local admission contract disagree. The single-request invalid-location regression now reaches and asserts the intended changed-side rejection instead of passing on an earlier status mismatch.
- **Require source-bound observed defect classes in Noema formal reviews (#1641).** Canonical changed-line coordinates now reject JSON booleans, material reviews must cover distinct classes from the executable external-finding corpus, and class witnesses bind to exact changed-side source text (including lexical-shape-independent blank/non-ASCII lines) with non-vacuous causal observations. A single parser now owns both source text and coordinates; bounded truncation drops the incomplete line instead of synthesizing a changed-line marker, so genuine source equal to the old marker remains reviewable. The prompt explicitly attacks workflow-event authority plus mutable-alias, TOCTOU, identity, oracle, contract, authority, dependency-context, coercion, and state-machine failure shapes without fabricating benchmark claims.
- **Fail closed on fabricated Noema execution and external-source provenance (#1641).** Model-authored claims that runtime behavior, command output, toolchain help, or authoritative external documentation confirmed a conclusion now require an out-of-band typed receipt and an exact receipt citation. The isolated reviewer may still reason from changed source and recommend toolchain-specific verification; it cannot present that recommendation as executed evidence. This regression is grounded in `ConceptWeave#35@a31ae0c2`, where review `5120903874` claimed Cargo runtime/documentation confirmation although required Noema run `33938445009` executed no Cargo or documentation lookup step.
- **Pin `opencode-review-dispatch.yml` off the starved floating `ubuntu-latest` image.**
The 2026-09-01 floating-image fix (see that entry below) pinned `strix.yml`,
`opencode-review.yml`, and `noema-review.yml` -- the three required-check
Expand Down Expand Up @@ -1458,3 +1464,5 @@ Semantic Versioning where the repository publishes a release.
- Added an organization-owned reusable exact-artifact SBOM attestation boundary that validates inert six-file wheel/sdist evidence, binds CycloneDX 1.7 predicates to exact SHA-256 subjects, signs through least-privilege GitHub artifact attestations, and exports online and offline verification bundles.
- Hardened exact-artifact SBOM verification with strict finite RFC 8259 JSON, integer CycloneDX document versions, deterministic UUIDv5 subject identities, exact filename properties and single SHA-256 root bindings, environment-only shell input transfer, pinned Ubuntu 24.04 quality runners, and checksum-sealed beginner-readable offline evidence. The decision record now cites Bray (2017) so NaN and Infinity cannot be treated as sealed SBOM numbers.
- Recorded the org control-plane architecture, including exact-artifact SBOM attestation, so agents reconstruct the signing trust boundary from the repo instead of private memory.

- Noema review evidence now uses exact class-and-field claim roles and source excerpts instead of a fixed English causal-word heuristic, preserving non-ASCII and symbol-only review evidence without treating keywords as proof.
27 changes: 27 additions & 0 deletions docs/doctoring/noema-observed-defect-corpus-current-main.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Noema observed-defect review corpus

The trusted Noema review gate treats externally demonstrated review misses as executable regression evidence, not as benchmark claims. Material source/test reviews must exercise at least two distinct observed defect classes and every admitted class witness remains bound to an exact changed-side source coordinate.

The current closed taxonomy is: `mutable_alias`, `time_of_check_time_of_use`, `execution_identity`, `coercion_boundary`, `test_oracle`, `cross_contract`, `authority_boundary`, `dependency_context`, and `state_machine_race`. Each class has class-specific witness keys. Witness values are `{path,line,side,source_excerpt,claim_role,observation}` records bound to the probe location. `source_excerpt` must equal the exact changed-side line, and `observation` must quote the exact source line (or `<blank>`) plus a causal/behavioral relation beyond taxonomy labels; ASCII token shape is not admission authority; repeated or differently worded generic labels do not satisfy the deterministic validator.

The outbound strict structured-output schema and the local validator share that same closed contract. Every probe is one nested `anyOf` variant that correlates a single `probe_kind` with exactly its required `class_evidence` keys; every witness field is required and unknown fields are rejected. Only the containing `adversarial_validation` value is nullable for a non-formal comment. This follows the strict structured-output rule that object properties are required (nullable when truly optional) and prevents the gateway from accepting a probe shape that deterministic admission must reject.

The model is explicitly asked to attack mutable/immutability escapes, changing getters/TOCTOU, request or tenant identity confusion, weak/vacuous oracles, cross-contract contradictions, authority overreach, missing causal dependency context, and reliability/security state-machine races. A falsified hypothesis is valid evidence and must not be promoted into a finding merely to satisfy taxonomy diversity. For CI/automation changes, the review prompt also requires checking whether the mutation credential can create the downstream events/checks the state machine depends on.

JSON booleans are rejected as line coordinates even though Python considers `True == 1`: changed-line evidence requires `type(line) is int` and a positive value. Production review calls always provide the complete changed-path manifest, which activates the observed taxonomy; direct validator unit tests may omit that manifest to exercise lower-level generic schema boundaries independently.

This repair is a narrow current-main successor to the heavily diverged PR #1589 evidence lineage. It does not copy CodeRabbitAI or Devin wording and makes no superiority claim.

Exact-head follow-up removes synthetic bounded-diff omission lines from the diff grammar entirely: truncation drops the incomplete final line and carries the separate `truncated` control flag. A genuine source line equal to the historical marker remains admissible, as do short identifiers, symbol-only lines, blank changed lines, and non-ASCII source through exact string equality rather than lexical guessing. Coordinates and source text now come from one parser so future diff fixes cannot desynchronize their trust boundaries.

The exact-head structural follow-up removes the fixed English relation-word list. Formal evidence now carries a schema-derived `claim_role` for each defect-class witness, while the deterministic gate verifies exact source identity, canonical coordinates, role identity, and distinct observations. Semantic causal adequacy remains a reviewer/evaluation responsibility; the validator does not pretend English keyword presence proves causality.

Workflow-local bootstrap or generated commits are not accepted as final review/check proof merely because their source transaction verified locally. The merge candidate must be a workflow-starting successor writer head produced through ordinary owner-side mutation, with the required review and quality checks observed on that exact unchanged head before merge.

Failed HTTP responses belong to the requesting transport. After bounded telemetry
extraction, close the response there; a secondary cleanup exception must not replace
the original typed transport failure. Process cancellation still propagates. Tests
that invoke a redirect handler directly own the resulting HTTPError and must close
it themselves rather than relying on garbage collection. Run the Noema regression
tests with `-W error`; tests in other HTTP consumers do not become passing evidence
merely because this transport was repaired.
21 changes: 21 additions & 0 deletions docs/product-technical-gap-baseline.md
Original file line number Diff line number Diff line change
Expand Up @@ -2626,6 +2626,27 @@ Higgins, S. S., Crepalde, N., & Fernandes, L. (2021). Segmented multiplexity: A

**Residual.** This closes the specific floating-image contribution from these three central workflows; it does not by itself guarantee the organization-wide Actions queue is fully drained, since other repositories' own workflows and any remaining unpinned central workflows may still request the floating image. Worth a follow-up sweep across the rest of `.github/workflows/` and sibling-repo workflows if queuing persists after this lands.


### 2026-09-02 — Noema observed-defect false-negative corpus (#1641)

- **Verified gap:** protected current main admitted Noema adversarial evidence by count/prose identity and compared model line coordinates with Python integers without excluding booleans. Thus `true` could alias line `1`, and two differently worded probes could satisfy material-change diversity without proving distinct observed defect shapes.
- **Repair:** exact changed-side coordinates now require canonical positive integers; production review verdicts use a closed observed-defect taxonomy with class-specific source-bound witnesses whose exact `source_excerpt` must match the cited changed line and whose observation must quote that exact source (or `<blank>`) plus causal behavior without ASCII/token-shape heuristics. Material changes require distinct classes, and the prompt explicitly checks workflow-starting mutation credentials before relying on downstream required checks.
- **Regression evidence:** `tests/test_noema_observed_defect_corpus_current_main.py` is committed before the causal production change and covers boolean aliasing, malformed/unknown class labels, duplicate-class diversity, witness/source binding, a valid multi-class verdict, and rendered prompt coverage.
- **Authority boundary:** no reviewer, provider, routing, merge, or repository-write authority is widened. The taxonomy is evaluation/admission evidence only.

- **Noema exact-source follow-up (PR #1641):** bounded truncation no longer synthesizes a +/- omission line; it drops the incomplete line and carries the separate `truncated` flag. Genuine source equal to the historical marker remains admissible. One parser now emits both changed coordinates and exact source text, while short, symbol-only, blank, and non-ASCII changed lines use exact equality and arbitrary source-adjacent words do not satisfy causal evidence.

- **Noema structural-causality follow-up (PR #1641):** removed fixed English relation-word admission. Each class witness now carries an exact schema-derived `claim_role` plus exact changed-line source text; deterministic validation stays language-neutral and semantic causality is tested through reviewer/evaluation regressions rather than guessed from keywords.

- **Noema strict-schema parity follow-up (PR #1641):** the outbound response schema now correlates every observed `probe_kind` with the exact required class-witness object that production validates. A realistic verdict is applied to both contracts in one regression, and invalid changed-line telemetry reaches the intended coordinate rejection before asserting the one-request boundary.

### 2026-09-05 — Noema executed-evidence provenance boundary (#1641)

- **Observed RED:** `ConceptualWisdomLab/ConceptWeave#35@a31ae0c2df920f2794f7ddb456795b04797ab472` received CHANGES_REQUESTED review `5120903874`, which stated that Cargo CLI documentation and runtime behavior confirmed `cargo generate-lockfile --locked` was unsupported. Required Noema run `33938445009`, job `101256294197`, used trusted workflow source `8272e4f95c253ab067592460cc9288581bf3a422`; its model phase invoked only the isolated Noema gateway client. No Cargo command, help lookup, or official-document retrieval step executed. Cargo 1.98.0's actual help is contrary evidence, but this central repair does not hard-code a Cargo verdict or remove the consumer lockfile guard.
- **Causal boundary:** exact changed-line and observed-defect-class validation proves that a model response is structurally reviewable; it does not prove that prose describing runtime or external documentation was observed. The trusted gate now inspects only model-authored evidence fields and rejects claims of executed/toolchain behavior or authoritative external sources unless an out-of-band typed receipt ID is supplied and cited in the same statement. The current workflow supplies no such receipts. Source-only reasoning and explicit verification directions remain admissible.
- **Fail-closed preservation:** a missing, wrong-type, or uncited receipt produces no usable verdict. Self-approval, blanket warning suppression, toolchain assumptions, and hard-coded consumer approval are not introduced. Future command/document preprocessors must bind receipt type and ID outside model-controlled context before enabling those claim classes.
- **Scope:** this is central reviewer-evidence provenance only. ConceptWeave source and PR state remain read-only to this owner; the prior CHANGES_REQUESTED review is not dismissed or converted to approval by this change.

## 2026-09-02 GitHub Actions review sidecar pool pinned to `orchestrator/free`; `auto` removed as an accepted value

**Problem.** `scripts/ci/contextual_orchestrator_review_sidecar.sh` — the script every central required review workflow (Strix, OpenCode Review, Noema Review, the PR-review autofix sidecar) provisions to talk to `contextual-orchestrator` — read an operator-settable `CONTEXTUAL_ORCHESTRATOR_POOL` environment variable, defaulted it to `free`, and validated it against exactly two accepted values: `free` or `auto` (`case "$orchestrator_pool" in free|auto) ...`). `auto` is a real, load-bearing value one layer down: `scripts/ci/contextual_orchestrator_review_launcher.py --pool auto` admits *priced* discovered routes as a fallback stage once the free pool is exhausted (`build_zdr_prioritized_catalog(..., pool="auto")`), by design, for callers that want that behavior. Nothing in this repository's own review-provisioning code path currently sets `CONTEXTUAL_ORCHESTRATOR_POOL=auto` — the only workflow that sets the variable at all, `strix.yml`, sets it to `free`; every other central review workflow simply relies on the script's own `:-free` default — so this was not a live incident, it was an unaudited, structurally-reachable escape hatch: a future edit to any of the four workflows above, or a manually-triggered `workflow_dispatch` with a custom env override, could set `CONTEXTUAL_ORCHESTRATOR_POOL=auto` and the sidecar would accept it silently, with no cost ceiling, no budget/authorization gate, and no reviewer visibility that priced models were now in scope for a required check.
Expand Down
Loading
Loading