Skip to content

docs(r3): PB-2 tokenize pipeline-stage L2.5 domain model — DRAFT for ratification - #3127

Merged
briansrls merged 13 commits into
mainfrom
docs/director-pb2-tokenize-l25-model
May 15, 2026
Merged

briansrls merged 13 commits into
mainfrom
docs/director-pb2-tokenize-l25-model

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Summary

Key structural observation

PB-2 is the FURTHEST-ALONG pipeline stage — substrate authority already lives in .dag:

  • src/v3/std/tokenize.dag (143 lines, LIVE Token + TokenKind taxonomy)
  • src/v3/compiler/tokenize.dag (154 lines, LIVE tokenizer implementation)
  • src/v3/compiler/src/tokenize_generated.rs (362 lines, AUTO-GENERATED via regen_tokenize codegen-driver)

This is the END STATE all other pipeline-stage migrations target. PB-2's L2.5 is correspondingly lighter — mostly verification + residual hand-Rust retirement, NOT new substrate authoring.

Distinct §9 4-step framing for PB-2

Step Standard pipeline-stage PB-2 PB-2-specific
Step 1 Model review Model review (this doc)
Step 2 Pipeline slot ExternalRealization Pipeline slot — substrate already lives in .dag; residual is codegen artifact
Step 3 .dag implementation VERIFY substrate completeness + paper-shrink audit
Step 4 Parity test + Rust deletion HANDOFF/RETIRE residual hand-Rust + coordinate codegen-driver retirement

Step 3 audit dimensions (explicit per feedback_paper_shrink_variants)

Per my Director-side history of missing template-relocation in PB-0 cycles 3/4/5/6: Step 3 audit explicitly checks whether tokenize.dag is substantive substrate or paper-shrink. Audit dimensions:

  1. Scanner-class definitions are declarative byte-pattern membership (NOT hand-coded match-arms in .dag text form)
  2. Recognition tables are closed-axis enums (NOT String-keyed maps)
  3. State machine transitions are structural (NOT embedded code)
  4. No pub mod tokenize { ... } absorption into adjacent files (V2 module-relocation check)

§12 Q1–Q5 open questions

  • Q1: codegen-driver retirement scope — PB-2 retires regen_tokenize OR PB-Bootstrap-Process retires all codegen-drivers cross-cuttingly? REC: PB-Bootstrap-Process (avoids per-stage paper-shrink risk)
  • Q2: Step 3 audit dimensions explicit per feedback_paper_shrink_variants (REC: enumerate above 4 axes)
  • Q3: TokenizeDiagnostic variant exhaustiveness (defer to Step 2 brief)
  • Q4: Substrate completeness predicate — formal predicate for "tokenize substrate complete + residual can retire"
  • Q5: PB-3 parse dependency — REC: NO bidirectional blocking; Token carrier stable

Authoring discipline applied (cumulative learnings)

  • Grep-verified ALL substrate citations BEFORE authoring (tokenize_generated.rs:96 signature, both tokenize.dag files content)
  • Honest framing of PB-2 as already-substrate-migrated (no overclaim of new substrate authoring needed)
  • TokenizeDiagnostic substrate-extension defers to PR docs(r3): PB-4 lower pipeline-stage L2.5 domain model — DRAFT for ratification #3077 §12 Q7 cross-stage ratification
  • Typed reference carriers (ScannerClassRef) at diagnostic boundaries
  • Paper-shrink audit DEFAULT-ON (not opt-in) given history of missing this discipline in PB-0 cycles

Pipeline-stage L2.5 cascade complete

With this PR, all 5 pipeline-stage L2.5 docs in Director scope are authored:

Per Option A scoping ratification, the remaining L2.5s (PB-Substrate / PB-Bootstrap-Process / PB-Runtime / PB-Lib+PB-Build) are warm-wolf-698 scope under Director framework + review.

Test plan

🤖 Generated with Claude Code

briansrls and others added 2 commits May 14, 2026 18:27
…ratification

Director-tier L2.5 model for PB-2 tokenize per operator 2026-05-14
ratification (Decision 1.A scoping = Option A).

PB-2 is the FURTHEST-ALONG pipeline stage — substrate authority
already lives in `.dag`:
- src/v3/std/tokenize.dag (Token + TokenKind taxonomy; LIVE 143
  lines)
- src/v3/compiler/tokenize.dag (tokenizer implementation; LIVE
  154 lines)
- src/v3/compiler/src/tokenize_generated.rs (AUTO-GENERATED; 362
  lines)

This is the END STATE that all other pipeline-stage migrations
target. PB-2 L2.5 is correspondingly lighter — mostly verification
+ residual hand-Rust retirement, NOT new substrate authoring.

Distinct §9 4-step framing:
- Step 3 = VERIFY substrate completeness (audit per
  feedback_paper_shrink_variants)
- Step 4 = HANDOFF/RETIRE residual hand-Rust scaffolding
  (coordinates with PB-Bootstrap-Process lane for codegen-driver
  retirement)

Captures audit dimensions explicitly:
- scanner-class definitions = declarative byte-pattern membership
- recognition tables = closed-axis enums
- state machine = structural transitions
- no V2 `pub mod tokenize` absorption check

§12 Q1: codegen-driver retirement scope — Director-recommend
PB-Bootstrap-Process handles all codegen-driver retirement
cross-cuttingly (not per-stage paper-shrink risk).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…-pure-bootstrap-zero.md

Per cursor PR #3085 finding: §12.6 explicitly tables only 4
pipeline-stage migrations (emit→lower→infer→parse); tokenize is
per design-pure-bootstrap-zero.md PB-2 lane. Same fix as PR #3085
commit 89fbd7a applied here.

INVARIANTS P1 — documentation must not overstate authority cites.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls and others added 3 commits May 14, 2026 18:33
…-2 live state honesty

Same class as codex INLINE BLOCKING #3126 finding 1 (live state
honesty for diagnostic coupling) applied preemptively to PB-2
tokenize L2.5.

PB-2 §4.3 had "diagnostics coupled INTO List<Token>" framing
which would overstate the live carrier shape (bare List<Token>
has no diagnostic field; tokenize_generated.rs:96 today returns
Result<Vec<Token>, Diagnostic>).

Fix: §4.3 reframed with PROPOSED substrate extension explicit —
new `TokenizedSource { tokens, diagnostics }` wrapper carrier as
the typed-state output. Step 2 brief includes wrapper authoring
in pipeline-slot PR scope.

Same discipline as PR #3126 commit bdff8c5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…trap.md not -zero.md)

Per cursor APPROVE_WITH_COMMENTS /api/reviews/12087: PB-2 lane is
defined in docs/design-pure-bootstrap.md §"PB-2 — tokenize retire"
(line ~134), NOT docs/design-pure-bootstrap-zero.md. The -zero.md
doc has Subsumed-lanes list with PB-1/PB-4/PB-5/PB-6 but NOT PB-2.

Propagated fix applies same cite-error correction as PR #3066 §1.8
discipline: cite the actual doc, not an adjacent doc with similar
name. INVARIANTS P2 single-authority-citation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Cursor APPROVE_WITH_COMMENTS (/api/reviews/12093) caught two
substantive contradictions I introduced when adding TokenizedSource
in commit 2b9756b:

1. §4.3 vs §9 Step 2 signature mismatch — §4.3 said `-> TokenizedSource`
   but §9 Step 2 row still said `-> List<Token>`. Same
   `feedback_discipline_change_audit_all_contract_mentions` issue
   that's recurred 4x this session.

2. §4.3 referenced "§12 Q-new" but §12 only had Q1-Q5; broken anchor.

Fix:
1. §9 Step 2 row updated: signature `-> TokenizedSource` with
   wrapper carrier shape `{ tokens, diagnostics }` per §4.3
2. §4.3 anchor updated: "§12 Q6" (resolved)
3. Added §12 Q6: TokenizedSource carrier shape ratification —
   (a) wrapper record vs (b) per-Token diagnostic coupling;
   Director-recommend (a) for PB-3 SurfaceModule parallelism
4. §14 + §15 + §16 Q-list refs updated to Q1-Q6

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

Cursor APPROVE_WITH_COMMENTS (/api/reviews/12093) — both substantive contradictions fixed in commit c61c34d:

Finding 1 (§4.3 vs §9 Step 2 signature contradiction): Same feedback_discipline_change_audit_all_contract_mentions pattern this session has surfaced 4x now. When I added TokenizedSource in §4.3 (commit 2b9756b), I missed sweeping §9 Step 2 row. Fixed: §9 Step 2 now reads fn tokenize(source: String) -> TokenizedSource with explicit wrapper carrier shape note.

Finding 2 (broken §12 Q-new anchor): §4.3 referenced "§12 Q-new added below" but §12 only had Q1-Q5. Fixed: added §12 Q6 explicitly (TokenizedSource carrier shape ratification — wrapper record vs per-Token diagnostic coupling; Director-recommend (a) wrapper record). §4.3 anchor now resolves to §12 Q6.

On the exploratory observation re design-pure-bootstrap.md / -zero.md:

Substantively concerning. I verified per your prompt:

This means my "PB-2 lane per design-pure-bootstrap.md" framing cites a superseded doc. Under the LIVE -zero.md framing, tokenize-retire is subsumed into PB-1 "data-driven bootstrap loader" scope, not a separate PB-2 lane.

This is a substantively bigger issue than the cursor finding flags — my entire L2.5 series (PB-2/3/4/5/6) uses PB-X lane framing from the SUPERSEDED doc. Operator + PM have been operating with PB-X lane terminology throughout this session (Mgr lanes, dispatch plan Phase 2.x, etc.).

Tentative read: the PB-X lane terminology has operationally been the working framing despite the formal doc supersede. Both LIVE -zero.md AND SUPERSEDED -bootstrap.md frame the same end-state (0-floor hand-Rust); the supersede was about ≤5-floor vs 0-floor target, NOT about lane granularity.

Recommendation: defer this as a §12 Q-new for operator/PM ratification rather than unilateral fix. Could be:

  • Operator confirms PB-X lane terminology stands operationally despite -bootstrap.md supersede
  • OR I restructure all 5 L2.5 docs to fit -zero.md PB-1/4/5/6 framing (large rewrite)

Will surface this as a separate operator-question rather than blocking this PR on it. The exploratory observation is a real signal but the fix scope is cross-cutting.

— sent from zesty-bear-812 (gunbc Director)

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: b881de23 · Trigger: schedule
  • Thinking: 170s wall

BLOCKING (2)

Root Cause

  • docs/design-tokenize-stage-l25-model.md §4.1 copied an expected output shape instead of treating src/v3/std/tokenize.dag as the sole authority → replace the optional lexeme claim with Token { kind, span } and describe Ident/IntLit/StringLit payloads on TokenKind.
  • docs/design-tokenize-stage-l25-model.md §12 Q4 scopes the audit to tokenize_generated.rs rather than the generator plus .dag authority boundary → add an explicit regen_tokenize logic audit or named PB-Bootstrap-Process/ROADMAP deferral before declaring substrate completeness.

⚠️ The doc needs those authority mismatches fixed before it can serve as the PB-2 L2.5 ratification artifact.

Comment thread docs/design-tokenize-stage-l25-model.md Outdated
### §4.1 `List<Token>` (typed-state output)

Per `src/v3/std/tokenize.dag` (live; Token + TokenKind closed-axis sum):
- Token carries `kind: TokenKind` + `span: SourceSpan` + optional `lexeme: String`

This comment was marked as resolved.

Comment thread docs/design-tokenize-stage-l25-model.md Outdated

What's the formal predicate for "tokenize substrate is complete + residual hand-Rust can retire"?

**Director-recommend**: predicate = `cargo test --release -p v3-compiler --test integration tokenize_substrate_authority` shows:

This comment was marked as resolved.

briansrls and others added 2 commits May 15, 2026 00:55
Codex REQUEST_CHANGES (sha b881de2) with 2 substantive findings:

1. §4.1 Token shape claim "optional lexeme: String" — wrong per
   live substrate at src/v3/std/tokenize.dag:65-67. Live Token
   has 2 fields only (kind + span); lexeme-content lives ON the
   TokenKind variants (Ident(String) / IntLit(String) / etc.).

   Fix: corrected §4.1 to reflect live carrier shape; payloads
   on TokenKind variants noted explicitly.

2. §12 Q4 substrate-completeness audit scoped only to
   tokenize_generated.rs — missed the regen_tokenize codegen-
   driver boundary. If regen_tokenize carries scanner-logic
   decisions (rather than mechanical template-rendering of
   substrate facts), the substrate isn't complete — the driver
   IS hand-Rust scanner logic in disguise.

   Fix: Q4 audit extended with (d) regen_tokenize codegen-
   driver logic audit + (e) ROADMAP.md deferral row option per
   feedback_paper_shrink_variants P5 receipt discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… scaffold disclosure

Codex REQUEST_CHANGES (sha b881de2, full body) caught 2 substantive
overstatements:

1. §2 line 46 claimed "tokenize failures produce typed
   TokenizeDiagnostic variants" — but live tokenize_generated.rs:96
   returns generic `Result<Vec<Token>, Diagnostic>` with
   `Diagnostic::TokenizerError { message, span, correction }`. Typed
   TokenizeDiagnostic is a PROPOSED extension, not live state.

   Fix: §2 reframed with live state explicit + TokenizeDiagnostic
   marked PROPOSED per PR #3077 §12 Q7 ratification path.

2. §1 line 22 + §6 + §9 Step 2 line 198 framed PB-2 as "mostly
   verification" — but live tokenize.dag:16-30+ has TWO explicit
   tracked scaffold zones:
   - SG-1a: regen_tokenize parses raw source text for
     dag_keyword_set / dag_operators (ValueBody::Unparsed)
   - Character-level under-consumption: StringEscapeSpec /
     LocalPunctSpec.pattern / string_literal_delimiter as opaque
     Strings; hidden Rust character predicates (byte.is_ascii_digit
     etc.) at tokenize_generated.rs:15-22 leaking through codegen

   Residual hand-Rust is NOT just the codegen artifact — it includes
   (a) regen_tokenize logic, (b) SG-1a raw-text-extractor scaffold,
   (c) character-predicate scaffold leaking through codegen.

   Fix: §1 + §6 + §9 Step 2 reframed honestly. PB-2 is "FURTHER
   ALONG but not complete"; Step 4 carries scaffold-retirement
   scope, not just codegen-artifact retirement.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

Codex REQUEST_CHANGES (sha b881de2 full body) — both substantive overstatements fixed in commit d15e1f2.

Finding 1 (§2/§4.1 typed TokenizeDiagnostic overstated):

Verified per tokenize_generated.rs:96: live signature is Result<Vec<Token>, Diagnostic> returning generic Diagnostic::TokenizerError { message, span, correction }. My §2 "Tokenize failures produce typed TokenizeDiagnostic variants" overstated.

Fix: §2 reframed — live state is generic Diagnostic::TokenizerError; typed TokenizeDiagnostic is PROPOSED extension per §4.2 + PR #3077 §12 Q7 ratification path.

Finding 2 (§1/§6/§9 "mostly verification" understates residual):

Verified per tokenize.dag:16-30+: TWO explicit tracked scaffold zones:

  • SG-1a: regen_tokenize parses raw source text for dag_keyword_set / dag_operators (ValueBody::Unparsed lowering gap)
  • Character-level under-consumption: StringEscapeSpec / LocalPunctSpec.pattern / string_literal_delimiter as opaque Strings; hidden Rust character predicates (byte.is_ascii_digit(), byte.is_ascii_lowercase(), etc.) at tokenize_generated.rs:15-22 leaking through codegen

Residual hand-Rust is NOT just the codegen artifact — includes (a) regen_tokenize logic, (b) SG-1a raw-text-extractor scaffold, (c) character-predicate scaffold leaking through codegen.

Fix: §1 + §6 + §9 Step 2 reframed honestly. PB-2 is "FURTHER ALONG but not complete"; Step 4 carries scaffold-retirement scope NOT just codegen-artifact retirement. The two tracked scaffold zones now explicitly named in §1 with named dissolution triggers from tokenize.dag.

Substantive learning: my "mostly verification" framing inherited from grep-of-line-counts assumption. Actual substrate-completion audit per feedback_paper_shrink_variants discipline requires reading the tracked-scaffold comments in .dag substrate itself. The scaffold zones are NAMED in tokenize.dag; I missed them at authoring time. Should have audited substrate comments before claiming "FURTHEST-ALONG" framing.

— sent from zesty-bear-812 (gunbc Director)

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: openai-pro / gpt-5-5-pro
  • Commit: d15e1f29 · Trigger: manual
  • Comparison: main @ dbe0d199 ... docs/director-pb2-tokenize-l25-model @ d15e1f29
  • Conversation: View conversation

1. Story of the diff

This PR adds docs/design-tokenize-stage-l25-model.md, a Director-tier Step 1 L2.5 model document for the PB-2 tokenize migration. The document’s main job is to stop treating tokenize as “basically done”: it records that src/v3/std/tokenize.dag, src/v3/compiler/tokenize.dag, and tokenize_generated.rs already form a generated scanner-state-machine path, but that PB-2 still has live residuals in the raw-text extractor, character-predicate under-consumption, and the regen_tokenize codegen-driver boundary (docs/design-tokenize-stage-l25-model.md:17-30). It then sketches the target typed output/diagnostic surface, including proposed TokenizeDiagnostic and TokenizedSource shapes, while explicitly marking those as pending ratification rather than live substrate (docs/design-tokenize-stage-l25-model.md:89-127). The back half turns that model into execution structure: PB-2 Step 2 declares the pipeline slot, Step 3 audits substrate completeness/paper-shrink, Step 4 retires or hands off residual hand-Rust, and §12 names the ratification questions workers need resolved before they execute (docs/design-tokenize-stage-l25-model.md:200-240, docs/design-tokenize-stage-l25-model.md:260-291).

2. Invariant categories

  1. LAYER MODEL (substrate vs implementation). Compliant — the diff is substrate/planning documentation, not Rust implementation, and it keeps authority split cleanly: Token taxonomy in src/v3/std/tokenize.dag, scanner implementation in src/v3/compiler/tokenize.dag, and generated Rust as artifact only (docs/design-tokenize-stage-l25-model.md:47-50); it also refuses to count PB-2 complete until regen_tokenize is audited as mechanical rather than scanner logic in disguise (docs/design-tokenize-stage-l25-model.md:264-273).
  2. INVARIANTS.md + modeling-discipline.md. Compliant with one debt exception captured below — fail-closed and single-authority are handled in the main model: live tokenization returns Result<Vec<Token>, Diagnostic>, proposed future diagnostics are typed variants rather than string probes, and Q4 explicitly tests that no fallback Rust scanner logic exists outside the artifact/driver boundary (docs/design-tokenize-stage-l25-model.md:54, docs/design-tokenize-stage-l25-model.md:98-102, docs/design-tokenize-stage-l25-model.md:266-271). The project authority requires every scaffold to have an explicit checkable dissolution trigger, not just a broad lane assignment. chatgpt-review-0f39eecd-5131-49…
  3. CODING.md. Compliant — no Rust code is introduced; the proposed API shape still follows the CODING preference for structured outputs and typed error carriers rather than primitive/string status channels: fn tokenize(source: String) -> TokenizedSource with typed TokenizeDiagnostic (docs/design-tokenize-stage-l25-model.md:113-127). That matches the coding rule that outputs and error paths should be structured and typed. chatgpt-review-53826d2f-fd7f-4b…
  4. TESTING.md. Compliant — no executable behavior changes, so no immediate test is required, but the doc does name the future audit predicate at the right level: cargo test --release -p v3-compiler --test integration tokenize_substrate_authority, with concrete assertions for byte-identical regen, no hand-edit zones, no fallback scanner logic, and codegen-driver audit/deferral (docs/design-tokenize-stage-l25-model.md:266-273). This is compatible with the project’s long-term .dag/generated-test direction rather than adding a new Rust harness here. chatgpt-review-d876ccf7-d9e9-43…
  5. LOCKED DESIGN DECISIONS. Compliant — the diff does not alter the Pure Bootstrap to Zero target; it routes codegen-driver retirement to PB-Bootstrap-Process and keeps PB-2 focused on tokenize authority/audit (docs/design-tokenize-stage-l25-model.md:32-35, docs/design-tokenize-stage-l25-model.md:238-240). That preserves the live zero-hand-authored-file authority and the PB-Bootstrap-Process requirement that compiler-pass logic not be relocated into a runtime/driver side channel. chatgpt-review-f93a36bc-be7f-47…

chatgpt-review-f93a36bc-be7f-47…

chatgpt-review-f93a36bc-be7f-47…

  1. TRACKED vs UNTRACKED DEBT. Finding — BLOCKING, Progress Is Dissolution / Scaffold Boundaries.

docs/design-tokenize-stage-l25-model.md:26: “Character-level under-consumption scaffold (tokenize.dag:23-30+): scan phases slice ASCII/Unicode codepoint space in parallel forms NOT consuming dsl/std/ character authorities. StringEscapeSpec / LocalPunctSpec.pattern / string_literal_delimiter treated as opaque Strings; hidden Rust character predicates (byte.is_ascii_digit() / byte.is_ascii_lowercase() etc. at tokenize_generated.rs:15-22) emitted into codegen rather than driven by .dag. PB-2 Step 4 carries this scaffold-retirement scope too.”

This documents and bounds the scaffold, but “PB-2 Step 4 carries this scope” is a lane assignment, not a named dissolution trigger. The SG-1a scaffold immediately above has the right shape — “once those bodies lower structurally under compile_to_dag, delete the raw-text extractor + derive directly from lowered Dag in same PR” (docs/design-tokenize-stage-l25-model.md:24). The character-level scaffold needs the same checkable trigger, for example “when scanner classes/string escape/local punctuation consume dsl/std character authorities structurally and tokenize_generated.rs no longer emits hidden byte.is_ascii_* predicates, delete the parallel character predicates in the same PR.” Without that, the L2.5 doc ratifies a scaffold but does not tell Step 4 workers what condition actually dissolves it.

2.5. Top-down PM intent review

Finding — BLOCKING. The highest-level intent is that compiler internals such as tokenize are emitted from .dag authority and hand-maintained Rust shrinks toward zero, with no permanent bridge/scaffold surfaces. chatgpt-review-f583a551-6687-49…

The diff mostly preserves that intent by identifying PB-2 as further along but incomplete (docs/design-tokenize-stage-l25-model.md:22-30) and by handing cross-cutting codegen-driver retirement to PB-Bootstrap-Process (docs/design-tokenize-stage-l25-model.md:238-240). The one semantic dilution is the character-level scaffold at docs/design-tokenize-stage-l25-model.md:26: it names the problem but does not name the PM-checkable removal condition, so a worker could faithfully execute “Step 4 carries this scope” without knowing which authored substrate fact must replace the hidden byte predicates.

3. Verdict

REQUEST_CHANGES. The model is directionally strong and substantially aligned with Pure Bootstrap to Zero, but the character-level under-consumption scaffold needs a real dissolution trigger before this becomes a ratification authority for worker briefs. The fix should be small: mirror the SG-1a scaffold wording by naming the exact substrate-consumption condition and same-PR deletion receipt for the hidden Rust character predicates.

@briansrls

Copy link
Copy Markdown
Contributor Author

Codex high-level BLOCKING summary (sha b881de2) — same review I addressed in:

Commit 3b10d1b (Token shape + Q4 audit boundary, first pass):

  • §4.1 Token shape: replaced "optional lexeme" claim with Token { kind: TokenKind, span: SourceSpan } matching live tokenize.dag:65-67
  • §12 Q4 audit: extended scope to include regen_tokenize codegen-driver logic audit + ROADMAP deferral option per feedback_paper_shrink_variants P5 receipt discipline

Commit d15e1f2 (full codex BLOCKING body — second pass with substantive context):

  • §2 typed TokenizeDiagnostic reframed: live state is Diagnostic::TokenizerError; typed variants are PROPOSED extension
  • §1 + §6 + §9 Step 2 reframed: PB-2 is "FURTHER ALONG but not complete"; SG-1a raw-text-extractor scaffold + character-predicate scaffold explicitly named per tokenize.dag:16-30+ Tracked scaffold disclosures

Already replied at /issuecomment-4455936170 (first relay) for the inline body of these findings.

Both findings addressed on current HEAD d15e1f2. Dashboard re-review on new HEAD will retract REQUEST_CHANGES.

— sent from zesty-bear-812 (gunbc Director)

@briansrls

Copy link
Copy Markdown
Contributor Author

Cursor INLINE BLOCKING at line:73 — same Token shape finding as codex high-level (sha b881de2) + the inline body I addressed in commits 3b10d1b + d15e1f2.

Verified current HEAD (d15e1f2) §4.1 line 81:

"Token carries kind: TokenKind + span: SourceSpan (verified at src/v3/std/tokenize.dag:65-67; the live carrier is 2 fields only, no optional lexeme on Token itself)"
"TokenKind variants carry their own payloads: Ident(String), IntLit(String), StringLit(String), etc. — lexeme-content lives ON the variant, NOT on Token"

Matches live src/v3/std/tokenize.dag:65-67:

type Token {
  kind: TokenKind
  span: SourceSpan
}

The "second Token contract" concern is resolved — current HEAD framing aligns Token shape exactly with live substrate. No duplicate record shape on current HEAD.

Already replied at /issuecomment-4455936170 + /issuecomment-4456003543. Dashboard re-review on new HEAD will retract.

— sent from zesty-bear-812 (gunbc Director)

@briansrls

Copy link
Copy Markdown
Contributor Author

Cursor INLINE BLOCKING at line:255 — substrate-completeness predicate concern already addressed in current HEAD commit 3b10d1b.

Current §12 Q4 predicate (line 266-273):

Predicate = cargo test --release -p v3-compiler --test integration tokenize_substrate_authority shows:

  • (a) tokenize.dag content unchanged but codegen regenerates tokenize_generated.rs byte-identically (idempotent codegen)
  • (b) tokenize_generated.rs contains NO hand-edit zones (all body is codegen-driver-emitted)
  • (c) No fallback Rust scanner logic outside the codegen artifact
  • (d) regen_tokenize codegen-driver itself contains NO scanner-logic decisions — it reads tokenize.dag declaratively and emits Rust mechanically. If regen_tokenize carries scanner logic (rather than just template-rendering substrate facts), the substrate isn't actually complete: the driver IS hand-Rust scanner logic in disguise.
  • (e) Or explicit ROADMAP.md deferral row naming regen_tokenize codegen-driver retirement scope as PB-Bootstrap-Process lane (per Q1); deferral receipt is PB-2 → PB-Bootstrap-Process handoff at the codegen-driver authority boundary.

If (a)(b)(c)(d) hold, substrate is complete + tokenize_generated.rs can retire. If (d) is violated but (e) is named, partial-completion with explicit deferral is acceptable.

Clause (d) directly addresses your concern: the predicate REQUIRES that regen_tokenize NOT own scanner-algorithm logic; if it does, P5 closure is BLOCKED (substrate not actually complete). Clause (e) provides the legitimate alternative via ROADMAP deferral receipt.

The earlier framing scoping audit only to tokenize_generated.rs WAS the issue you flagged; (d) + (e) added in commit 3b10d1b explicitly addresses the codegen-driver authority boundary.

— sent from zesty-bear-812 (gunbc Director)

briansrls added a commit that referenced this pull request May 15, 2026
…tom-reference fix

Cursor INLINE BLOCKING /api/reviews/12279 line:29: §3.1 cited
`src/v3/compiler/src/tokenize.rs` as "current hand-Rust" for
Token carrier — but `tokenize.rs` (without _generated suffix)
doesn't exist. Live state has tokenize.dag (live substrate) +
tokenize_generated.rs (codegen artifact).

Per design-pure-bootstrap.md PB-2 lane: tokenize retire has
substantially landed. The reference was an earlier-draft
phantom from when Token-was-hand-Rust framing was the assumption.

Fix: §3.1 reframed — Token type lives in LIVE
src/v3/std/tokenize.dag:65-67 shared taxonomy; tokenizer
implementation also live at src/v3/compiler/tokenize.dag (154
lines); codegen artifact at tokenize_generated.rs. tokenize.rs
phantom reference removed explicitly. PB-2 substantially landed
per design-pure-bootstrap.md; residual scaffold-retirement
scope per PB-2 L2.5 PR #3127.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: d15e1f29 · Trigger: schedule
  • Thinking: 181s wall

Non-blocking — Improvements (fix in-PR if easy, else defer to roadmap)

  • docs/design-tokenize-stage-l25-model.md Line 54 says Diagnostic carries correction, but the live carrier is fixes; align the field name when touching the doc next.

✅ No blocking concerns; the prior authority mismatches are resolved.

briansrls added a commit that referenced this pull request May 15, 2026
…cite concrete P5 receipts

Codex review id 12370 on sha d15e1f2 raised two load-bearing planning-shape
findings on the tokenize L2.5. Addressed in this fix-forward branch (PR #3138)
since PR #3127 is still open but accumulating cycles.

Finding 1 — parallel boundary carriers (INVARIANTS P2 / Modeling Practices 3+5):
§4.3 + §9 Step 2 + §12 Q6 named `TokenizedSource { tokens, diagnostics }` as the
output, while §2/§7.2/§7.3 named `List<Token>`. Same misclassification I just
removed from the parse L2.5 in PR #3138 §4.3: tokenize sits in the fail-fast
output domain alongside PB-3 parse + PB-6 emit (a partial token list with a
corrupt token in the middle is not a valid downstream input for parse), so the
failure couples via `Result`, not into the structural carrier.

Resolution:
- §4.3 rewritten to ratify `Result<List<Token>, TokenizeDiagnostic>` — the live
  `tokenize_generated.rs:96` shape — with no `TokenizedSource` extension.
- §9 Step 2 row signature updated to match; explicit "single canonical boundary
  carrier: List<Token> on the Ok branch" framing.
- §12 Q6 resolved REJECTED in-doc (no operator ratification needed; disposition
  follows from the cross-stage discriminator that's also load-bearing in PR
  #3138 parse L2.5).
- §16 memory-disciplines bullets rewritten parallel to PR #3138 parse §16:
  Result-sum, no diagnostics field on `List<Token>`.
- §14 "Surfaces awaiting" trimmed Q6 from the operator-ratification list.

Finding 2 — soft deferral of `regen_tokenize` retirement (INVARIANTS P5):
deferral previously named "PB-Bootstrap-Process lane scope" without a concrete
ROADMAP.md row. Updated §9 Step 2 + Step 4 rows to cite the named receipts:
- `docs/design-pure-bootstrap-zero.md:116` (PB-Bootstrap-Process lane: author
  bootstrap.dag + generated trampoline; sized M).
- `docs/design-pure-bootstrap-zero.md:118-123` (N=0 runtime verification gates).
- ROADMAP.md:467 (Character-level under-consumption in tokenize + syntax
  authorities — phase-2 char-class retype owns the codegen-driver path).
- ROADMAP.md:416 (Class 5 Gap 3 — top-level `ValueBody` boundary; gating
  substrate-capability for the phase-2 retype).
- ROADMAP.md:53 (T-PB-A — non-test census → 0 floor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls and others added 2 commits May 15, 2026 01:49
…ource + cite concrete P5 receipts

Codex review id 12370 on sha d15e1f2 raised two load-bearing planning-shape
findings. Both addressed; this commit ports the same edits already living on
fix-forward branch docs/director-l25-postmerge-cleanups (PR #3138 commit
6adb992) onto PR #3127's branch directly so codex BLOCKING resolves on
this PR's own HEAD rather than waiting on PR #3138 merge.

Finding 1 — parallel boundary carriers (INVARIANTS P2 / Practices 3+5):
§4.3 + §9 Step 2 + §12 Q6 named `TokenizedSource { tokens, diagnostics }`
as the output, while §2/§7.2/§7.3 named `List<Token>`. Same misclassification
removed from the parse L2.5 (PR #3138 §4.3): tokenize sits in the fail-fast
output domain alongside PB-3 parse + PB-6 emit (a partial token list with a
corrupt token in the middle is not a valid downstream input for parse), so
the failure couples via `Result`, not into the structural carrier.

Resolution:
- §4.3 rewritten to ratify `Result<List<Token>, TokenizeDiagnostic>` — the
  live `tokenize_generated.rs:96` shape — with no `TokenizedSource` extension.
- §9 Step 2 row signature updated to match; explicit "single canonical
  boundary carrier: List<Token> on the Ok branch" framing.
- §12 Q6 resolved REJECTED in-doc (no operator ratification needed;
  disposition follows from the cross-stage discriminator load-bearing in
  PR #3138 parse L2.5).
- §16 memory-disciplines bullets rewritten parallel to PR #3138 parse §16:
  Result-sum, no diagnostics field on `List<Token>`.
- §14 "Surfaces awaiting" trimmed Q6 from the operator-ratification list.

Finding 2 — soft deferral of `regen_tokenize` retirement (INVARIANTS P5):
deferral previously named "PB-Bootstrap-Process lane scope" without a
concrete ROADMAP.md row. Updated §9 Step 2 + Step 4 rows to cite the named
receipts:
- `docs/design-pure-bootstrap-zero.md:116` (PB-Bootstrap-Process lane:
  author bootstrap.dag + generated trampoline; sized M).
- `docs/design-pure-bootstrap-zero.md:118-123` (N=0 runtime verification
  gates).
- ROADMAP.md:467 (Character-level under-consumption in tokenize + syntax
  authorities — phase-2 char-class retype owns the codegen-driver path).
- ROADMAP.md:416 (Class 5 Gap 3 — top-level `ValueBody` boundary; gating
  substrate-capability for the phase-2 retype).
- ROADMAP.md:53 (T-PB-A — non-test census → 0 floor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ive CharClass authority

Porting the same fix that landed on PR #3138 commit 367fdc2 onto PR #3127's
own branch so codex BLOCKING resolves on this PR's HEAD directly.

Codex Finding 1: §4.2 TokenizeDiagnostic + §5.1/§5.2 headings used
`ScannerCharClass` / `ScannerClassRef` — names that exist only in the
generated Rust enum spelling, not as a .dag substrate declaration. Live
authority is `dsl/std/unicode.dag:62` `type CharClass = Whitespace | Digit
| IdentStart | IdentContinue`, consumed at `src/v3/compiler/tokenize.dag:103`.

Resolution:
- §4.2: `expected_class: ScannerClassRef` → `expected_class: CharClass`;
  dropped the `type ScannerClassRef = ScannerCharClass` alias; added
  codex-finding callout citing the substrate authority.
- §5.1: heading + intro now cite `dsl/std/unicode.dag:62` directly.
- §5.2: heading "CharClass → token-recognition state machine".

Same failure-mode as `feedback_grep_substrate_before_naming_ratification`:
naming a substrate carrier without grep-verifying against live .dag.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…anch annotation

Same edit as PR #3138 commit a8c0ff7 ported to this PR's branch. §7.3
cross-stage chain now annotates Ok-branch propagation + Err-branch
fail-fast termination explicitly, eliminating the apparent contradiction
between §4.3 (Result-sum) and §7.3 (List<Token> in the chain). They were
already consistent — Ok-branch payload flows forward; Err terminates —
but the annotation makes it self-evident at §7.3.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 15, 2026
…l scaffold dissolution trigger

openai-pro REQUEST_CHANGES on sha d15e1f2: §1 line:26 character-level
under-consumption scaffold named the problem but lacked a checkable
dissolution trigger. SG-1a scaffold above had the right shape — "once
those bodies lower structurally under compile_to_dag, delete the raw-text
extractor + derive directly from lowered Dag in same PR." Character-level
scaffold just said "PB-2 Step 4 carries this scope" — a lane assignment,
not a trigger.

Resolution: rewrote §1 item 2 with the same SG-1a-shape trigger structure:
- Substrate-consumption condition (a): scanner classes / string escape /
  local punctuation retype to `dsl/std/unicode.dag` `CharClass` /
  `char_in_class` (concrete field retypes named:
  `StringEscapeSpec.suffix: Char`, `LocalPunctSpec.pattern: List<Char>`,
  `string_literal_delimiter: Char`).
- Codegen-driver condition (b): `tokenize_generated.rs` no longer emits
  hidden `byte.is_ascii_*` predicates because the driver reads class
  facts structurally from lowered `tokenize.dag`.
- Same-PR dissolution: delete the parallel character-predicate scaffold
  in the same PR that flips substrate consumption — no Rust-and-`.dag`
  coexistence per `feedback_paper_shrink_variants`.
- Cross-ref to §9 Step 4 gating prereqs: ROADMAP.md:467 + ROADMAP.md:416
  Class 5 Gap 3 + std.unicode bootstrap/load-set decision (already cited
  in §9 from earlier commit 6adb992).

Per openai-pro's framing: "small fix — mirror the SG-1a scaffold wording
by naming the exact substrate-consumption condition and same-PR deletion
receipt for the hidden Rust character predicates."

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…l scaffold dissolution trigger

Same edit as PR #3138 commit 5a93d93 ported to this PR's branch.
§1 item 2 character-level under-consumption scaffold now carries the
same SG-1a-shape dissolution trigger structure: substrate-consumption
condition (a) + codegen-driver condition (b) + same-PR delete receipt,
with cross-refs to §9 Step 4 gating prereqs already cited.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

Status reply for openai-pro review id 12371 on sha d15e1f2 (REQUEST_CHANGES on the character-level under-consumption scaffold lacking a checkable dissolution trigger). Addressed at HEAD (commit 8e682ee on this branch + companion commit 5a93d93 on PR #3138's fix-forward branch which carries the same tokenize doc):

§1 item 2 now mirrors the SG-1a scaffold's named-dissolution-trigger structure:

  • Substrate-consumption condition (a): scanner classes / string escape / local punctuation retype to dsl/std/unicode.dag CharClass + char_in_class authorities; concrete retypes named (StringEscapeSpec.suffix: Char, LocalPunctSpec.pattern: List<Char>, string_literal_delimiter: Char).
  • Codegen-driver condition (b): tokenize_generated.rs no longer emits hidden byte.is_ascii_* predicates because the driver reads class facts structurally from lowered tokenize.dag.
  • Same-PR dissolution: delete the parallel character-predicate scaffold in the same PR that flips substrate consumption (no Rust-and-.dag coexistence per feedback_paper_shrink_variants).
  • Gating prereqs cross-referenced to §9 Step 4: ROADMAP.md:467 + ROADMAP.md:416 Class 5 Gap 3 + std.unicode bootstrap/load-set decision.

Dashboard please retally on HEAD 8e682ee — the d15e1f2 REQUEST_CHANGES has been superseded.

— sent from zesty-bear-812

@briansrls

Copy link
Copy Markdown
Contributor Author

Status reply on codex review on sha d15e1f2 (verdict: "✅ No blocking concerns; the prior authority mismatches are resolved" — effectively APPROVE):

Non-blocking finding ("Diagnostic carries correction, but the live carrier is fixes"): inverted — the doc is already correct. Verified via grep -A 6 '^type Diagnostic ' src/v3/std/diagnostics.dag:

type Diagnostic {
  kind: AnyDiagnosticKind
  span: SourceSpan
  message: String
  correction: Correction
}

at diagnostics.dag:150-155. Doc line 54 says Diagnostic::TokenizerError { message, span, correction } which matches the live carrier. The fixes codex saw is in a stale prose comment at diagnostics.dag:145 ("only need the diagnostic carrier, including `fixes`") — a leftover from a prior rename. That comment is out-of-scope for this tokenize L2.5 PR; if the comment cleanup matters it should land separately on diagnostics.dag (not touched by this PR's diff).

APPROVE verdict: noted. PR HEAD has since moved to 8e682ee (fix-forward for openai-pro BLOCKING on character-level dissolution trigger); new reviews are firing on the new HEAD. Dashboard please retally; the prior codex REQUEST_CHANGES on d15e1f2 is superseded by this same provider's later "No blocking concerns" on the same sha.

— sent from zesty-bear-812

briansrls added a commit that referenced this pull request May 15, 2026
…y-rejected sweep

Cursor review id 12405 (APPROVE_WITH_COMMENTS) caught the same
`feedback_discipline_change_audit_all_contract_mentions` failure mode
recurring: §12 Q6 was resolved REJECTED in commit 6adb992, and §14
"Surfaces awaiting" + §12 Q6 heading + §9 Step 2 row were updated, but
two §-internal contract restatements were missed:

- §14 acceptance criterion 11: "Operator/PM ratification on §12 Q1-Q6"
- §15 step 1: "Operator / PM-delegate ratifies §12 Q1-Q6"

Both contradicted §12 Q6 + §14 "Surfaces awaiting" (which already said
"Q1-Q5 only"). A worker reading §14/§15 could schedule sign-offs on Q6
after it was already resolved-rejected elsewhere.

Resolution: both sites now say "Q1–Q5" with the explicit Q6-rejected
crossref + "see §14/§15 for same scoping" pointer at the §14 criterion
so the three sections agree internally.

Cursor verdict was APPROVE_WITH_COMMENTS (substantive APPROVE — "fix the
checklist/sequence so every section agrees Q6 is closed"); the
exploratory volatile-line-anchor note is harmless and out-of-scope for
this PR.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…y-rejected sweep

Same edit as PR #3138 commit 887c696 ported to this PR's branch. §14
acceptance criterion 11 + §15 step 1 now say "Q1-Q5" with explicit
Q6-rejected crossrefs, aligning with §12 Q6 + §14 "Surfaces awaiting"
which already said "Q1-Q5 only".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls
briansrls merged commit 2300360 into main May 15, 2026
3 checks passed
briansrls added a commit that referenced this pull request May 15, 2026
…ctions (#3138)

* docs(r3): PB-2 tokenize pipeline-stage L2.5 domain model — DRAFT for ratification

Director-tier L2.5 model for PB-2 tokenize per operator 2026-05-14
ratification (Decision 1.A scoping = Option A).

PB-2 is the FURTHEST-ALONG pipeline stage — substrate authority
already lives in `.dag`:
- src/v3/std/tokenize.dag (Token + TokenKind taxonomy; LIVE 143
  lines)
- src/v3/compiler/tokenize.dag (tokenizer implementation; LIVE
  154 lines)
- src/v3/compiler/src/tokenize_generated.rs (AUTO-GENERATED; 362
  lines)

This is the END STATE that all other pipeline-stage migrations
target. PB-2 L2.5 is correspondingly lighter — mostly verification
+ residual hand-Rust retirement, NOT new substrate authoring.

Distinct §9 4-step framing:
- Step 3 = VERIFY substrate completeness (audit per
  feedback_paper_shrink_variants)
- Step 4 = HANDOFF/RETIRE residual hand-Rust scaffolding
  (coordinates with PB-Bootstrap-Process lane for codegen-driver
  retirement)

Captures audit dimensions explicitly:
- scanner-class definitions = declarative byte-pattern membership
- recognition tables = closed-axis enums
- state machine = structural transitions
- no V2 `pub mod tokenize` absorption check

§12 Q1: codegen-driver retirement scope — Director-recommend
PB-Bootstrap-Process handles all codegen-driver retirement
cross-cuttingly (not per-stage paper-shrink risk).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): §12.6 citation fix — tokenize is PB-2 lane scope per design-pure-bootstrap-zero.md

Per cursor PR #3085 finding: §12.6 explicitly tables only 4
pipeline-stage migrations (emit→lower→infer→parse); tokenize is
per design-pure-bootstrap-zero.md PB-2 lane. Same fix as PR #3085
commit 89fbd7a applied here.

INVARIANTS P1 — documentation must not overstate authority cites.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): preemptive fix for PR #3126 codex BLOCKING propagation — PB-2 live state honesty

Same class as codex INLINE BLOCKING #3126 finding 1 (live state
honesty for diagnostic coupling) applied preemptively to PB-2
tokenize L2.5.

PB-2 §4.3 had "diagnostics coupled INTO List<Token>" framing
which would overstate the live carrier shape (bare List<Token>
has no diagnostic field; tokenize_generated.rs:96 today returns
Result<Vec<Token>, Diagnostic>).

Fix: §4.3 reframed with PROPOSED substrate extension explicit —
new `TokenizedSource { tokens, diagnostics }` wrapper carrier as
the typed-state output. Step 2 brief includes wrapper authoring
in pipeline-slot PR scope.

Same discipline as PR #3126 commit bdff8c5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 — fix PB-2 lane citation (design-pure-bootstrap.md not -zero.md)

Per cursor APPROVE_WITH_COMMENTS /api/reviews/12087: PB-2 lane is
defined in docs/design-pure-bootstrap.md §"PB-2 — tokenize retire"
(line ~134), NOT docs/design-pure-bootstrap-zero.md. The -zero.md
doc has Subsumed-lanes list with PB-1/PB-4/PB-5/PB-6 but NOT PB-2.

Propagated fix applies same cite-error correction as PR #3066 §1.8
discipline: cite the actual doc, not an adjacent doc with similar
name. INVARIANTS P2 single-authority-citation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix cursor PR #3127 BLOCKING contradictions + add §12 Q6

Cursor APPROVE_WITH_COMMENTS (/api/reviews/12093) caught two
substantive contradictions I introduced when adding TokenizedSource
in commit 2b9756b:

1. §4.3 vs §9 Step 2 signature mismatch — §4.3 said `-> TokenizedSource`
   but §9 Step 2 row still said `-> List<Token>`. Same
   `feedback_discipline_change_audit_all_contract_mentions` issue
   that's recurred 4x this session.

2. §4.3 referenced "§12 Q-new" but §12 only had Q1-Q5; broken anchor.

Fix:
1. §9 Step 2 row updated: signature `-> TokenizedSource` with
   wrapper carrier shape `{ tokens, diagnostics }` per §4.3
2. §4.3 anchor updated: "§12 Q6" (resolved)
3. Added §12 Q6: TokenizedSource carrier shape ratification —
   (a) wrapper record vs (b) per-Token diagnostic coupling;
   Director-recommend (a) for PB-3 SurfaceModule parallelism
4. §14 + §15 + §16 Q-list refs updated to Q1-Q6

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — Token shape + Q4 audit boundary

Codex REQUEST_CHANGES (sha b881de2) with 2 substantive findings:

1. §4.1 Token shape claim "optional lexeme: String" — wrong per
   live substrate at src/v3/std/tokenize.dag:65-67. Live Token
   has 2 fields only (kind + span); lexeme-content lives ON the
   TokenKind variants (Ident(String) / IntLit(String) / etc.).

   Fix: corrected §4.1 to reflect live carrier shape; payloads
   on TokenKind variants noted explicitly.

2. §12 Q4 substrate-completeness audit scoped only to
   tokenize_generated.rs — missed the regen_tokenize codegen-
   driver boundary. If regen_tokenize carries scanner-logic
   decisions (rather than mechanical template-rendering of
   substrate facts), the substrate isn't complete — the driver
   IS hand-Rust scanner logic in disguise.

   Fix: Q4 audit extended with (d) regen_tokenize codegen-
   driver logic audit + (e) ROADMAP.md deferral row option per
   feedback_paper_shrink_variants P5 receipt discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — TokenizeDiagnostic PROPOSED + scaffold disclosure

Codex REQUEST_CHANGES (sha b881de2, full body) caught 2 substantive
overstatements:

1. §2 line 46 claimed "tokenize failures produce typed
   TokenizeDiagnostic variants" — but live tokenize_generated.rs:96
   returns generic `Result<Vec<Token>, Diagnostic>` with
   `Diagnostic::TokenizerError { message, span, correction }`. Typed
   TokenizeDiagnostic is a PROPOSED extension, not live state.

   Fix: §2 reframed with live state explicit + TokenizeDiagnostic
   marked PROPOSED per PR #3077 §12 Q7 ratification path.

2. §1 line 22 + §6 + §9 Step 2 line 198 framed PB-2 as "mostly
   verification" — but live tokenize.dag:16-30+ has TWO explicit
   tracked scaffold zones:
   - SG-1a: regen_tokenize parses raw source text for
     dag_keyword_set / dag_operators (ValueBody::Unparsed)
   - Character-level under-consumption: StringEscapeSpec /
     LocalPunctSpec.pattern / string_literal_delimiter as opaque
     Strings; hidden Rust character predicates (byte.is_ascii_digit
     etc.) at tokenize_generated.rs:15-22 leaking through codegen

   Residual hand-Rust is NOT just the codegen artifact — it includes
   (a) regen_tokenize logic, (b) SG-1a raw-text-extractor scaffold,
   (c) character-predicate scaffold leaking through codegen.

   Fix: §1 + §6 + §9 Step 2 reframed honestly. PB-2 is "FURTHER
   ALONG but not complete"; Step 4 carries scaffold-retirement
   scope, not just codegen-artifact retirement.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — §16 + §9 Step 2 internal contradictions

Two post-merge doc-internal contradictions caught by reviewers
after operator merged PR #3077 / #3126 / #3085 at 2026-05-15T00:21Z.

**PR #3077 PB-4 lower §16 fix**:
§16 "Memory disciplines applied" bullet said "diagnostics coupled
INTO PreInferDag via biconditional" — but §4.3 (per openai-pro
DiagnosticAnchor fix commit b812db9) constrains biconditional
to PortAnchor-only. Other anchor kinds (DeclarationAnchor /
RecordFieldAnchor / SurfaceFormAnchor) couple without port-state.
Fix: §16 bullet honors §4.3 anchor-typed framing.

**PR #3126 PB-3 parse §9 Step 2 fix**:
§9 Step 2 row described diagnostics as "coupled INTO SurfaceModule"
as if live — §4.3 correctly marks it PROPOSED. Same
feedback_discipline_change_audit_all_contract_mentions pattern
that's recurred this session.
Fix: §9 Step 2 row clarified — "PROPOSED extension per §4.3";
Step 2 PR scope includes authoring the diagnostics field
extension, NOT a live coupling.

Per feedback_discipline_change_audit_all_contract_mentions: when
a substantive fix changes a discipline framing, audit ALL sections
(framing + contract + handoff). Post-merge audit surfaced these
residual contradictions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — PR #3126 codex BLOCKING (GrammarSpec parallel-authority + fail-closed weakening)

Codex REQUEST_CHANGES on already-merged PR #3126 (/api/reviews/12175):
2 substantive findings on the post-merge doc.

**Finding 1 (P2 violation — GrammarSpec parallel authority)**:

§3.2 says GrammarSpec is compile-time-only, NOT runtime-
interpreted (per Decision 3.B (b) operator override). But the
proposed stage contract still took `grammar: GrammarSpec` as
runtime input. Creates two authorities (compiled parser tables
+ runtime GrammarSpec value).

Fix: §4.3 signature reframed to `fn parse(tokens: List<Token>)
-> Result<SurfaceModule, ParseDiagnostic>` — NO runtime
GrammarSpec input. Compile-time generated parser tables consumed
via internal dispatch. Step 2 + Step 4 rows updated.

**Finding 2 (P3 + Practices 1/2 — fail-closed weakening)**:

Live parser at parse_generated.rs:138 returns `Result<SurfaceModule,
Diagnostic>` (fail-closed; aborts on first error). Earlier draft
proposed `SurfaceModule` with embedded diagnostics — would let
partial-parse states be constructible + let downstream observe
"success" output after parse failure. Violation of fail-closed
discipline.

Fix: signature preserves Result-sum (matches live + emit's
pattern). Distinguished cross-stage:
- Result-sum (parse + emit): fail-fast output domain
- Typed-state-with-coupled-diagnostics (lower + infer): structural
  output domain where partial-failure IS valid intermediate

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 post-merge APPROVE_WITH_COMMENTS — §4.1 + §16 fail-closed honesty

Cursor caught 2 more post-merge contradictions on PR #3126:

1. §4.1 line 81 "construction-time invariant" talks about
   "ParseDiagnostic in the diagnostic stream" tied to SurfaceModule
   path — but §5.2 + live parse_generated.rs:138 use
   Result<SurfaceModule, Diagnostic>. §4.1 reads as claim about
   today's plumbing.

   Fix: §4.1 reframed — live boundary explicit (Result-sum);
   construction-time invariant scoped to Ok-arm SurfaceModule +
   Err-arm ParseDiagnostic, no partial-parse with embedded
   diagnostics.

2. §16 line 397 cites C-8 as "ParseDiagnostic coupled INTO
   SurfaceModule" without qualifier — but §4.3 (post codex
   REQUEST_CHANGES fix) constrains to Result-sum.

   Fix: §16 bullet honors §4.3 Result-sum framing; cross-stage
   discriminator named.

Same recurring feedback_discipline_change_audit_all_contract_mentions
pattern — substantive fix to §4.3 + §9 Step 2 left §4.1 + §16
inconsistent. Post-merge audit catches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — TypeConnective extension stop-signal

Cursor INLINE BLOCKING caught §3.2 line 66 "rules extend
automatically" weakens substrate-extension stop-signal. Thesis
discipline: a 7th TypeConnective variant requires explicit C1
audit + named infer-rule receipt.

Fix: §3.2 reframed. New TypeConnective variants do NOT extend
automatically; require explicit C1 substrate-extension audit +
named infer-rule receipt for the new variant's structural
inference behavior.

Per-variant structural facts means new variants need new
per-variant facts, NOT silent inheritance.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — AlgebraAxis + InferDiagnostic Practice 4 classification

Cursor INLINE BLOCKING /api/reviews/12218 (line:112): proposed
substrate coproducts (AlgebraAxis + InferDiagnostic) lack
🟢/🟡/🔴 classification + ledger/trigger per modeling-discipline
Practice 4 (Coproduct dissolution).

Fix: added 🟡 SCAFFOLD classification + named dissolution
trigger for both:

AlgebraAxis 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full algebra-axiom set
  against infer.rs check sites + verifies coverage parity
  with live verification.dag:146 AlgebraicLawKind 3-variant
  subset → promote to 🟢 TERMINAL.

InferDiagnostic 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full variant set against
  parse_generated.rs / lower.rs / infer.rs diagnostic emission
  sites → promote to 🟢 TERMINAL.
- Anti-bridge per Q6.5: does NOT collapse into
  CompilerDiagnosticKind without substrate-extension
  ratification per PR #3077 §12 Q7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — §5.1 SurfaceItem allocation correction

Cursor INLINE BLOCKING /api/reviews/12251 line:130: §5.1 said
"every SurfaceItem variant maps 1:1 to a Declaration placeholder"
but live lower.rs:2950-2958 explicitly skips Let/Module/Import
in collect_symbols.

Verified via Read of lower.rs:2956-2958:
  SurfaceItem::Let { .. } => continue,
  SurfaceItem::Module { .. } => continue,
  SurfaceItem::Import { .. } => continue,

Fix: §5.1 reframed — DeclarationAllocating variants (Fn /
FnExternalBody / Data / TypeAtom / TypeRecord) map to
placeholders; NonDeclarationAllocating variants (Let / Module /
Import) skip allocation per live lower.rs behavior.

Let-bodies lower to Bind expressions in Pass 2; Module/Import
are parsed-facts preserved but un-declared.

Earlier "every SurfaceItem variant" framing overstated; would
have steered Step 2/3 worker into wrong allocation contract.
Per INVARIANTS P1/P2 live-state honesty + facts-flow-forward.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — ElaborationSpec scope broadened

Cursor INLINE BLOCKING line:33: ElaborationSpec was defined only
as Surface→Behavior recipe mapping, but lower constructs
Declarations / TypeConnectives / BranchPatterns / Bindings as
well. Non-Behavior lowering decisions outside declared authority
violates THESIS substrate ownership + INVARIANTS P2.

Fix: §3.2 ElaborationSpec scope broadened to ALL lowering
decisions:

1. SurfaceItem → Declaration recipes (Fn / Data / Type variants +
   Let/Module/Import skip-allocation per §5.1)
2. SurfaceType → TypeConnective recipes (Atom / Arrow / Compose /
   Disj construction)
3. SurfaceExpr → Behavior recipes (Value / Transform / Branch /
   Loop / Bind construction)
4. SurfacePattern → BranchPattern recipes (ResolvedVariant /
   UnresolvedVariant / record-pattern construction)
5. Binding-site rules (Bind params + result_port construction)

ElaborationSpec is single-authority across ALL axes; no axis
lives in implementation-tier hand-Rust.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — §7.1 ordering vs independence clarification

Cursor APPROVE_WITH_COMMENTS /api/reviews/12265 line 197: §7.1
mixed two claims:
- "parse migrates AFTER tokenize substrate-side stable" (ordering)
- "PB-3 parse migration is independent of PB-2 tokenize migration
  status" (independence)

Read as contradictory by reviewers. Need one coherent story.

Fix: §7.1 reframed with two distinct axes explicit:

1. Substrate-stability ordering (SELF_HOSTING.md §2 bottom-up):
   tokenize Token carrier shape must be stable BEFORE parse
   migrates. Already true at HEAD (tokenize.dag:65-67 declares
   live carrier). ✓

2. Migration-timing independence (parallel-dispatch axis): PB-3
   parse migration ships in parallel with PB-2 residual-retirement
   work (SG-1a + character-level scaffold + codegen-driver
   retirement per PB-2 L2.5 §1). What parse needs is the stable
   Token CARRIER; PB-2's migration is about retiring residual
   hand-Rust, not changing the carrier.

Both claims coherent on the axis split; not contradictory.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — ParseDiagnostic SourceSpan structural requirement

Cursor INLINE BLOCKING /api/reviews/12277 line:65: ParseDiagnostic
variants like UnexpectedToken lacked SourceSpan field;
List<ParseDiagnostic> cannot satisfy fail-closed source
attribution structurally without span on every variant.
INVARIANTS P2/P3 violation.

Fix: every ParseDiagnostic variant now carries SourceSpan
structurally:
- UnexpectedToken: added span: SourceSpan
- UnterminatedConstruct: opener_span: SourceSpan (already present)
- InvalidLiteral: added span: SourceSpan
- DuplicateRecordFieldLabel: added span: SourceSpan (current site)
  + prior_span: SourceSpan (prior site; both required)

Per INVARIANTS P2/P3 fail-closed source attribution discipline:
every diagnostic emission carries structural source-span
provenance; not optional.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Token authority cite phantom-reference fix

Cursor INLINE BLOCKING /api/reviews/12279 line:29: §3.1 cited
`src/v3/compiler/src/tokenize.rs` as "current hand-Rust" for
Token carrier — but `tokenize.rs` (without _generated suffix)
doesn't exist. Live state has tokenize.dag (live substrate) +
tokenize_generated.rs (codegen artifact).

Per design-pure-bootstrap.md PB-2 lane: tokenize retire has
substantially landed. The reference was an earlier-draft
phantom from when Token-was-hand-Rust framing was the assumption.

Fix: §3.1 reframed — Token type lives in LIVE
src/v3/std/tokenize.dag:65-67 shared taxonomy; tokenizer
implementation also live at src/v3/compiler/tokenize.dag (154
lines); codegen artifact at tokenize_generated.rs. tokenize.rs
phantom reference removed explicitly. PB-2 substantially landed
per design-pure-bootstrap.md; residual scaffold-retirement
scope per PB-2 L2.5 PR #3127.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 2 pipeline-slot target fix (compiler.dag → pipeline.dag)

Cursor INLINE BLOCKING /api/reviews/12281 line:138: §9 Step 2
referenced generic "compiler.dag" but dsl/gunbc/compiler.dag:24
explicitly directs internal pipeline (Tokenize → Parse → ...)
to src/v3/compiler/pipeline.dag, NOT generic compiler.dag.

Worker briefs authored against this doc would target the wrong
file for pipeline-slot declaration. P2 single-authority violation.

Fix: §9 Step 2 row in ALL 4 L2.5 docs (PB-2 / PB-3 / PB-4 / PB-5)
updated:
- "declared in compiler.dag" → "declared in src/v3/compiler/pipeline.dag (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag)"
- substrate column: "compiler.dag refinement" → "pipeline.dag refinement"
- §13 "Step 2 (pipeline-slot in compiler.dag)" → "pipeline-slot in src/v3/compiler/pipeline.dag"

Same phantom-citation class as the tokenize.rs phantom (commit
8ae37b4): I cited generic file path without verifying which
specific file is authoritative per project structure. Should
have grep'd dsl/gunbc/compiler.dag header notes before authoring.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 3b/4 phasing P5 receipt unambiguity

Cursor INLINE BLOCKING /api/reviews/12283 line:181: Q5 said
Step 3b "full parser body .dag migration + parser body deletion
in same PR", but §15 sequence schedules Step 4 parity/deletion
AFTER Step 3b merges. P5 dissolution receipt ambiguous.

If Step 3b lands .dag parser body BEFORE Step 4 deletes Rust
parse() body, Rust + .dag parser bodies coexist temporarily —
paper-shrink-relocation risk per feedback_paper_shrink_variants.

Fix:
1. §9 Step 3b row reframed as "Step 3b/4 COMBINED" — atomic
   single PR (full .dag parser body + parity TestClaim +
   parse_generated.rs:138 deletion + census shrink). Cannot land
   .dag parser body before Rust deletion.
2. §12 Q5 phasing clarified:
   - Phase 3a (separate PR): grammar table extension; P5 receipt
     = ROADMAP deferral row naming Step 3b/4 as future-receipt
   - Phase 3b/4 COMBINED (single PR): atomic substrate substitution
3. §9 + §15 update notes: sequence collapses steps 12-17 into
   single dispatch+merge for combined phase

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 high-level BLOCKING findings 1 + 2

Codex high-level BLOCKING (sha 140eb6b) had 4 findings:
- Findings 3 + 4 already addressed in PR #3138 commits 0159773
  + a1607a8
- Findings 1 + 2 addressed in this commit

**Finding 1 (Diagnostic kind/record conflated)**:
§4.2 ParseDiagnostic was modeled as variants-directly with span
+ kind-fields mixed. Live Diagnostic at diagnostics.dag:150 uses
record-wraps-kind pattern (`{ kind, span, message, correction }`).
Need consistent shape.

Fix: refactored to record-wraps-kind:
- `type ParseDiagnostic { kind: ParseDiagnosticKind, span: SourceSpan }`
- `type ParseDiagnosticKind = UnexpectedToken | UnterminatedConstruct | InvalidLiteral | DuplicateRecordFieldLabel | ...`

Span lives on ParseDiagnostic record (single source of truth);
variant-specific spans (opener_span / prior_span) remain on kind
variants where meaningful.

**Finding 2 (PB-2 §6 obsolete sibling-lane assumption)**:
§6 line 198 said "PB-2 Tokenize | src/v3/std/tokenize.dag (NEW
per PB-2 L2.5)" — but tokenize.dag is ALREADY LIVE per
design-pure-bootstrap.md §"PB-2 — tokenize retire" (substantially
landed).

Fix: §6 prereq table row updated to reflect LIVE state. PB-2's
residual scope is scaffold-retirement (SG-1a + character-level +
codegen-driver per PB-2 L2.5 §1), not carrier authoring. PB-3
consumes the live Token carrier; carrier shape stable across
PB-2 residual-retirement timing.

Findings 3 + 4 already addressed:
- Finding 3 (pipeline.dag target): commit 0159773
- Finding 4 (Step 3b/4 atomic vs sequential): commit a1607a8

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — phantom "module-level metadata" removed

Cursor APPROVE_WITH_COMMENTS line:79: §4.1 said "SurfaceModule
(verified live) carries List<SurfaceItem> + module-level
metadata" — but live parse_surface.dag:29 has ONLY
`{ items: List<SurfaceItem> }`. No metadata fields.

Phantom addition violated INVARIANTS P1 live-state honesty.

Fix: §4.1 reframed to match live carrier exactly. Same
phantom-addition class as tokenize.rs phantom (commit 8ae37b4).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3085 high-level BLOCKING findings — AtomPayload 5-variant + diagnostic-table PROPOSED

Codex high-level BLOCKING (sha bf4d315) — 2 substantive findings:

**Finding 1 (AtomPayload stale)**:
§2 + §5 cited infer.rs:10 top-comment "Atom(Identifier { name,
resolved })" — but live substrate.dag:87 has 5-variant
AtomPayload sum: Literal | UnresolvedIdentifier |
ResolvedByStructure | ResolvedByName | TypeParam.

infer.rs top-comment is STALE vs live substrate. My doc inherited
the drift.

Fix: §2 enumeration corrected to all 5 AtomPayload variants per
live substrate.dag:87.

**Finding 2 (diagnostic-table PROPOSED, not live)**:
§2 + §4.1 + §4.3 said "diagnostics.contains(port_id)
biconditional" as if live — but Dag at substrate.dag:525 has
ONLY { declarations, nodes, ports, clusters }. NO diagnostics
field. Diagnostic-table is PROPOSED substrate extension.

Fix: §2 fail-closed note flagged PROPOSED — Step 2 PR scope
includes `diagnostics: Map<PortId, Diagnostic>` field extension
to Dag, OR PB-Substrate prereq adds it before PB-5 dispatch.

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline (PR #3126 SurfaceModule analogous case); needed to
grep type Dag fields BEFORE claiming the diagnostics field
exists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING line:94 — §4.1 + §4.3 PROPOSED diagnostic-table marking

Cursor INLINE BLOCKING line:94: §4.1 + §4.3 said "diagnostics
coupled INTO InferredDag" without flagging the diagnostic-table
as a substrate extension. Live Dag at substrate.dag:525 has
{ declarations, nodes, ports, clusters } — NO diagnostics field.
Earlier commit c0d96af added PROPOSED marking only in §2;
§4.1 + §4.3 needed same treatment.

Fix: §4.1 + §4.3 reframed with explicit PROPOSED substrate-
extension marking + reference to §2 for extension scope.
Construction-time invariant + structural coupling are both
contingent on the substrate-extension landing (Step 2 PR scope
or PB-Substrate prereq).

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline applied to §4.1 + §4.3 consistently with §2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3077 high-level BLOCKING — Practice 4 classifications for PB-4 coproducts

Codex high-level BLOCKING (sha b812db9): "Diagnostic substrate
shape was repaired around anchoring but not re-audited as new
substrate type declarations" — meaning DiagnosticAnchor +
LowerDiagnostic + SurfaceFormRef + IdentifierRef need
🟢/🟡/🔴 classifications per modeling-discipline Practice 4
(Coproduct dissolution).

Same pattern as PB-5 fix in commit 040681f (AlgebraAxis +
InferDiagnostic) applied here.

Fix: added classifications + dissolution triggers:

- SurfaceFormRef 🟢 TERMINAL: closed-axis sum over live Surface*
  carriers; no further dissolution.
- IdentifierRef 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates full identifier-kind set against lower.rs
  identifier-resolution sites; promote to 🟢 TERMINAL when
  SurfaceVarRef + TypePathRef + ModulePathRef cover actual axes.
- LowerDiagnostic 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates variant set against lower.rs Diagnostic emission
  sites + Q6.5 anti-bridge preserved + PR #3077 §12 Q7
  ratification path.
- DiagnosticAnchor 🟢 TERMINAL: closed-axis covering all
  lowering-stage anchor kinds; no further dissolution.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING line:161 — DiagnosticSource Practice 4 classification

Cursor inline finding adds DiagnosticSource to Practice 4
classification scope. Earlier commit aba79a1 classified
SurfaceFormRef + IdentifierRef + LowerDiagnostic + DiagnosticAnchor
but missed DiagnosticSource.

Fix: DiagnosticSource 🟢 TERMINAL at pipeline-stage
discrimination scope. Closed-axis sum (Parse | Lower | Infer |
Emit); adding new pipeline stage requires explicit substrate-
extension audit per Practice 4 + stop-signal discipline (same
shape as PB-5 §3.2 TypeConnective stop-signal).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING line:151 — §11 SurfaceItem::Let + SurfaceLiteral span correction

Cursor INLINE BLOCKING line:151: §11 claimed "every Surface
variant carries SourceSpan" but live parse_surface.dag has:
- SurfaceItem::Let { name, type_ann, expr } — no direct span
- SurfaceLiteral = Int(String) | Bool(Bool) | String(String) —
  plain-tuple variants with no direct span

Source-span provenance for these cases is via enclosing carrier:
SurfaceLiteral wraps within `Literal { value, span }` at
parse_surface.dag:150. Let-item inherits container span.

INVARIANTS P2/P3 source provenance is structurally guaranteed
via direct-OR-enclosing carrier, but my "every variant"
overstatement obscured this.

Fix: §11 corrected to "most Surface variants carry SourceSpan
directly" + explicit Let + SurfaceLiteral exceptions noted +
Step 2 PR scope audits whether exceptions are structural-honest
(enclosing-carrier-provides-span) OR require substrate extension.

Same recurring overstatement class as earlier "every SurfaceItem
maps 1:1 to Declaration" (commit 61e2b67) — need to grep live
substrate variant fields BEFORE claiming uniform shape.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3126 INLINE BLOCKING line:32 — GrammarSpec is concept, not carrier

Earlier draft (now-merged PR #3126) framed `GrammarSpec` as a second declared
`.dag` input type alongside `List<Token>`. Verified via grep that no
`type GrammarSpec` exists anywhere under src/v3/ or dsl/ — the repo carries
only `parse_tables.dag`'s 6 SG-2c table-families. Per cursor 2026-05-14T23:30:04Z
inline finding, this violates INVARIANTS P2 (the Step-2 signature names a
carrier the substrate doesn't declare).

Reframed §3 (and downstream mentions in §3.2, §3 preamble line:33, §4.2
codex-correction recap line:165, §5.1 line:181, §12 Q6 line:344) so that:

- Parse has ONE input at the API boundary: `List<Token>`.
- "GrammarSpec" is a concept-level grouping for the 6 compile-time table-families
  in `parse_tables.dag`, not a substrate carrier and not a runtime parameter.
- Step-2 signature stays `fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
  per cfe842b (already applied in #3126); these edits remove the lingering
  "two input types" framing that contradicted that signature.

Per Decision 3.B (b) compile-time parser tables: substrate authority is
parse_tables.dag (6 table-families) consumed via direct table lookups inside
the parser body — no runtime grammar value, no `GrammarSpec` carrier needed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING — §4.3 + §16 internal contradictions resolved

Codex review id 12391 on sha cd6e8d1 flagged two live statements of the
already-rejected typed-state-with-coupled-diagnostics model still sitting
inside the parse L2.5, contradicting the corrected Result-sum framing in §4.2
and §6 Step 2:

1. §4.3 still proposed a `SurfaceModule { items, diagnostics }` extension
   modeled on PB-4/PB-5 patterns. Reframed: §4.3 now states explicitly that
   parse-stage uses Result-sum (no diagnostics field on SurfaceModule),
   restates the cross-stage discriminator (Result-sum for fail-fast output
   domains: parse + emit; typed-state for structural output domains: lower +
   infer), and cites `parse_generated.rs:138` as the live shape.

2. §16 "Memory disciplines applied" bullet read
   "feedback_state_space_vs_behavioral_invariants (typed-state SurfaceModule
   at output)". Rewritten to "parse output is `Result<SurfaceModule,
   ParseDiagnostic>` — the type rules out partial-parse states by
   construction; SurfaceModule itself carries no diagnostic field per §4.3".

Per `feedback_discipline_change_audit_all_contract_mentions`: when a
contract changes (here: SurfaceModule extension dropped in favor of
Result-sum), all §-internal restatements must be swept in the same diff
or they leak through as authoritative parallel claims.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3127 BLOCKING — drop TokenizedSource extension + cite concrete P5 receipts

Codex review id 12370 on sha d15e1f2 raised two load-bearing planning-shape
findings on the tokenize L2.5. Addressed in this fix-forward branch (PR #3138)
since PR #3127 is still open but accumulating cycles.

Finding 1 — parallel boundary carriers (INVARIANTS P2 / Modeling Practices 3+5):
§4.3 + §9 Step 2 + §12 Q6 named `TokenizedSource { tokens, diagnostics }` as the
output, while §2/§7.2/§7.3 named `List<Token>`. Same misclassification I just
removed from the parse L2.5 in PR #3138 §4.3: tokenize sits in the fail-fast
output domain alongside PB-3 parse + PB-6 emit (a partial token list with a
corrupt token in the middle is not a valid downstream input for parse), so the
failure couples via `Result`, not into the structural carrier.

Resolution:
- §4.3 rewritten to ratify `Result<List<Token>, TokenizeDiagnostic>` — the live
  `tokenize_generated.rs:96` shape — with no `TokenizedSource` extension.
- §9 Step 2 row signature updated to match; explicit "single canonical boundary
  carrier: List<Token> on the Ok branch" framing.
- §12 Q6 resolved REJECTED in-doc (no operator ratification needed; disposition
  follows from the cross-stage discriminator that's also load-bearing in PR
  #3138 parse L2.5).
- §16 memory-disciplines bullets rewritten parallel to PR #3138 parse §16:
  Result-sum, no diagnostics field on `List<Token>`.
- §14 "Surfaces awaiting" trimmed Q6 from the operator-ratification list.

Finding 2 — soft deferral of `regen_tokenize` retirement (INVARIANTS P5):
deferral previously named "PB-Bootstrap-Process lane scope" without a concrete
ROADMAP.md row. Updated §9 Step 2 + Step 4 rows to cite the named receipts:
- `docs/design-pure-bootstrap-zero.md:116` (PB-Bootstrap-Process lane: author
  bootstrap.dag + generated trampoline; sized M).
- `docs/design-pure-bootstrap-zero.md:118-123` (N=0 runtime verification gates).
- ROADMAP.md:467 (Character-level under-consumption in tokenize + syntax
  authorities — phase-2 char-class retype owns the codegen-driver path).
- ROADMAP.md:416 (Class 5 Gap 3 — top-level `ValueBody` boundary; gating
  substrate-capability for the phase-2 retype).
- ROADMAP.md:53 (T-PB-A — non-test census → 0 floor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 BLOCKING line:119 — §6/§7.2/§15 Q7 sequencing resolved by PR #3077 merge

Cursor inline BLOCKING at docs/design-parse-stage-l25-model.md:119 (sha at
merge time of PR #3126) flagged a real internal contradiction:
- §6 (lines 191, 207, 308) claimed "Step 2 (pipeline-slot declaration) is
  unblocked"
- §7.2 (line 227) claimed "PR #3077 §12 Q7 must ratify before any Step 2
  worker brief authoring"

A worker reading the doc could land Step 2 (pipeline boundary) before the
diagnostic carrier's P2/P3 failure shape was fixed.

Resolution: PR #3077 (PB-4 lower L2.5) merged at 2026-05-15T00:21:19Z,
carrying the §12 Q7 ratification of the Decision 2.B per-stage diagnostic
extension path. The gate IS now satisfied at HEAD, so the resolution is
fact-update (annotate Q7 as DONE with the merge timestamp) rather than
retracting either §6 or §7.2.

Edits:
- §15 step 4: annotated "DONE 2026-05-15T00:21:19Z when PR #3077 merged"
  and added the explicit "Step 2 is now genuinely unblocked, not just
  procedurally next" framing so workers reading the sequence don't bypass
  the gate.
- §7.2 line 227: rewritten from "Q7 must ratify before Step 2 brief
  authoring" (future tense, the contradiction surface) to "Gate satisfied
  2026-05-15T00:21:19Z when PR #3077 merged; Step 2 worker brief authoring
  is unblocked at HEAD per §15 step 4." Cites the cursor finding as the
  resolution path.

§6 unblocking statements stay as-is — they were correct at HEAD; the
contradiction lived in §7.2's pre-merge framing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 16d21f4) — Q7 cited at every Step 2 unblocking claim

Codex BLOCKING finding 2 (sha 16d21f4, 216s thinking): the Q7 dependency
was recorded in §7.2/§15 but not in the other Step-2-unblocking sites at
§6 lines 191, 207, 308 — risking a worker reading "Step 2 unblocked" without
also reading the Q7 prerequisite.

Codex framing: "make Q7 a hard precondition wherever Step 2 is called
unblocked, or split Step 2 into pre-Q7 and post-Q7 scopes with separate
receipts."

Chose the first option since PR #3077 has already merged (2026-05-15T00:21:19Z)
and splitting into pre/post-Q7 scopes is no longer load-bearing. Annotated
all three §6 sites:
- Line 191 (Implication for PB-3 migration): cites gate + merge timestamp
  + explicit "must NOT be brief-authored before that merge timestamp."
- Line 207 (Critical observation): cites the Step 2 gate as PR #3077 §12 Q7
  + merge timestamp + P3 failure-shape consequence if violated.
- Line 308 (Director-recommend phase list): cites gate + §7.2/§15 step 4
  cross-refs + merge timestamp.

Codex BLOCKING finding 1 (GrammarSpec carrier non-existence) verified
already resolved at HEAD via commit c97dc15 — every GrammarSpec mention
now explicitly marks it as a concept-not-carrier; Step 2 signature is
`fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
with no GrammarSpec parameter.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3126 BLOCKING (sha 16d21f4) — Step 3b/4 atomicity + Q1 ratification-pending status

openai-pro REQUEST_CHANGES on PR #3126 sha 16d21f4 flagged three findings.
Findings 1 + 2 (signature vs realization mismatch under LAYER MODEL + P2/P3
boundary discipline) were already resolved at HEAD by earlier fix-forward
commits — Step 2 signature is now `fn parse(tokens: List<Token>) ->
Result<SurfaceModule, ParseDiagnostic>` matching live `parse_generated.rs:138`
with no GrammarSpec parameter and no diagnostics-coupled SurfaceModule
extension (§4.3 Result-sum disposition).

Finding 6 (TRACKED vs UNTRACKED DEBT — Step 3b/4 same-PR vs two-PR):
§9 had both a "Step 3b/4 COMBINED" row (line 250) and a leftover separate
"Step 4: Parity test" row (line 251) — internally contradictory. §15 also
still sequenced Steps 3b + 4 as four separate authoring/dispatch/ratify
beats (steps 12-17), contradicting §12 Q5's "same PR" decision and the
explicit "Update to §15" note at §12 line 325.

Resolution: collapsed §9 to one COMBINED row absorbing the parity-TestClaim
mechanics + P5 dissolution receipt from the deleted Step 4 row; collapsed
§15 steps 12-17 into single COMBINED authoring + dispatch + ratify (steps
12-14). Added explicit `feedback_paper_shrink_variants` reasoning in both
sections.

Finding 2.5 (PM intent — substrate-capability bundled vs separate):
§12 Q1 line 288 said "Director-recommend: (b) bundled" while §13 line 340
listed substrate-capability as a non-goal of PB-3 and §15 step 11 had
"WAIT for substrate-capability landing" — three sections, two different
execution paths.

Resolution: annotated Q1 as "PENDING operator/PM ratification" with
explicit default-execution clause: until operator ratifies, the doc
treats substrate-capability as path (a) separate lane (matching §13 + §15
+ §9 row). If/when ratified to (b), §13 drops the non-goal and §15 step 11
collapses into the COMBINED Step 3b/4 brief.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 5619afa) — parse_tables.dag as single enumerated authority

Codex caught me copying the prose summary at parse_tables.dag:23-29 (which
enumerates only SG-2c-numbered families) instead of grepping the live
`^type [A-Z]` declarations. Result: SoftKeywordIdentRow (line 334) was
missing from §3.2 / §5.1 / §6 / §12 because it lacks an SG-2c-N number
in the prose summary.

Per `feedback_parallel_representation_debt`: structural fix is to stop
hand-enumerating in the doc — cite parse_tables.dag itself as the single
enumerated authority and use `type`-declaration line-anchors for the
worked example, not a hand-maintained count.

Edits:
- §3.2 §"Live substrate authority": replaced the SG-2c-numbered bullet
  list with `type`-declaration line-anchor enumeration including
  SoftKeywordIdentRow at parse_tables.dag:334 + the supporting enum
  BinaryOpLevel at line 133. Added codex-finding callout explaining the
  miss + the discipline shift.
- §3 preamble line:33, §3.2 line:45 callout, §3.2 line:57 framing,
  §3.2 §"Substrate authority" line:71, §5.1 line:171-180, §12 Q2 line:298,
  §12 Q6 line:334: all hardcoded "6 table-families" counts dropped; doc
  now points readers to §3.2 enumeration / `parse_tables.dag` directly.
- §5.1 sub-enumeration list (the parallel 6-item list at lines 173-178)
  deleted; replaced with redirect to §3.2 + restated 3a-vs-3b/4 scope split.

Code-level check before commit: `grep -nE '^type [A-Z]' src/v3/compiler/parse_tables.dag`
returns 7 types: BinaryOpLevel (133), BinaryOpRow (167), TopLevelItemKwRow (289),
SoftKeywordIdentRow (334), BracketRow (385), PrimaryPrefixRow (449),
PrimaryAtomRow (486). Doc enumeration matches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING (sha f08b952) — bind TokenizeDiagnostic to live CharClass authority

Codex Finding 1 (sha f08b952, 216s thinking): the §4.2 TokenizeDiagnostic
draft and the §5.1/§5.2 headings used `ScannerCharClass` / `ScannerClassRef`
— names that exist only as the *generated Rust enum spelling*, not as a
declared .dag substrate type. Verified via grep:

  grep -rn '^type CharClass\|^type ScannerC' dsl/ src/v3/

returns ONE authority: `dsl/std/unicode.dag:62`
  type CharClass = Whitespace | Digit | IdentStart | IdentContinue

consumed at `src/v3/compiler/tokenize.dag:103`
  data ascii_scan_order: List<CharClass> = [Whitespace, Digit, IdentStart, IdentContinue]

There is no `ScannerCharClass` declaration anywhere — that name was copied
from generated Rust without grep-verification, the same failure mode as
`feedback_grep_substrate_before_naming_ratification` (carrier-name
collision discipline).

Resolution:
- §4.2 TokenizeDiagnostic carrier: `expected_class: ScannerClassRef` →
  `expected_class: CharClass`, dropped the `type ScannerClassRef =
  ScannerCharClass` alias entirely; added a codex-finding callout citing
  the substrate authority + naming the failure mode.
- §5.1 heading "Byte → ScannerCharClass dispatch" → "Byte → CharClass
  dispatch"; bullets unchanged; added line-anchor cites for the substrate
  authority + explicit "NOT ScannerCharClass" disclaimer.
- §5.2 heading "ScannerCharClass → token-recognition state machine" →
  "CharClass → token-recognition state machine".

Finding 2 (TokenizedSource not reconciled with parse input contract): no
new fix required — already resolved by commit 6adb992 (TokenizedSource
extension dropped entirely; tokenize uses Result<List<Token>,
TokenizeDiagnostic>; List<Token> is the single canonical boundary carrier
consumed by parse). Codex was reviewing sha f08b952, which predated the
TokenizedSource drop.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3138 INLINE BLOCKING line:188 — clarify §7.3 chain is Ok-branch propagation

Cursor INLINE at sha f08b952 worried that §7.3's cross-stage chain
"tokenize → List<Token> → parse → ..." was inconsistent with §4.3's
TokenizedSource carrier (diagnostics not flowing forward). At HEAD the
TokenizedSource extension is dropped (commit 6adb992); tokenize uses
Result<List<Token>, TokenizeDiagnostic>, so List<Token> IS the canonical
Ok-branch payload that flows forward and Err branches terminate the
pipeline fail-fast.

To make this explicit at §7.3 (instead of leaving readers to infer it
from §4.3), annotated the chain with:
- "Ok-branch propagation; Err branches are stage-terminal fail-fast per
  §4.3 Result-sum discriminator" framing prefix.
- Per-stage Result/typed-state annotations: tokenize/parse show
  Result<Ok, Err>; lower/infer show typed-state structural-output;
  emit shows Result<EmittedArtifact, EmissionDiagnostic>.
- Explicit "on any stage's Err branch the pipeline aborts at that stage
  (no partial-output propagation across boundaries)" trailer.

This makes the chain self-consistent vis-a-vis §4.3 without requiring
the reader to walk back-and-forth, and prevents future readers from
re-introducing a TokenizedSource-shaped extension to "make diagnostics
flow forward" — they already do, just via the Err branch terminating
the pipeline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3127 BLOCKING (sha d15e1f2) — character-level scaffold dissolution trigger

openai-pro REQUEST_CHANGES on sha d15e1f2: §1 line:26 character-level
under-consumption scaffold named the problem but lacked a checkable
dissolution trigger. SG-1a scaffold above had the right shape — "once
those bodies lower structurally under compile_to_dag, delete the raw-text
extractor + derive directly from lowered Dag in same PR." Character-level
scaffold just said "PB-2 Step 4 carries this scope" — a lane assignment,
not a trigger.

Resolution: rewrote §1 item 2 with the same SG-1a-shape trigger structure:
- Substrate-consumption condition (a): scanner classes / string escape /
  local punctuation retype to `dsl/std/unicode.dag` `CharClass` /
  `char_in_class` (concrete field retypes named:
  `StringEscapeSpec.suffix: Char`, `LocalPunctSpec.pattern: List<Char>`,
  `string_literal_delimiter: Char`).
- Codegen-driver condition (b): `tokenize_generated.rs` no longer emits
  hidden `byte.is_ascii_*` predicates because the driver reads class
  facts structurally from lowered `tokenize.dag`.
- Same-PR dissolution: delete the parallel character-predicate scaffold
  in the same PR that flips substrate consumption — no Rust-and-`.dag`
  coexistence per `feedback_paper_shrink_variants`.
- Cross-ref to §9 Step 4 gating prereqs: ROADMAP.md:467 + ROADMAP.md:416
  Class 5 Gap 3 + std.unicode bootstrap/load-set decision (already cited
  in §9 from earlier commit 6adb992).

Per openai-pro's framing: "small fix — mirror the SG-1a scaffold wording
by naming the exact substrate-consumption condition and same-PR deletion
receipt for the hidden Rust character predicates."

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3127 APPROVE_WITH_COMMENTS — §14 + §15 Q6-already-rejected sweep

Cursor review id 12405 (APPROVE_WITH_COMMENTS) caught the same
`feedback_discipline_change_audit_all_contract_mentions` failure mode
recurring: §12 Q6 was resolved REJECTED in commit 6adb992, and §14
"Surfaces awaiting" + §12 Q6 heading + §9 Step 2 row were updated, but
two §-internal contract restatements were missed:

- §14 acceptance criterion 11: "Operator/PM ratification on §12 Q1-Q6"
- §15 step 1: "Operator / PM-delegate ratifies §12 Q1-Q6"

Both contradicted §12 Q6 + §14 "Surfaces awaiting" (which already said
"Q1-Q5 only"). A worker reading §14/§15 could schedule sign-offs on Q6
after it was already resolved-rejected elsewhere.

Resolution: both sites now say "Q1–Q5" with the explicit Q6-rejected
crossref + "see §14/§15 for same scoping" pointer at the §14 criterion
so the three sections agree internally.

Cursor verdict was APPROVE_WITH_COMMENTS (substantive APPROVE — "fix the
checklist/sequence so every section agrees Q6 is closed"); the
exploratory volatile-line-anchor note is harmless and out-of-scope for
this PR.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 15, 2026
…ross 3 L2.5 docs

Codex review id (sha 887c696, 333s thinking) flagged three findings,
all `feedback_discipline_change_audit_all_contract_mentions` recurrences
where partial sweeps left §-internal contradictions.

Finding 1 — infer doc cross-stage trigger leak:
§4.2 InferDiagnostic dissolution trigger said "when Step 2 worker brief
enumerates the full variant set against `parse_generated.rs`
Diagnostic::ParseError, lower.rs Diagnostic construction sites, and
infer.rs Dag::mark_unresolved emission sites." That cross-stage trigger
surface is wrong for the infer-specific scaffold. Narrowed to infer-only:
"against `src/v3/compiler/src/infer.rs` Diagnostic construction sites +
`Dag::mark_unresolved` emission sites (infer-stage only)." Parse and lower
have their own per-stage carriers + own Q7 mapping; this doc no longer
reaches into their dissolution-trigger surface.

Finding 2 — parse + tokenize coproduct receipts:
`ParseDiagnosticKind` (parse §4.2) and `TokenizeDiagnostic` (tokenize §4.2)
sums were declared without the 🟡 SCAFFOLD / 🟢 TERMINAL classification +
dissolution trigger that the infer doc carries (per modeling-discipline
Practice 4 + `feedback_coproduct_dissolution`). Added matching scaffold
receipts to both: 🟡 SCAFFOLD at PROPOSED stage with stage-local dissolution
trigger (Step 2 brief enumerates against parse_generated.rs / tokenize_generated.rs
Diagnostic::* construction sites — stage-only, not cross-stage); promote
to 🟢 TERMINAL when full variant set lands. Both also carry the anti-bridge
note per Q6.5.

Finding 3 — tokenize Q7 reconciliation:
The parse doc has the "Q7 DONE 2026-05-15T00:21:19Z" annotation at §15 step 4
(commit f85fa1f) but the tokenize doc still framed Q7 as a pending lane
dependency at §4.2 line:113 ("Lane dependency"), §15 step 3, and §14
"Surfaces awaiting". Annotated all three:
- §4.2 "Lane dependency": Q7 DONE timestamp + "TokenizeDiagnostic per-stage
  variant authoring is the remaining lane work."
- §15 step 3: full DONE annotation matching parse doc §15 step 4 shape +
  cross-ref to the parse doc.
- §14 "Surfaces awaiting": strikethrough + DONE annotation.

PR #3127 also carries the tokenize doc; same edits will port there in a
companion commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: 8e682eef · Trigger: schedule
  • Thinking: 263s wall

BLOCKING (2)

Root Cause

  • docs/design-tokenize-stage-l25-model.md §3.1 models tokenize input as bare String → model the source identity in the tokenize boundary so SourceSpan can be constructed without fabrication.
  • docs/design-tokenize-stage-l25-model.md diagnostic section blurs live Diagnostic/CompilerDiagnosticKind with proposed per-stage TokenizeDiagnostic → cite the actual live carrier or mark the per-stage carrier as pending on PR #3077 §12 Q7.

⚠️ The prior issues are resolved, but the new model still drops source provenance at the pipeline boundary and misstates the diagnostic authority surface.


**Live state matches**: `tokenize_generated.rs:96` already returns `pub fn tokenize(source: &str, file: &str) -> Result<Vec<Token>, Diagnostic>`. The L2.5 ratifies this shape rather than replacing it.

Step 2 signature: `fn tokenize(source: String) -> Result<List<Token>, TokenizeDiagnostic>` — `List<Token>` is the Ok-branch payload, `TokenizeDiagnostic` (§4.2 substrate extension) is the Err-branch variant. This is the **single canonical boundary carrier** between PB-2 tokenize and PB-3 parse: `List<Token>` on the Ok branch, no parallel `TokenizedSource`.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BLOCKING: The proposed Step 2 boundary drops the file/source identity even though live tokenize takes source+file and Token/diagnostic SourceSpan requires file provenance, violating INVARIANTS P2 facts-flow-forward.

| Token + TokenKind taxonomy | `src/v3/std/tokenize.dag` (Token + TokenKind closed-axis sum) | PB-Substrate | LIVE at HEAD; complete |
| Tokenizer implementation | `src/v3/compiler/tokenize.dag` (scanner-class + recognition tables) | PB-Substrate | LIVE at HEAD; 154 lines |
| Codegen pipeline | `regen_tokenize` codegen-driver → `tokenize_generated.rs` | PB-Bootstrap-Process lane | LIVE at HEAD; codegen-driver retirement is PB-Bootstrap-Process scope (NOT PB-2 scope) |
| TokenizeDiagnostic substrate extension | extension of `src/v3/std/diagnostics.dag:150` per PR #3077 §12 Q7 | PB-Substrate + Director-tier per-stage authoring | Carrier LIVE; per-stage variant authoring NEW per Q7 ratification |

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BLOCKING: The prereq table says the TokenizeDiagnostic carrier is live at diagnostics.dag:150, but that line is EmissionDiagnostic receipt prose and no TokenizeDiagnostic carrier exists, so the design has not named a single diagnostic authority per INVARIANTS P2.

briansrls added a commit that referenced this pull request May 15, 2026
…r inline Q7 sweep (#3140)

* docs(r3): PB-2 tokenize pipeline-stage L2.5 domain model — DRAFT for ratification

Director-tier L2.5 model for PB-2 tokenize per operator 2026-05-14
ratification (Decision 1.A scoping = Option A).

PB-2 is the FURTHEST-ALONG pipeline stage — substrate authority
already lives in `.dag`:
- src/v3/std/tokenize.dag (Token + TokenKind taxonomy; LIVE 143
  lines)
- src/v3/compiler/tokenize.dag (tokenizer implementation; LIVE
  154 lines)
- src/v3/compiler/src/tokenize_generated.rs (AUTO-GENERATED; 362
  lines)

This is the END STATE that all other pipeline-stage migrations
target. PB-2 L2.5 is correspondingly lighter — mostly verification
+ residual hand-Rust retirement, NOT new substrate authoring.

Distinct §9 4-step framing:
- Step 3 = VERIFY substrate completeness (audit per
  feedback_paper_shrink_variants)
- Step 4 = HANDOFF/RETIRE residual hand-Rust scaffolding
  (coordinates with PB-Bootstrap-Process lane for codegen-driver
  retirement)

Captures audit dimensions explicitly:
- scanner-class definitions = declarative byte-pattern membership
- recognition tables = closed-axis enums
- state machine = structural transitions
- no V2 `pub mod tokenize` absorption check

§12 Q1: codegen-driver retirement scope — Director-recommend
PB-Bootstrap-Process handles all codegen-driver retirement
cross-cuttingly (not per-stage paper-shrink risk).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): §12.6 citation fix — tokenize is PB-2 lane scope per design-pure-bootstrap-zero.md

Per cursor PR #3085 finding: §12.6 explicitly tables only 4
pipeline-stage migrations (emit→lower→infer→parse); tokenize is
per design-pure-bootstrap-zero.md PB-2 lane. Same fix as PR #3085
commit 89fbd7a applied here.

INVARIANTS P1 — documentation must not overstate authority cites.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): preemptive fix for PR #3126 codex BLOCKING propagation — PB-2 live state honesty

Same class as codex INLINE BLOCKING #3126 finding 1 (live state
honesty for diagnostic coupling) applied preemptively to PB-2
tokenize L2.5.

PB-2 §4.3 had "diagnostics coupled INTO List<Token>" framing
which would overstate the live carrier shape (bare List<Token>
has no diagnostic field; tokenize_generated.rs:96 today returns
Result<Vec<Token>, Diagnostic>).

Fix: §4.3 reframed with PROPOSED substrate extension explicit —
new `TokenizedSource { tokens, diagnostics }` wrapper carrier as
the typed-state output. Step 2 brief includes wrapper authoring
in pipeline-slot PR scope.

Same discipline as PR #3126 commit bdff8c5.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 — fix PB-2 lane citation (design-pure-bootstrap.md not -zero.md)

Per cursor APPROVE_WITH_COMMENTS /api/reviews/12087: PB-2 lane is
defined in docs/design-pure-bootstrap.md §"PB-2 — tokenize retire"
(line ~134), NOT docs/design-pure-bootstrap-zero.md. The -zero.md
doc has Subsumed-lanes list with PB-1/PB-4/PB-5/PB-6 but NOT PB-2.

Propagated fix applies same cite-error correction as PR #3066 §1.8
discipline: cite the actual doc, not an adjacent doc with similar
name. INVARIANTS P2 single-authority-citation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix cursor PR #3127 BLOCKING contradictions + add §12 Q6

Cursor APPROVE_WITH_COMMENTS (/api/reviews/12093) caught two
substantive contradictions I introduced when adding TokenizedSource
in commit 2b9756b:

1. §4.3 vs §9 Step 2 signature mismatch — §4.3 said `-> TokenizedSource`
   but §9 Step 2 row still said `-> List<Token>`. Same
   `feedback_discipline_change_audit_all_contract_mentions` issue
   that's recurred 4x this session.

2. §4.3 referenced "§12 Q-new" but §12 only had Q1-Q5; broken anchor.

Fix:
1. §9 Step 2 row updated: signature `-> TokenizedSource` with
   wrapper carrier shape `{ tokens, diagnostics }` per §4.3
2. §4.3 anchor updated: "§12 Q6" (resolved)
3. Added §12 Q6: TokenizedSource carrier shape ratification —
   (a) wrapper record vs (b) per-Token diagnostic coupling;
   Director-recommend (a) for PB-3 SurfaceModule parallelism
4. §14 + §15 + §16 Q-list refs updated to Q1-Q6

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — Token shape + Q4 audit boundary

Codex REQUEST_CHANGES (sha b881de2) with 2 substantive findings:

1. §4.1 Token shape claim "optional lexeme: String" — wrong per
   live substrate at src/v3/std/tokenize.dag:65-67. Live Token
   has 2 fields only (kind + span); lexeme-content lives ON the
   TokenKind variants (Ident(String) / IntLit(String) / etc.).

   Fix: corrected §4.1 to reflect live carrier shape; payloads
   on TokenKind variants noted explicitly.

2. §12 Q4 substrate-completeness audit scoped only to
   tokenize_generated.rs — missed the regen_tokenize codegen-
   driver boundary. If regen_tokenize carries scanner-logic
   decisions (rather than mechanical template-rendering of
   substrate facts), the substrate isn't complete — the driver
   IS hand-Rust scanner logic in disguise.

   Fix: Q4 audit extended with (d) regen_tokenize codegen-
   driver logic audit + (e) ROADMAP.md deferral row option per
   feedback_paper_shrink_variants P5 receipt discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — TokenizeDiagnostic PROPOSED + scaffold disclosure

Codex REQUEST_CHANGES (sha b881de2, full body) caught 2 substantive
overstatements:

1. §2 line 46 claimed "tokenize failures produce typed
   TokenizeDiagnostic variants" — but live tokenize_generated.rs:96
   returns generic `Result<Vec<Token>, Diagnostic>` with
   `Diagnostic::TokenizerError { message, span, correction }`. Typed
   TokenizeDiagnostic is a PROPOSED extension, not live state.

   Fix: §2 reframed with live state explicit + TokenizeDiagnostic
   marked PROPOSED per PR #3077 §12 Q7 ratification path.

2. §1 line 22 + §6 + §9 Step 2 line 198 framed PB-2 as "mostly
   verification" — but live tokenize.dag:16-30+ has TWO explicit
   tracked scaffold zones:
   - SG-1a: regen_tokenize parses raw source text for
     dag_keyword_set / dag_operators (ValueBody::Unparsed)
   - Character-level under-consumption: StringEscapeSpec /
     LocalPunctSpec.pattern / string_literal_delimiter as opaque
     Strings; hidden Rust character predicates (byte.is_ascii_digit
     etc.) at tokenize_generated.rs:15-22 leaking through codegen

   Residual hand-Rust is NOT just the codegen artifact — it includes
   (a) regen_tokenize logic, (b) SG-1a raw-text-extractor scaffold,
   (c) character-predicate scaffold leaking through codegen.

   Fix: §1 + §6 + §9 Step 2 reframed honestly. PB-2 is "FURTHER
   ALONG but not complete"; Step 4 carries scaffold-retirement
   scope, not just codegen-artifact retirement.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — §16 + §9 Step 2 internal contradictions

Two post-merge doc-internal contradictions caught by reviewers
after operator merged PR #3077 / #3126 / #3085 at 2026-05-15T00:21Z.

**PR #3077 PB-4 lower §16 fix**:
§16 "Memory disciplines applied" bullet said "diagnostics coupled
INTO PreInferDag via biconditional" — but §4.3 (per openai-pro
DiagnosticAnchor fix commit b812db9) constrains biconditional
to PortAnchor-only. Other anchor kinds (DeclarationAnchor /
RecordFieldAnchor / SurfaceFormAnchor) couple without port-state.
Fix: §16 bullet honors §4.3 anchor-typed framing.

**PR #3126 PB-3 parse §9 Step 2 fix**:
§9 Step 2 row described diagnostics as "coupled INTO SurfaceModule"
as if live — §4.3 correctly marks it PROPOSED. Same
feedback_discipline_change_audit_all_contract_mentions pattern
that's recurred this session.
Fix: §9 Step 2 row clarified — "PROPOSED extension per §4.3";
Step 2 PR scope includes authoring the diagnostics field
extension, NOT a live coupling.

Per feedback_discipline_change_audit_all_contract_mentions: when
a substantive fix changes a discipline framing, audit ALL sections
(framing + contract + handoff). Post-merge audit surfaced these
residual contradictions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — PR #3126 codex BLOCKING (GrammarSpec parallel-authority + fail-closed weakening)

Codex REQUEST_CHANGES on already-merged PR #3126 (/api/reviews/12175):
2 substantive findings on the post-merge doc.

**Finding 1 (P2 violation — GrammarSpec parallel authority)**:

§3.2 says GrammarSpec is compile-time-only, NOT runtime-
interpreted (per Decision 3.B (b) operator override). But the
proposed stage contract still took `grammar: GrammarSpec` as
runtime input. Creates two authorities (compiled parser tables
+ runtime GrammarSpec value).

Fix: §4.3 signature reframed to `fn parse(tokens: List<Token>)
-> Result<SurfaceModule, ParseDiagnostic>` — NO runtime
GrammarSpec input. Compile-time generated parser tables consumed
via internal dispatch. Step 2 + Step 4 rows updated.

**Finding 2 (P3 + Practices 1/2 — fail-closed weakening)**:

Live parser at parse_generated.rs:138 returns `Result<SurfaceModule,
Diagnostic>` (fail-closed; aborts on first error). Earlier draft
proposed `SurfaceModule` with embedded diagnostics — would let
partial-parse states be constructible + let downstream observe
"success" output after parse failure. Violation of fail-closed
discipline.

Fix: signature preserves Result-sum (matches live + emit's
pattern). Distinguished cross-stage:
- Result-sum (parse + emit): fail-fast output domain
- Typed-state-with-coupled-diagnostics (lower + infer): structural
  output domain where partial-failure IS valid intermediate

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 post-merge APPROVE_WITH_COMMENTS — §4.1 + §16 fail-closed honesty

Cursor caught 2 more post-merge contradictions on PR #3126:

1. §4.1 line 81 "construction-time invariant" talks about
   "ParseDiagnostic in the diagnostic stream" tied to SurfaceModule
   path — but §5.2 + live parse_generated.rs:138 use
   Result<SurfaceModule, Diagnostic>. §4.1 reads as claim about
   today's plumbing.

   Fix: §4.1 reframed — live boundary explicit (Result-sum);
   construction-time invariant scoped to Ok-arm SurfaceModule +
   Err-arm ParseDiagnostic, no partial-parse with embedded
   diagnostics.

2. §16 line 397 cites C-8 as "ParseDiagnostic coupled INTO
   SurfaceModule" without qualifier — but §4.3 (post codex
   REQUEST_CHANGES fix) constrains to Result-sum.

   Fix: §16 bullet honors §4.3 Result-sum framing; cross-stage
   discriminator named.

Same recurring feedback_discipline_change_audit_all_contract_mentions
pattern — substantive fix to §4.3 + §9 Step 2 left §4.1 + §16
inconsistent. Post-merge audit catches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — TypeConnective extension stop-signal

Cursor INLINE BLOCKING caught §3.2 line 66 "rules extend
automatically" weakens substrate-extension stop-signal. Thesis
discipline: a 7th TypeConnective variant requires explicit C1
audit + named infer-rule receipt.

Fix: §3.2 reframed. New TypeConnective variants do NOT extend
automatically; require explicit C1 substrate-extension audit +
named infer-rule receipt for the new variant's structural
inference behavior.

Per-variant structural facts means new variants need new
per-variant facts, NOT silent inheritance.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — AlgebraAxis + InferDiagnostic Practice 4 classification

Cursor INLINE BLOCKING /api/reviews/12218 (line:112): proposed
substrate coproducts (AlgebraAxis + InferDiagnostic) lack
🟢/🟡/🔴 classification + ledger/trigger per modeling-discipline
Practice 4 (Coproduct dissolution).

Fix: added 🟡 SCAFFOLD classification + named dissolution
trigger for both:

AlgebraAxis 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full algebra-axiom set
  against infer.rs check sites + verifies coverage parity
  with live verification.dag:146 AlgebraicLawKind 3-variant
  subset → promote to 🟢 TERMINAL.

InferDiagnostic 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full variant set against
  parse_generated.rs / lower.rs / infer.rs diagnostic emission
  sites → promote to 🟢 TERMINAL.
- Anti-bridge per Q6.5: does NOT collapse into
  CompilerDiagnosticKind without substrate-extension
  ratification per PR #3077 §12 Q7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — §5.1 SurfaceItem allocation correction

Cursor INLINE BLOCKING /api/reviews/12251 line:130: §5.1 said
"every SurfaceItem variant maps 1:1 to a Declaration placeholder"
but live lower.rs:2950-2958 explicitly skips Let/Module/Import
in collect_symbols.

Verified via Read of lower.rs:2956-2958:
  SurfaceItem::Let { .. } => continue,
  SurfaceItem::Module { .. } => continue,
  SurfaceItem::Import { .. } => continue,

Fix: §5.1 reframed — DeclarationAllocating variants (Fn /
FnExternalBody / Data / TypeAtom / TypeRecord) map to
placeholders; NonDeclarationAllocating variants (Let / Module /
Import) skip allocation per live lower.rs behavior.

Let-bodies lower to Bind expressions in Pass 2; Module/Import
are parsed-facts preserved but un-declared.

Earlier "every SurfaceItem variant" framing overstated; would
have steered Step 2/3 worker into wrong allocation contract.
Per INVARIANTS P1/P2 live-state honesty + facts-flow-forward.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — ElaborationSpec scope broadened

Cursor INLINE BLOCKING line:33: ElaborationSpec was defined only
as Surface→Behavior recipe mapping, but lower constructs
Declarations / TypeConnectives / BranchPatterns / Bindings as
well. Non-Behavior lowering decisions outside declared authority
violates THESIS substrate ownership + INVARIANTS P2.

Fix: §3.2 ElaborationSpec scope broadened to ALL lowering
decisions:

1. SurfaceItem → Declaration recipes (Fn / Data / Type variants +
   Let/Module/Import skip-allocation per §5.1)
2. SurfaceType → TypeConnective recipes (Atom / Arrow / Compose /
   Disj construction)
3. SurfaceExpr → Behavior recipes (Value / Transform / Branch /
   Loop / Bind construction)
4. SurfacePattern → BranchPattern recipes (ResolvedVariant /
   UnresolvedVariant / record-pattern construction)
5. Binding-site rules (Bind params + result_port construction)

ElaborationSpec is single-authority across ALL axes; no axis
lives in implementation-tier hand-Rust.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — §7.1 ordering vs independence clarification

Cursor APPROVE_WITH_COMMENTS /api/reviews/12265 line 197: §7.1
mixed two claims:
- "parse migrates AFTER tokenize substrate-side stable" (ordering)
- "PB-3 parse migration is independent of PB-2 tokenize migration
  status" (independence)

Read as contradictory by reviewers. Need one coherent story.

Fix: §7.1 reframed with two distinct axes explicit:

1. Substrate-stability ordering (SELF_HOSTING.md §2 bottom-up):
   tokenize Token carrier shape must be stable BEFORE parse
   migrates. Already true at HEAD (tokenize.dag:65-67 declares
   live carrier). ✓

2. Migration-timing independence (parallel-dispatch axis): PB-3
   parse migration ships in parallel with PB-2 residual-retirement
   work (SG-1a + character-level scaffold + codegen-driver
   retirement per PB-2 L2.5 §1). What parse needs is the stable
   Token CARRIER; PB-2's migration is about retiring residual
   hand-Rust, not changing the carrier.

Both claims coherent on the axis split; not contradictory.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — ParseDiagnostic SourceSpan structural requirement

Cursor INLINE BLOCKING /api/reviews/12277 line:65: ParseDiagnostic
variants like UnexpectedToken lacked SourceSpan field;
List<ParseDiagnostic> cannot satisfy fail-closed source
attribution structurally without span on every variant.
INVARIANTS P2/P3 violation.

Fix: every ParseDiagnostic variant now carries SourceSpan
structurally:
- UnexpectedToken: added span: SourceSpan
- UnterminatedConstruct: opener_span: SourceSpan (already present)
- InvalidLiteral: added span: SourceSpan
- DuplicateRecordFieldLabel: added span: SourceSpan (current site)
  + prior_span: SourceSpan (prior site; both required)

Per INVARIANTS P2/P3 fail-closed source attribution discipline:
every diagnostic emission carries structural source-span
provenance; not optional.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Token authority cite phantom-reference fix

Cursor INLINE BLOCKING /api/reviews/12279 line:29: §3.1 cited
`src/v3/compiler/src/tokenize.rs` as "current hand-Rust" for
Token carrier — but `tokenize.rs` (without _generated suffix)
doesn't exist. Live state has tokenize.dag (live substrate) +
tokenize_generated.rs (codegen artifact).

Per design-pure-bootstrap.md PB-2 lane: tokenize retire has
substantially landed. The reference was an earlier-draft
phantom from when Token-was-hand-Rust framing was the assumption.

Fix: §3.1 reframed — Token type lives in LIVE
src/v3/std/tokenize.dag:65-67 shared taxonomy; tokenizer
implementation also live at src/v3/compiler/tokenize.dag (154
lines); codegen artifact at tokenize_generated.rs. tokenize.rs
phantom reference removed explicitly. PB-2 substantially landed
per design-pure-bootstrap.md; residual scaffold-retirement
scope per PB-2 L2.5 PR #3127.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 2 pipeline-slot target fix (compiler.dag → pipeline.dag)

Cursor INLINE BLOCKING /api/reviews/12281 line:138: §9 Step 2
referenced generic "compiler.dag" but dsl/gunbc/compiler.dag:24
explicitly directs internal pipeline (Tokenize → Parse → ...)
to src/v3/compiler/pipeline.dag, NOT generic compiler.dag.

Worker briefs authored against this doc would target the wrong
file for pipeline-slot declaration. P2 single-authority violation.

Fix: §9 Step 2 row in ALL 4 L2.5 docs (PB-2 / PB-3 / PB-4 / PB-5)
updated:
- "declared in compiler.dag" → "declared in src/v3/compiler/pipeline.dag (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag)"
- substrate column: "compiler.dag refinement" → "pipeline.dag refinement"
- §13 "Step 2 (pipeline-slot in compiler.dag)" → "pipeline-slot in src/v3/compiler/pipeline.dag"

Same phantom-citation class as the tokenize.rs phantom (commit
8ae37b4): I cited generic file path without verifying which
specific file is authoritative per project structure. Should
have grep'd dsl/gunbc/compiler.dag header notes before authoring.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 3b/4 phasing P5 receipt unambiguity

Cursor INLINE BLOCKING /api/reviews/12283 line:181: Q5 said
Step 3b "full parser body .dag migration + parser body deletion
in same PR", but §15 sequence schedules Step 4 parity/deletion
AFTER Step 3b merges. P5 dissolution receipt ambiguous.

If Step 3b lands .dag parser body BEFORE Step 4 deletes Rust
parse() body, Rust + .dag parser bodies coexist temporarily —
paper-shrink-relocation risk per feedback_paper_shrink_variants.

Fix:
1. §9 Step 3b row reframed as "Step 3b/4 COMBINED" — atomic
   single PR (full .dag parser body + parity TestClaim +
   parse_generated.rs:138 deletion + census shrink). Cannot land
   .dag parser body before Rust deletion.
2. §12 Q5 phasing clarified:
   - Phase 3a (separate PR): grammar table extension; P5 receipt
     = ROADMAP deferral row naming Step 3b/4 as future-receipt
   - Phase 3b/4 COMBINED (single PR): atomic substrate substitution
3. §9 + §15 update notes: sequence collapses steps 12-17 into
   single dispatch+merge for combined phase

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 high-level BLOCKING findings 1 + 2

Codex high-level BLOCKING (sha 140eb6b) had 4 findings:
- Findings 3 + 4 already addressed in PR #3138 commits 0159773
  + a1607a8
- Findings 1 + 2 addressed in this commit

**Finding 1 (Diagnostic kind/record conflated)**:
§4.2 ParseDiagnostic was modeled as variants-directly with span
+ kind-fields mixed. Live Diagnostic at diagnostics.dag:150 uses
record-wraps-kind pattern (`{ kind, span, message, correction }`).
Need consistent shape.

Fix: refactored to record-wraps-kind:
- `type ParseDiagnostic { kind: ParseDiagnosticKind, span: SourceSpan }`
- `type ParseDiagnosticKind = UnexpectedToken | UnterminatedConstruct | InvalidLiteral | DuplicateRecordFieldLabel | ...`

Span lives on ParseDiagnostic record (single source of truth);
variant-specific spans (opener_span / prior_span) remain on kind
variants where meaningful.

**Finding 2 (PB-2 §6 obsolete sibling-lane assumption)**:
§6 line 198 said "PB-2 Tokenize | src/v3/std/tokenize.dag (NEW
per PB-2 L2.5)" — but tokenize.dag is ALREADY LIVE per
design-pure-bootstrap.md §"PB-2 — tokenize retire" (substantially
landed).

Fix: §6 prereq table row updated to reflect LIVE state. PB-2's
residual scope is scaffold-retirement (SG-1a + character-level +
codegen-driver per PB-2 L2.5 §1), not carrier authoring. PB-3
consumes the live Token carrier; carrier shape stable across
PB-2 residual-retirement timing.

Findings 3 + 4 already addressed:
- Finding 3 (pipeline.dag target): commit 0159773
- Finding 4 (Step 3b/4 atomic vs sequential): commit a1607a8

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — phantom "module-level metadata" removed

Cursor APPROVE_WITH_COMMENTS line:79: §4.1 said "SurfaceModule
(verified live) carries List<SurfaceItem> + module-level
metadata" — but live parse_surface.dag:29 has ONLY
`{ items: List<SurfaceItem> }`. No metadata fields.

Phantom addition violated INVARIANTS P1 live-state honesty.

Fix: §4.1 reframed to match live carrier exactly. Same
phantom-addition class as tokenize.rs phantom (commit 8ae37b4).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3085 high-level BLOCKING findings — AtomPayload 5-variant + diagnostic-table PROPOSED

Codex high-level BLOCKING (sha bf4d315) — 2 substantive findings:

**Finding 1 (AtomPayload stale)**:
§2 + §5 cited infer.rs:10 top-comment "Atom(Identifier { name,
resolved })" — but live substrate.dag:87 has 5-variant
AtomPayload sum: Literal | UnresolvedIdentifier |
ResolvedByStructure | ResolvedByName | TypeParam.

infer.rs top-comment is STALE vs live substrate. My doc inherited
the drift.

Fix: §2 enumeration corrected to all 5 AtomPayload variants per
live substrate.dag:87.

**Finding 2 (diagnostic-table PROPOSED, not live)**:
§2 + §4.1 + §4.3 said "diagnostics.contains(port_id)
biconditional" as if live — but Dag at substrate.dag:525 has
ONLY { declarations, nodes, ports, clusters }. NO diagnostics
field. Diagnostic-table is PROPOSED substrate extension.

Fix: §2 fail-closed note flagged PROPOSED — Step 2 PR scope
includes `diagnostics: Map<PortId, Diagnostic>` field extension
to Dag, OR PB-Substrate prereq adds it before PB-5 dispatch.

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline (PR #3126 SurfaceModule analogous case); needed to
grep type Dag fields BEFORE claiming the diagnostics field
exists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING line:94 — §4.1 + §4.3 PROPOSED diagnostic-table marking

Cursor INLINE BLOCKING line:94: §4.1 + §4.3 said "diagnostics
coupled INTO InferredDag" without flagging the diagnostic-table
as a substrate extension. Live Dag at substrate.dag:525 has
{ declarations, nodes, ports, clusters } — NO diagnostics field.
Earlier commit c0d96af added PROPOSED marking only in §2;
§4.1 + §4.3 needed same treatment.

Fix: §4.1 + §4.3 reframed with explicit PROPOSED substrate-
extension marking + reference to §2 for extension scope.
Construction-time invariant + structural coupling are both
contingent on the substrate-extension landing (Step 2 PR scope
or PB-Substrate prereq).

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline applied to §4.1 + §4.3 consistently with §2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3077 high-level BLOCKING — Practice 4 classifications for PB-4 coproducts

Codex high-level BLOCKING (sha b812db9): "Diagnostic substrate
shape was repaired around anchoring but not re-audited as new
substrate type declarations" — meaning DiagnosticAnchor +
LowerDiagnostic + SurfaceFormRef + IdentifierRef need
🟢/🟡/🔴 classifications per modeling-discipline Practice 4
(Coproduct dissolution).

Same pattern as PB-5 fix in commit 040681f (AlgebraAxis +
InferDiagnostic) applied here.

Fix: added classifications + dissolution triggers:

- SurfaceFormRef 🟢 TERMINAL: closed-axis sum over live Surface*
  carriers; no further dissolution.
- IdentifierRef 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates full identifier-kind set against lower.rs
  identifier-resolution sites; promote to 🟢 TERMINAL when
  SurfaceVarRef + TypePathRef + ModulePathRef cover actual axes.
- LowerDiagnostic 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates variant set against lower.rs Diagnostic emission
  sites + Q6.5 anti-bridge preserved + PR #3077 §12 Q7
  ratification path.
- DiagnosticAnchor 🟢 TERMINAL: closed-axis covering all
  lowering-stage anchor kinds; no further dissolution.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING line:161 — DiagnosticSource Practice 4 classification

Cursor inline finding adds DiagnosticSource to Practice 4
classification scope. Earlier commit aba79a1 classified
SurfaceFormRef + IdentifierRef + LowerDiagnostic + DiagnosticAnchor
but missed DiagnosticSource.

Fix: DiagnosticSource 🟢 TERMINAL at pipeline-stage
discrimination scope. Closed-axis sum (Parse | Lower | Infer |
Emit); adding new pipeline stage requires explicit substrate-
extension audit per Practice 4 + stop-signal discipline (same
shape as PB-5 §3.2 TypeConnective stop-signal).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING line:151 — §11 SurfaceItem::Let + SurfaceLiteral span correction

Cursor INLINE BLOCKING line:151: §11 claimed "every Surface
variant carries SourceSpan" but live parse_surface.dag has:
- SurfaceItem::Let { name, type_ann, expr } — no direct span
- SurfaceLiteral = Int(String) | Bool(Bool) | String(String) —
  plain-tuple variants with no direct span

Source-span provenance for these cases is via enclosing carrier:
SurfaceLiteral wraps within `Literal { value, span }` at
parse_surface.dag:150. Let-item inherits container span.

INVARIANTS P2/P3 source provenance is structurally guaranteed
via direct-OR-enclosing carrier, but my "every variant"
overstatement obscured this.

Fix: §11 corrected to "most Surface variants carry SourceSpan
directly" + explicit Let + SurfaceLiteral exceptions noted +
Step 2 PR scope audits whether exceptions are structural-honest
(enclosing-carrier-provides-span) OR require substrate extension.

Same recurring overstatement class as earlier "every SurfaceItem
maps 1:1 to Declaration" (commit 61e2b67) — need to grep live
substrate variant fields BEFORE claiming uniform shape.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3126 INLINE BLOCKING line:32 — GrammarSpec is concept, not carrier

Earlier draft (now-merged PR #3126) framed `GrammarSpec` as a second declared
`.dag` input type alongside `List<Token>`. Verified via grep that no
`type GrammarSpec` exists anywhere under src/v3/ or dsl/ — the repo carries
only `parse_tables.dag`'s 6 SG-2c table-families. Per cursor 2026-05-14T23:30:04Z
inline finding, this violates INVARIANTS P2 (the Step-2 signature names a
carrier the substrate doesn't declare).

Reframed §3 (and downstream mentions in §3.2, §3 preamble line:33, §4.2
codex-correction recap line:165, §5.1 line:181, §12 Q6 line:344) so that:

- Parse has ONE input at the API boundary: `List<Token>`.
- "GrammarSpec" is a concept-level grouping for the 6 compile-time table-families
  in `parse_tables.dag`, not a substrate carrier and not a runtime parameter.
- Step-2 signature stays `fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
  per cfe842b (already applied in #3126); these edits remove the lingering
  "two input types" framing that contradicted that signature.

Per Decision 3.B (b) compile-time parser tables: substrate authority is
parse_tables.dag (6 table-families) consumed via direct table lookups inside
the parser body — no runtime grammar value, no `GrammarSpec` carrier needed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING — §4.3 + §16 internal contradictions resolved

Codex review id 12391 on sha cd6e8d1 flagged two live statements of the
already-rejected typed-state-with-coupled-diagnostics model still sitting
inside the parse L2.5, contradicting the corrected Result-sum framing in §4.2
and §6 Step 2:

1. §4.3 still proposed a `SurfaceModule { items, diagnostics }` extension
   modeled on PB-4/PB-5 patterns. Reframed: §4.3 now states explicitly that
   parse-stage uses Result-sum (no diagnostics field on SurfaceModule),
   restates the cross-stage discriminator (Result-sum for fail-fast output
   domains: parse + emit; typed-state for structural output domains: lower +
   infer), and cites `parse_generated.rs:138` as the live shape.

2. §16 "Memory disciplines applied" bullet read
   "feedback_state_space_vs_behavioral_invariants (typed-state SurfaceModule
   at output)". Rewritten to "parse output is `Result<SurfaceModule,
   ParseDiagnostic>` — the type rules out partial-parse states by
   construction; SurfaceModule itself carries no diagnostic field per §4.3".

Per `feedback_discipline_change_audit_all_contract_mentions`: when a
contract changes (here: SurfaceModule extension dropped in favor of
Result-sum), all §-internal restatements must be swept in the same diff
or they leak through as authoritative parallel claims.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3127 BLOCKING — drop TokenizedSource extension + cite concrete P5 receipts

Codex review id 12370 on sha d15e1f2 raised two load-bearing planning-shape
findings on the tokenize L2.5. Addressed in this fix-forward branch (PR #3138)
since PR #3127 is still open but accumulating cycles.

Finding 1 — parallel boundary carriers (INVARIANTS P2 / Modeling Practices 3+5):
§4.3 + §9 Step 2 + §12 Q6 named `TokenizedSource { tokens, diagnostics }` as the
output, while §2/§7.2/§7.3 named `List<Token>`. Same misclassification I just
removed from the parse L2.5 in PR #3138 §4.3: tokenize sits in the fail-fast
output domain alongside PB-3 parse + PB-6 emit (a partial token list with a
corrupt token in the middle is not a valid downstream input for parse), so the
failure couples via `Result`, not into the structural carrier.

Resolution:
- §4.3 rewritten to ratify `Result<List<Token>, TokenizeDiagnostic>` — the live
  `tokenize_generated.rs:96` shape — with no `TokenizedSource` extension.
- §9 Step 2 row signature updated to match; explicit "single canonical boundary
  carrier: List<Token> on the Ok branch" framing.
- §12 Q6 resolved REJECTED in-doc (no operator ratification needed; disposition
  follows from the cross-stage discriminator that's also load-bearing in PR
  #3138 parse L2.5).
- §16 memory-disciplines bullets rewritten parallel to PR #3138 parse §16:
  Result-sum, no diagnostics field on `List<Token>`.
- §14 "Surfaces awaiting" trimmed Q6 from the operator-ratification list.

Finding 2 — soft deferral of `regen_tokenize` retirement (INVARIANTS P5):
deferral previously named "PB-Bootstrap-Process lane scope" without a concrete
ROADMAP.md row. Updated §9 Step 2 + Step 4 rows to cite the named receipts:
- `docs/design-pure-bootstrap-zero.md:116` (PB-Bootstrap-Process lane: author
  bootstrap.dag + generated trampoline; sized M).
- `docs/design-pure-bootstrap-zero.md:118-123` (N=0 runtime verification gates).
- ROADMAP.md:467 (Character-level under-consumption in tokenize + syntax
  authorities — phase-2 char-class retype owns the codegen-driver path).
- ROADMAP.md:416 (Class 5 Gap 3 — top-level `ValueBody` boundary; gating
  substrate-capability for the phase-2 retype).
- ROADMAP.md:53 (T-PB-A — non-test census → 0 floor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 BLOCKING line:119 — §6/§7.2/§15 Q7 sequencing resolved by PR #3077 merge

Cursor inline BLOCKING at docs/design-parse-stage-l25-model.md:119 (sha at
merge time of PR #3126) flagged a real internal contradiction:
- §6 (lines 191, 207, 308) claimed "Step 2 (pipeline-slot declaration) is
  unblocked"
- §7.2 (line 227) claimed "PR #3077 §12 Q7 must ratify before any Step 2
  worker brief authoring"

A worker reading the doc could land Step 2 (pipeline boundary) before the
diagnostic carrier's P2/P3 failure shape was fixed.

Resolution: PR #3077 (PB-4 lower L2.5) merged at 2026-05-15T00:21:19Z,
carrying the §12 Q7 ratification of the Decision 2.B per-stage diagnostic
extension path. The gate IS now satisfied at HEAD, so the resolution is
fact-update (annotate Q7 as DONE with the merge timestamp) rather than
retracting either §6 or §7.2.

Edits:
- §15 step 4: annotated "DONE 2026-05-15T00:21:19Z when PR #3077 merged"
  and added the explicit "Step 2 is now genuinely unblocked, not just
  procedurally next" framing so workers reading the sequence don't bypass
  the gate.
- §7.2 line 227: rewritten from "Q7 must ratify before Step 2 brief
  authoring" (future tense, the contradiction surface) to "Gate satisfied
  2026-05-15T00:21:19Z when PR #3077 merged; Step 2 worker brief authoring
  is unblocked at HEAD per §15 step 4." Cites the cursor finding as the
  resolution path.

§6 unblocking statements stay as-is — they were correct at HEAD; the
contradiction lived in §7.2's pre-merge framing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 16d21f4) — Q7 cited at every Step 2 unblocking claim

Codex BLOCKING finding 2 (sha 16d21f4, 216s thinking): the Q7 dependency
was recorded in §7.2/§15 but not in the other Step-2-unblocking sites at
§6 lines 191, 207, 308 — risking a worker reading "Step 2 unblocked" without
also reading the Q7 prerequisite.

Codex framing: "make Q7 a hard precondition wherever Step 2 is called
unblocked, or split Step 2 into pre-Q7 and post-Q7 scopes with separate
receipts."

Chose the first option since PR #3077 has already merged (2026-05-15T00:21:19Z)
and splitting into pre/post-Q7 scopes is no longer load-bearing. Annotated
all three §6 sites:
- Line 191 (Implication for PB-3 migration): cites gate + merge timestamp
  + explicit "must NOT be brief-authored before that merge timestamp."
- Line 207 (Critical observation): cites the Step 2 gate as PR #3077 §12 Q7
  + merge timestamp + P3 failure-shape consequence if violated.
- Line 308 (Director-recommend phase list): cites gate + §7.2/§15 step 4
  cross-refs + merge timestamp.

Codex BLOCKING finding 1 (GrammarSpec carrier non-existence) verified
already resolved at HEAD via commit c97dc15 — every GrammarSpec mention
now explicitly marks it as a concept-not-carrier; Step 2 signature is
`fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
with no GrammarSpec parameter.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3126 BLOCKING (sha 16d21f4) — Step 3b/4 atomicity + Q1 ratification-pending status

openai-pro REQUEST_CHANGES on PR #3126 sha 16d21f4 flagged three findings.
Findings 1 + 2 (signature vs realization mismatch under LAYER MODEL + P2/P3
boundary discipline) were already resolved at HEAD by earlier fix-forward
commits — Step 2 signature is now `fn parse(tokens: List<Token>) ->
Result<SurfaceModule, ParseDiagnostic>` matching live `parse_generated.rs:138`
with no GrammarSpec parameter and no diagnostics-coupled SurfaceModule
extension (§4.3 Result-sum disposition).

Finding 6 (TRACKED vs UNTRACKED DEBT — Step 3b/4 same-PR vs two-PR):
§9 had both a "Step 3b/4 COMBINED" row (line 250) and a leftover separate
"Step 4: Parity test" row (line 251) — internally contradictory. §15 also
still sequenced Steps 3b + 4 as four separate authoring/dispatch/ratify
beats (steps 12-17), contradicting §12 Q5's "same PR" decision and the
explicit "Update to §15" note at §12 line 325.

Resolution: collapsed §9 to one COMBINED row absorbing the parity-TestClaim
mechanics + P5 dissolution receipt from the deleted Step 4 row; collapsed
§15 steps 12-17 into single COMBINED authoring + dispatch + ratify (steps
12-14). Added explicit `feedback_paper_shrink_variants` reasoning in both
sections.

Finding 2.5 (PM intent — substrate-capability bundled vs separate):
§12 Q1 line 288 said "Director-recommend: (b) bundled" while §13 line 340
listed substrate-capability as a non-goal of PB-3 and §15 step 11 had
"WAIT for substrate-capability landing" — three sections, two different
execution paths.

Resolution: annotated Q1 as "PENDING operator/PM ratification" with
explicit default-execution clause: until operator ratifies, the doc
treats substrate-capability as path (a) separate lane (matching §13 + §15
+ §9 row). If/when ratified to (b), §13 drops the non-goal and §15 step 11
collapses into the COMBINED Step 3b/4 brief.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 5619afa) — parse_tables.dag as single enumerated authority

Codex caught me copying the prose summary at parse_tables.dag:23-29 (which
enumerates only SG-2c-numbered families) instead of grepping the live
`^type [A-Z]` declarations. Result: SoftKeywordIdentRow (line 334) was
missing from §3.2 / §5.1 / §6 / §12 because it lacks an SG-2c-N number
in the prose summary.

Per `feedback_parallel_representation_debt`: structural fix is to stop
hand-enumerating in the doc — cite parse_tables.dag itself as the single
enumerated authority and use `type`-declaration line-anchors for the
worked example, not a hand-maintained count.

Edits:
- §3.2 §"Live substrate authority": replaced the SG-2c-numbered bullet
  list with `type`-declaration line-anchor enumeration including
  SoftKeywordIdentRow at parse_tables.dag:334 + the supporting enum
  BinaryOpLevel at line 133. Added codex-finding callout explaining the
  miss + the discipline shift.
- §3 preamble line:33, §3.2 line:45 callout, §3.2 line:57 framing,
  §3.2 §"Substrate authority" line:71, §5.1 line:171-180, §12 Q2 line:298,
  §12 Q6 line:334: all hardcoded "6 table-families" counts dropped; doc
  now points readers to §3.2 enumeration / `parse_tables.dag` directly.
- §5.1 sub-enumeration list (the parallel 6-item list at lines 173-178)
  deleted; replaced with redirect to §3.2 + restated 3a-vs-3b/4 scope split.

Code-level check before commit: `grep -nE '^type [A-Z]' src/v3/compiler/parse_tables.dag`
returns 7 types: BinaryOpLevel (133), BinaryOpRow (167), TopLevelItemKwRow (289),
SoftKeywordIdentRow (334), BracketRow (385), PrimaryPrefixRow (449),
PrimaryAtomRow (486). Doc enumeration matches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING (sha f08b952) — bind TokenizeDiagnostic to live CharClass authority

Codex Finding 1 (sha f08b952, 216s thinking): the §4.2 TokenizeDiagnostic
draft and the §5.1/§5.2 headings used `ScannerCharClass` / `ScannerClassRef`
— names that exist only as the *generated Rust enum spelling*, not as a
declared .dag substrate type. Verified via grep:

  grep -rn '^type CharClass\|^type ScannerC' dsl/ src/v3/

returns ONE authority: `dsl/std/unicode.dag:62`
  type CharClass = Whitespace | Digit | IdentStart | IdentContinue

consumed at `src/v3/compiler/tokenize.dag:103`
  data ascii_scan_order: List<CharClass> = [Whitespace, Digit, IdentStart, IdentContinue]

There is no `ScannerCharClass` declaration anywhere — that name was copied
from generated Rust without grep-verification, the same failure mode as
`feedback_grep_substrate_before_naming_ratification` (carrier-name
collision discipline).

Resolution:
- §4.2 TokenizeDiagnostic carrier: `expected_class: ScannerClassRef` →
  `expected_class: CharClass`, dropped the `type ScannerClassRef =
  ScannerCharClass` alias entirely; added a codex-finding callout citing
  the substrate authority + naming the failure mode.
- §5.1 heading "Byte → ScannerCharClass dispatch" → "Byte → CharClass
  dispatch"; bullets unchanged; added line-anchor cites for the substrate
  authority + explicit "NOT ScannerCharClass" disclaimer.
- §5.2 heading "ScannerCharClass → token-recognition state machine" →
  "CharClass → token-recognition state machine".

Finding 2 (TokenizedSource not reconciled with parse input contract): no
new fix required — already resolved by commit 6adb992 (TokenizedSource
extension dropped entirely; tokenize uses Result<List<Token>,
TokenizeDiagnostic>; List<Token> is the single canonical boundary carrier
consumed by parse). Codex was reviewing sha f08b952, which predated the
TokenizedSource drop.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3138 INLINE BLOCKING line:188 — clarify §7.3 chain is Ok-branch propagation

Cursor INLINE at sha f08b952 worried that §7.3's cross-stage chain
"tokenize → List<Token> → parse → ..." was inconsistent with §4.3's
TokenizedSource carrier (diagnostics not flowing forward). At HEAD the
TokenizedSource extension is dropped (commit 6adb992); tokenize uses
Result<List<Token>, TokenizeDiagnostic>, so List<Token> IS the canonical
Ok-branch payload that flows forward and Err branches terminate the
pipeline fail-fast.

To make this explicit at §7.3 (instead of leaving readers to infer it
from §4.3), annotated the chain with:
- "Ok-branch propagation; Err branches are stage-terminal fail-fast per
  §4.3 Result-sum discriminator" framing prefix.
- Per-stage Result/typed-state annotations: tokenize/parse show
  Result<Ok, Err>; lower/infer show typed-state structural-output;
  emit shows Result<EmittedArtifact, EmissionDiagnostic>.
- Explicit "on any stage's Err branch the pipeline aborts at that stage
  (no partial-output propagation across boundaries)" trailer.

This makes the chain self-consistent vis-a-vis §4.3 without requiring
the reader to walk back-and-forth, and prevents future readers from
re-introducing a TokenizedSource-shaped extension to "make diagnostics
flow forward" — they already do, just via the Err branch terminating
the pipeline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3127 BLOCKING (sha d15e1f2) — character-level scaffold dissolution trigger

openai-pro REQUEST_CHANGES on sha d15e1f2: §1 line:26 character-level
under-consumption scaffold named the problem but lacked a checkable
dissolution trigger. SG-1a scaffold above had the right shape — "once
those bodies lower structurally under compile_to_dag, delete the raw-text
extractor + derive directly from lowered Dag in same PR." Character-level
scaffold just said "PB-2 Step 4 carries this scope" — a lane assignment,
not a trigger.

Resolution: rewrote §1 item 2 with the same SG-1a-shape trigger structure:
- Substrate-consumption condition (a): scanner classes / string escape /
  local punctuation retype to `dsl/std/unicode.dag` `CharClass` /
  `char_in_class` (concrete field retypes named:
  `StringEscapeSpec.suffix: Char`, `LocalPunctSpec.pattern: List<Char>`,
  `string_literal_delimiter: Char`).
- Codegen-driver condition (b): `tokenize_generated.rs` no longer emits
  hidden `byte.is_ascii_*` predicates because the driver reads class
  facts structurally from lowered `tokenize.dag`.
- Same-PR dissolution: delete the parallel character-predicate scaffold
  in the same PR that flips substrate consumption — no Rust-and-`.dag`
  coexistence per `feedback_paper_shrink_variants`.
- Cross-ref to §9 Step 4 gating prereqs: ROADMAP.md:467 + ROADMAP.md:416
  Class 5 Gap 3 + std.unicode bootstrap/load-set decision (already cited
  in §9 from earlier commit 6adb992).

Per openai-pro's framing: "small fix — mirror the SG-1a scaffold wording
by naming the exact substrate-consumption condition and same-PR deletion
receipt for the hidden Rust character predicates."

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3127 APPROVE_WITH_COMMENTS — §14 + §15 Q6-already-rejected sweep

Cursor review id 12405 (APPROVE_WITH_COMMENTS) caught the same
`feedback_discipline_change_audit_all_contract_mentions` failure mode
recurring: §12 Q6 was resolved REJECTED in commit 6adb992, and §14
"Surfaces awaiting" + §12 Q6 heading + §9 Step 2 row were updated, but
two §-internal contract restatements were missed:

- §14 acceptance criterion 11: "Operator/PM ratification on §12 Q1-Q6"
- §15 step 1: "Operator / PM-delegate ratifies §12 Q1-Q6"

Both contradicted §12 Q6 + §14 "Surfaces awaiting" (which already said
"Q1-Q5 only"). A worker reading §14/§15 could schedule sign-offs on Q6
after it was already resolved-rejected elsewhere.

Resolution: both sites now say "Q1–Q5" with the explicit Q6-rejected
crossref + "see §14/§15 for same scoping" pointer at the §14 criterion
so the three sections agree internally.

Cursor verdict was APPROVE_WITH_COMMENTS (substantive APPROVE — "fix the
checklist/sequence so every section agrees Q6 is closed"); the
exploratory volatile-line-anchor note is harmless and out-of-scope for
this PR.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING (sha 887c696) — 3 findings swept across 3 L2.5 docs

Codex review id (sha 887c696, 333s thinking) flagged three findings,
all `feedback_discipline_change_audit_all_contract_mentions` recurrences
where partial sweeps left §-internal contradictions.

Finding 1 — infer doc cross-stage trigger leak:
§4.2 InferDiagnostic dissolution trigger said "when Step 2 worker brief
enumerates the full variant set against `parse_generated.rs`
Diagnostic::ParseError, lower.rs Diagnostic construction sites, and
infer.rs Dag::mark_unresolved emission sites." That cross-stage trigger
surface is wrong for the infer-specific scaffold. Narrowed to infer-only:
"against `src/v3/compiler/src/infer.rs` Diagnostic construction sites +
`Dag::mark_unresolved` emission sites (infer-stage only)." Parse and lower
have their own per-stage carriers + own Q7 mapping; this doc no longer
reaches into their dissolution-trigger surface.

Finding 2 — parse + tokenize coproduct receipts:
`ParseDiagnosticKind` (parse §4.2) and `TokenizeDiagnostic` (tokenize §4.2)
sums were declared without the 🟡 SCAFFOLD / 🟢 TERMINAL classification +
dissolution trigger that the infer doc carries (per modeling-discipline
Practice 4 + `feedback_coproduct_dissolution`). Added matching scaffold
receipts to both: 🟡 SCAFFOLD at PROPOSED stage with stage-local dissolution
trigger (Step 2 brief enumerates against parse_generated.rs / tokenize_generated.rs
Diagnostic::* construction sites — stage-only, not cross-stage); promote
to 🟢 TERMINAL when full variant set lands. Both also carry the anti-bridge
note per Q6.5.

Finding 3 — tokenize Q7 reconciliation:
The parse doc has the "Q7 DONE 2026-05-15T00:21:19Z" annotation at §15 step 4
(commit f85fa1f) but the tokenize doc still framed Q7 as a pending lane
dependency at §4.2 line:113 ("Lane dependency"), §15 step 3, and §14
"Surfaces awaiting". Annotated all three:
- §4.2 "Lane dependency": Q7 DONE timestamp + "TokenizeDiagnostic per-stage
  variant authoring is the remaining lane work."
- §15 step 3: full DONE annotation matching parse doc §15 step 4 shape +
  cross-ref to the parse doc.
- §14 "Surfaces awaiting": strikethrough + DONE annotation.

PR #3127 also carries the tokenize doc; same edits will port there in a
companion commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3138 INLINE BLOCKING line:372 — §16 Surfaces awaiting Q7 contradiction

Cursor inline at sha 887c696 caught the symmetric finding to the codex
5-finding BLOCKING #3: parse-doc §15 step 4 already annotated Q7 as DONE
(commit f85fa1f) but parse-doc §16 "Surfaces awaiting" still listed Q7
as pending. Same `feedback_discipline_change_audit_all_contract_mentions`
sweep failure — the tokenize doc had three sites carrying the pending
framing (commit e9739ea fixed those) but the parse doc's §16 site was
missed in the original Q7-DONE sweep.

Resolution: §16 bullet now strikes through + "DONE 2026-05-15T00:21:19Z
(PR #3077 merged carrying Q7 ratification; see §15 step 4)" — same shape
as the tokenize doc's §14 "Surfaces awaiting" Q7-DONE annotation landed
in e9739ea.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 15, 2026
…y + §4.2 live-vs-proposed clarity

PR #3127 was merged at 02:04:22Z but codex's later review on sha 8e682ee
caught two findings that still live in main. Both addressed in this PR
(PR #3140) since the tokenize doc lives on main and this is the open
post-merge cleanup branch.

Finding 1 — §3.1 source-identity drop:
Earlier draft framed tokenize input as bare `String → List<Token>`. That
contradicts live `tokenize_generated.rs:96`:
  pub fn tokenize(source: &str, file: &str) -> Result<Vec<Token>, Diagnostic>
The `file` parameter is load-bearing because `SourceSpan` requires a
file/source-id field for byte ranges to be attributable; without it the
Token + Diagnostic span fields would have to fabricate source-id at the
pipeline boundary (P3 fail-closed violation). §3.1 now ratifies the live
two-input shape:
- `source: String` — UTF-8 source text (primitive).
- `file: SourceFileId` — source-identity carrier; Step 2 brief ratifies
  the appropriate `.dag` shape (NonEmptyStr newtype or richer sum if
  multiple source-class kinds).

Finding 2 — §4.2 live-vs-proposed blur:
Prior framing said "TokenizeDiagnostic (substrate extension per Decision
2.B)" without making clear that this is a PROPOSED per-stage carrier,
not the live one. The live carrier at `diagnostics.dag:150` is generic
`Diagnostic { kind: AnyDiagnosticKind, ... }` with kind sum at `:139-142`
discriminating CompilerKind vs LensInstanceKind; tokenize emits today
via `Diagnostic::TokenizerError`-shaped sites carrying
`kind: CompilerKind(...)`. §4.2 heading now reads "PROPOSED — NOT yet
live" + a codex-finding callout explicitly distinguishing the live
carrier from the proposed per-stage refinement + Q7 status DONE.

Both fixes preserve the e9739ea 🟡 SCAFFOLD coproduct receipt that
was already addressing codex sha 887c696 Finding 2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 15, 2026
… signature sweep

Cursor inline at sha 8e682ee line:121 (the §4.3 Step 2 signature site)
caught the same source-identity drop my §3.1 fix already addressed at the
input-type framing — but two companion sites still had the single-input
signature:
- §4.3 line:151 Step 2 signature recap
- §9 Step 2 row (4-step migration table)

Both now read:
  fn tokenize(source: String, file: SourceFileId) -> Result<List<Token>, TokenizeDiagnostic>

matching the §3.1 source-identity discipline + live `tokenize_generated.rs:96`
`pub fn tokenize(source: &str, file: &str)` shape.

Same `feedback_discipline_change_audit_all_contract_mentions` recurrence
— my Finding 1 fix at §3.1 (commit 18af49f) didn't sweep the §4.3 + §9
contract-signature restatements.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 15, 2026
…iagnostic LIVE/PROPOSED clarity

Cursor inline caught ambiguity in the §6 prereq-table row for
"TokenizeDiagnostic substrate extension." Old wording in column 4 read
"Carrier LIVE; per-stage variant authoring NEW" which could be parsed
as "TokenizeDiagnostic carrier itself is LIVE" — but no `type
TokenizeDiagnostic` exists in `dsl/std/` or `src/v3/std/`. Only the
generic `type Diagnostic` at `diagnostics.dag:150` is live.

Cursor's secondary claim that "line 150 is EmissionDiagnostic receipt
prose" is INCORRECT — verified via grep + read: `diagnostics.dag:150`
is the `type Diagnostic { kind: AnyDiagnosticKind, ... }` declaration
header. `type EmissionDiagnostic` lives separately at `diagnostics.dag:201`.
The §6 row's citation of `:150` for the underlying Diagnostic carrier
is correct.

But the ambiguity-in-wording finding stands. Reframed the row:
- Column 1: explicitly "PROPOSED — no `type TokenizeDiagnostic` exists yet"
- Column 2: clarifies the LIVE underlying carrier is `type Diagnostic`
  with `kind: AnyDiagnosticKind` sum at `:139-142`; cites Q7 ratification
  done timestamp.
- Column 4: explicit "Underlying Diagnostic carrier LIVE at HEAD;
  per-stage TokenizeDiagnostic variant authoring is NEW work, NOT live
  substrate" with cursor-finding callout.

Consistent with §4.2's "PROPOSED — NOT yet live" framing from commit
18af49f.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

Status reply for cursor INLINE BLOCKING at docs/design-tokenize-stage-l25-model.md:166 (sha 8e682ee). PR #3127 is merged (02:04:22Z); fix landed on post-merge PR #3140 commit dc1a8b4.

Two parts to the finding:

  1. "That line is EmissionDiagnostic receipt prose" — INCORRECT. Verified via grep + read of src/v3/std/diagnostics.dag:

    • :150 is the live type Diagnostic { kind: AnyDiagnosticKind, span: SourceSpan, message: String, correction: Correction } declaration header.
    • :201 is type EmissionDiagnostic, a distinct carrier.

    The §6 row's citation of :150 for the underlying Diagnostic carrier is correct.

  2. "No TokenizeDiagnostic carrier exists, so the design has not named a single diagnostic authority per INVARIANTS P2" — CORRECT as a wording finding. No type TokenizeDiagnostic exists. The old §6 row read "Carrier LIVE; per-stage variant authoring NEW" which could be parsed as "TokenizeDiagnostic carrier is LIVE" rather than "the underlying generic Diagnostic carrier at :150 is LIVE, and the per-stage TokenizeDiagnostic refinement is NEW work."

    PR docs(r3): L2.5 post-merge cleanups — codex 5-finding BLOCKING + cursor inline Q7 sweep #3140 dc1a8b4 reframes the row explicitly: column 1 marks TokenizeDiagnostic as "PROPOSED — no type TokenizeDiagnostic exists yet"; column 2 names type Diagnostic as the live underlying carrier with kind: AnyDiagnosticKind sum at :139-142; column 4 disambiguates "Underlying Diagnostic carrier LIVE at HEAD; per-stage TokenizeDiagnostic variant authoring is NEW work, NOT live substrate." Consistent with §4.2's "PROPOSED — NOT yet live" framing landed in PR docs(r3): L2.5 post-merge cleanups — codex 5-finding BLOCKING + cursor inline Q7 sweep #3140 18af49f earlier.

— sent from zesty-bear-812

briansrls added a commit that referenced this pull request May 15, 2026
…+ cursor inlines (#3141)

* docs(r3): PB-2 tokenize pipeline-stage L2.5 domain model — DRAFT for ratification

Director-tier L2.5 model for PB-2 tokenize per operator 2026-05-14
ratification (Decision 1.A scoping = Option A).

PB-2 is the FURTHEST-ALONG pipeline stage — substrate authority
already lives in `.dag`:
- src/v3/std/tokenize.dag (Token + TokenKind taxonomy; LIVE 143
  lines)
- src/v3/compiler/tokenize.dag (tokenizer implementation; LIVE
  154 lines)
- src/v3/compiler/src/tokenize_generated.rs (AUTO-GENERATED; 362
  lines)

This is the END STATE that all other pipeline-stage migrations
target. PB-2 L2.5 is correspondingly lighter — mostly verification
+ residual hand-Rust retirement, NOT new substrate authoring.

Distinct §9 4-step framing:
- Step 3 = VERIFY substrate completeness (audit per
  feedback_paper_shrink_variants)
- Step 4 = HANDOFF/RETIRE residual hand-Rust scaffolding
  (coordinates with PB-Bootstrap-Process lane for codegen-driver
  retirement)

Captures audit dimensions explicitly:
- scanner-class definitions = declarative byte-pattern membership
- recognition tables = closed-axis enums
- state machine = structural transitions
- no V2 `pub mod tokenize` absorption check

§12 Q1: codegen-driver retirement scope — Director-recommend
PB-Bootstrap-Process handles all codegen-driver retirement
cross-cuttingly (not per-stage paper-shrink risk).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): §12.6 citation fix — tokenize is PB-2 lane scope per design-pure-bootstrap-zero.md

Per cursor PR #3085 finding: §12.6 explicitly tables only 4
pipeline-stage migrations (emit→lower→infer→parse); tokenize is
per design-pure-bootstrap-zero.md PB-2 lane. Same fix as PR #3085
commit 89fbd7a2a applied here.

INVARIANTS P1 — documentation must not overstate authority cites.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): preemptive fix for PR #3126 codex BLOCKING propagation — PB-2 live state honesty

Same class as codex INLINE BLOCKING #3126 finding 1 (live state
honesty for diagnostic coupling) applied preemptively to PB-2
tokenize L2.5.

PB-2 §4.3 had "diagnostics coupled INTO List<Token>" framing
which would overstate the live carrier shape (bare List<Token>
has no diagnostic field; tokenize_generated.rs:96 today returns
Result<Vec<Token>, Diagnostic>).

Fix: §4.3 reframed with PROPOSED substrate extension explicit —
new `TokenizedSource { tokens, diagnostics }` wrapper carrier as
the typed-state output. Step 2 brief includes wrapper authoring
in pipeline-slot PR scope.

Same discipline as PR #3126 commit bdff8c500.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 — fix PB-2 lane citation (design-pure-bootstrap.md not -zero.md)

Per cursor APPROVE_WITH_COMMENTS /api/reviews/12087: PB-2 lane is
defined in docs/design-pure-bootstrap.md §"PB-2 — tokenize retire"
(line ~134), NOT docs/design-pure-bootstrap-zero.md. The -zero.md
doc has Subsumed-lanes list with PB-1/PB-4/PB-5/PB-6 but NOT PB-2.

Propagated fix applies same cite-error correction as PR #3066 §1.8
discipline: cite the actual doc, not an adjacent doc with similar
name. INVARIANTS P2 single-authority-citation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix cursor PR #3127 BLOCKING contradictions + add §12 Q6

Cursor APPROVE_WITH_COMMENTS (/api/reviews/12093) caught two
substantive contradictions I introduced when adding TokenizedSource
in commit 2b9756bb5:

1. §4.3 vs §9 Step 2 signature mismatch — §4.3 said `-> TokenizedSource`
   but §9 Step 2 row still said `-> List<Token>`. Same
   `feedback_discipline_change_audit_all_contract_mentions` issue
   that's recurred 4x this session.

2. §4.3 referenced "§12 Q-new" but §12 only had Q1-Q5; broken anchor.

Fix:
1. §9 Step 2 row updated: signature `-> TokenizedSource` with
   wrapper carrier shape `{ tokens, diagnostics }` per §4.3
2. §4.3 anchor updated: "§12 Q6" (resolved)
3. Added §12 Q6: TokenizedSource carrier shape ratification —
   (a) wrapper record vs (b) per-Token diagnostic coupling;
   Director-recommend (a) for PB-3 SurfaceModule parallelism
4. §14 + §15 + §16 Q-list refs updated to Q1-Q6

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — Token shape + Q4 audit boundary

Codex REQUEST_CHANGES (sha b881de23) with 2 substantive findings:

1. §4.1 Token shape claim "optional lexeme: String" — wrong per
   live substrate at src/v3/std/tokenize.dag:65-67. Live Token
   has 2 fields only (kind + span); lexeme-content lives ON the
   TokenKind variants (Ident(String) / IntLit(String) / etc.).

   Fix: corrected §4.1 to reflect live carrier shape; payloads
   on TokenKind variants noted explicitly.

2. §12 Q4 substrate-completeness audit scoped only to
   tokenize_generated.rs — missed the regen_tokenize codegen-
   driver boundary. If regen_tokenize carries scanner-logic
   decisions (rather than mechanical template-rendering of
   substrate facts), the substrate isn't complete — the driver
   IS hand-Rust scanner logic in disguise.

   Fix: Q4 audit extended with (d) regen_tokenize codegen-
   driver logic audit + (e) ROADMAP.md deferral row option per
   feedback_paper_shrink_variants P5 receipt discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): fix codex BLOCKING PR #3127 — TokenizeDiagnostic PROPOSED + scaffold disclosure

Codex REQUEST_CHANGES (sha b881de23, full body) caught 2 substantive
overstatements:

1. §2 line 46 claimed "tokenize failures produce typed
   TokenizeDiagnostic variants" — but live tokenize_generated.rs:96
   returns generic `Result<Vec<Token>, Diagnostic>` with
   `Diagnostic::TokenizerError { message, span, correction }`. Typed
   TokenizeDiagnostic is a PROPOSED extension, not live state.

   Fix: §2 reframed with live state explicit + TokenizeDiagnostic
   marked PROPOSED per PR #3077 §12 Q7 ratification path.

2. §1 line 22 + §6 + §9 Step 2 line 198 framed PB-2 as "mostly
   verification" — but live tokenize.dag:16-30+ has TWO explicit
   tracked scaffold zones:
   - SG-1a: regen_tokenize parses raw source text for
     dag_keyword_set / dag_operators (ValueBody::Unparsed)
   - Character-level under-consumption: StringEscapeSpec /
     LocalPunctSpec.pattern / string_literal_delimiter as opaque
     Strings; hidden Rust character predicates (byte.is_ascii_digit
     etc.) at tokenize_generated.rs:15-22 leaking through codegen

   Residual hand-Rust is NOT just the codegen artifact — it includes
   (a) regen_tokenize logic, (b) SG-1a raw-text-extractor scaffold,
   (c) character-predicate scaffold leaking through codegen.

   Fix: §1 + §6 + §9 Step 2 reframed honestly. PB-2 is "FURTHER
   ALONG but not complete"; Step 4 carries scaffold-retirement
   scope, not just codegen-artifact retirement.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — §16 + §9 Step 2 internal contradictions

Two post-merge doc-internal contradictions caught by reviewers
after operator merged PR #3077 / #3126 / #3085 at 2026-05-15T00:21Z.

**PR #3077 PB-4 lower §16 fix**:
§16 "Memory disciplines applied" bullet said "diagnostics coupled
INTO PreInferDag via biconditional" — but §4.3 (per openai-pro
DiagnosticAnchor fix commit b812db91b) constrains biconditional
to PortAnchor-only. Other anchor kinds (DeclarationAnchor /
RecordFieldAnchor / SurfaceFormAnchor) couple without port-state.
Fix: §16 bullet honors §4.3 anchor-typed framing.

**PR #3126 PB-3 parse §9 Step 2 fix**:
§9 Step 2 row described diagnostics as "coupled INTO SurfaceModule"
as if live — §4.3 correctly marks it PROPOSED. Same
feedback_discipline_change_audit_all_contract_mentions pattern
that's recurred this session.
Fix: §9 Step 2 row clarified — "PROPOSED extension per §4.3";
Step 2 PR scope includes authoring the diagnostics field
extension, NOT a live coupling.

Per feedback_discipline_change_audit_all_contract_mentions: when
a substantive fix changes a discipline framing, audit ALL sections
(framing + contract + handoff). Post-merge audit surfaced these
residual contradictions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): post-merge fix-forward — PR #3126 codex BLOCKING (GrammarSpec parallel-authority + fail-closed weakening)

Codex REQUEST_CHANGES on already-merged PR #3126 (/api/reviews/12175):
2 substantive findings on the post-merge doc.

**Finding 1 (P2 violation — GrammarSpec parallel authority)**:

§3.2 says GrammarSpec is compile-time-only, NOT runtime-
interpreted (per Decision 3.B (b) operator override). But the
proposed stage contract still took `grammar: GrammarSpec` as
runtime input. Creates two authorities (compiled parser tables
+ runtime GrammarSpec value).

Fix: §4.3 signature reframed to `fn parse(tokens: List<Token>)
-> Result<SurfaceModule, ParseDiagnostic>` — NO runtime
GrammarSpec input. Compile-time generated parser tables consumed
via internal dispatch. Step 2 + Step 4 rows updated.

**Finding 2 (P3 + Practices 1/2 — fail-closed weakening)**:

Live parser at parse_generated.rs:138 returns `Result<SurfaceModule,
Diagnostic>` (fail-closed; aborts on first error). Earlier draft
proposed `SurfaceModule` with embedded diagnostics — would let
partial-parse states be constructible + let downstream observe
"success" output after parse failure. Violation of fail-closed
discipline.

Fix: signature preserves Result-sum (matches live + emit's
pattern). Distinguished cross-stage:
- Result-sum (parse + emit): fail-fast output domain
- Typed-state-with-coupled-diagnostics (lower + infer): structural
  output domain where partial-failure IS valid intermediate

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 post-merge APPROVE_WITH_COMMENTS — §4.1 + §16 fail-closed honesty

Cursor caught 2 more post-merge contradictions on PR #3126:

1. §4.1 line 81 "construction-time invariant" talks about
   "ParseDiagnostic in the diagnostic stream" tied to SurfaceModule
   path — but §5.2 + live parse_generated.rs:138 use
   Result<SurfaceModule, Diagnostic>. §4.1 reads as claim about
   today's plumbing.

   Fix: §4.1 reframed — live boundary explicit (Result-sum);
   construction-time invariant scoped to Ok-arm SurfaceModule +
   Err-arm ParseDiagnostic, no partial-parse with embedded
   diagnostics.

2. §16 line 397 cites C-8 as "ParseDiagnostic coupled INTO
   SurfaceModule" without qualifier — but §4.3 (post codex
   REQUEST_CHANGES fix) constrains to Result-sum.

   Fix: §16 bullet honors §4.3 Result-sum framing; cross-stage
   discriminator named.

Same recurring feedback_discipline_change_audit_all_contract_mentions
pattern — substantive fix to §4.3 + §9 Step 2 left §4.1 + §16
inconsistent. Post-merge audit catches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — TypeConnective extension stop-signal

Cursor INLINE BLOCKING caught §3.2 line 66 "rules extend
automatically" weakens substrate-extension stop-signal. Thesis
discipline: a 7th TypeConnective variant requires explicit C1
audit + named infer-rule receipt.

Fix: §3.2 reframed. New TypeConnective variants do NOT extend
automatically; require explicit C1 substrate-extension audit +
named infer-rule receipt for the new variant's structural
inference behavior.

Per-variant structural facts means new variants need new
per-variant facts, NOT silent inheritance.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING — AlgebraAxis + InferDiagnostic Practice 4 classification

Cursor INLINE BLOCKING /api/reviews/12218 (line:112): proposed
substrate coproducts (AlgebraAxis + InferDiagnostic) lack
🟢/🟡/🔴 classification + ledger/trigger per modeling-discipline
Practice 4 (Coproduct dissolution).

Fix: added 🟡 SCAFFOLD classification + named dissolution
trigger for both:

AlgebraAxis 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full algebra-axiom set
  against infer.rs check sites + verifies coverage parity
  with live verification.dag:146 AlgebraicLawKind 3-variant
  subset → promote to 🟢 TERMINAL.

InferDiagnostic 🟡 SCAFFOLD:
- Trigger: Step 2 brief enumerates full variant set against
  parse_generated.rs / lower.rs / infer.rs diagnostic emission
  sites → promote to 🟢 TERMINAL.
- Anti-bridge per Q6.5: does NOT collapse into
  CompilerDiagnosticKind without substrate-extension
  ratification per PR #3077 §12 Q7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — §5.1 SurfaceItem allocation correction

Cursor INLINE BLOCKING /api/reviews/12251 line:130: §5.1 said
"every SurfaceItem variant maps 1:1 to a Declaration placeholder"
but live lower.rs:2950-2958 explicitly skips Let/Module/Import
in collect_symbols.

Verified via Read of lower.rs:2956-2958:
  SurfaceItem::Let { .. } => continue,
  SurfaceItem::Module { .. } => continue,
  SurfaceItem::Import { .. } => continue,

Fix: §5.1 reframed — DeclarationAllocating variants (Fn /
FnExternalBody / Data / TypeAtom / TypeRecord) map to
placeholders; NonDeclarationAllocating variants (Let / Module /
Import) skip allocation per live lower.rs behavior.

Let-bodies lower to Bind expressions in Pass 2; Module/Import
are parsed-facts preserved but un-declared.

Earlier "every SurfaceItem variant" framing overstated; would
have steered Step 2/3 worker into wrong allocation contract.
Per INVARIANTS P1/P2 live-state honesty + facts-flow-forward.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING — ElaborationSpec scope broadened

Cursor INLINE BLOCKING line:33: ElaborationSpec was defined only
as Surface→Behavior recipe mapping, but lower constructs
Declarations / TypeConnectives / BranchPatterns / Bindings as
well. Non-Behavior lowering decisions outside declared authority
violates THESIS substrate ownership + INVARIANTS P2.

Fix: §3.2 ElaborationSpec scope broadened to ALL lowering
decisions:

1. SurfaceItem → Declaration recipes (Fn / Data / Type variants +
   Let/Module/Import skip-allocation per §5.1)
2. SurfaceType → TypeConnective recipes (Atom / Arrow / Compose /
   Disj construction)
3. SurfaceExpr → Behavior recipes (Value / Transform / Branch /
   Loop / Bind construction)
4. SurfacePattern → BranchPattern recipes (ResolvedVariant /
   UnresolvedVariant / record-pattern construction)
5. Binding-site rules (Bind params + result_port construction)

ElaborationSpec is single-authority across ALL axes; no axis
lives in implementation-tier hand-Rust.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — §7.1 ordering vs independence clarification

Cursor APPROVE_WITH_COMMENTS /api/reviews/12265 line 197: §7.1
mixed two claims:
- "parse migrates AFTER tokenize substrate-side stable" (ordering)
- "PB-3 parse migration is independent of PB-2 tokenize migration
  status" (independence)

Read as contradictory by reviewers. Need one coherent story.

Fix: §7.1 reframed with two distinct axes explicit:

1. Substrate-stability ordering (SELF_HOSTING.md §2 bottom-up):
   tokenize Token carrier shape must be stable BEFORE parse
   migrates. Already true at HEAD (tokenize.dag:65-67 declares
   live carrier). ✓

2. Migration-timing independence (parallel-dispatch axis): PB-3
   parse migration ships in parallel with PB-2 residual-retirement
   work (SG-1a + character-level scaffold + codegen-driver
   retirement per PB-2 L2.5 §1). What parse needs is the stable
   Token CARRIER; PB-2's migration is about retiring residual
   hand-Rust, not changing the carrier.

Both claims coherent on the axis split; not contradictory.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — ParseDiagnostic SourceSpan structural requirement

Cursor INLINE BLOCKING /api/reviews/12277 line:65: ParseDiagnostic
variants like UnexpectedToken lacked SourceSpan field;
List<ParseDiagnostic> cannot satisfy fail-closed source
attribution structurally without span on every variant.
INVARIANTS P2/P3 violation.

Fix: every ParseDiagnostic variant now carries SourceSpan
structurally:
- UnexpectedToken: added span: SourceSpan
- UnterminatedConstruct: opener_span: SourceSpan (already present)
- InvalidLiteral: added span: SourceSpan
- DuplicateRecordFieldLabel: added span: SourceSpan (current site)
  + prior_span: SourceSpan (prior site; both required)

Per INVARIANTS P2/P3 fail-closed source attribution discipline:
every diagnostic emission carries structural source-span
provenance; not optional.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Token authority cite phantom-reference fix

Cursor INLINE BLOCKING /api/reviews/12279 line:29: §3.1 cited
`src/v3/compiler/src/tokenize.rs` as "current hand-Rust" for
Token carrier — but `tokenize.rs` (without _generated suffix)
doesn't exist. Live state has tokenize.dag (live substrate) +
tokenize_generated.rs (codegen artifact).

Per design-pure-bootstrap.md PB-2 lane: tokenize retire has
substantially landed. The reference was an earlier-draft
phantom from when Token-was-hand-Rust framing was the assumption.

Fix: §3.1 reframed — Token type lives in LIVE
src/v3/std/tokenize.dag:65-67 shared taxonomy; tokenizer
implementation also live at src/v3/compiler/tokenize.dag (154
lines); codegen artifact at tokenize_generated.rs. tokenize.rs
phantom reference removed explicitly. PB-2 substantially landed
per design-pure-bootstrap.md; residual scaffold-retirement
scope per PB-2 L2.5 PR #3127.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 2 pipeline-slot target fix (compiler.dag → pipeline.dag)

Cursor INLINE BLOCKING /api/reviews/12281 line:138: §9 Step 2
referenced generic "compiler.dag" but dsl/gunbc/compiler.dag:24
explicitly directs internal pipeline (Tokenize → Parse → ...)
to src/v3/compiler/pipeline.dag, NOT generic compiler.dag.

Worker briefs authored against this doc would target the wrong
file for pipeline-slot declaration. P2 single-authority violation.

Fix: §9 Step 2 row in ALL 4 L2.5 docs (PB-2 / PB-3 / PB-4 / PB-5)
updated:
- "declared in compiler.dag" → "declared in src/v3/compiler/pipeline.dag (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag)"
- substrate column: "compiler.dag refinement" → "pipeline.dag refinement"
- §13 "Step 2 (pipeline-slot in compiler.dag)" → "pipeline-slot in src/v3/compiler/pipeline.dag"

Same phantom-citation class as the tokenize.rs phantom (commit
8ae37b4a5): I cited generic file path without verifying which
specific file is authoritative per project structure. Should
have grep'd dsl/gunbc/compiler.dag header notes before authoring.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING — Step 3b/4 phasing P5 receipt unambiguity

Cursor INLINE BLOCKING /api/reviews/12283 line:181: Q5 said
Step 3b "full parser body .dag migration + parser body deletion
in same PR", but §15 sequence schedules Step 4 parity/deletion
AFTER Step 3b merges. P5 dissolution receipt ambiguous.

If Step 3b lands .dag parser body BEFORE Step 4 deletes Rust
parse() body, Rust + .dag parser bodies coexist temporarily —
paper-shrink-relocation risk per feedback_paper_shrink_variants.

Fix:
1. §9 Step 3b row reframed as "Step 3b/4 COMBINED" — atomic
   single PR (full .dag parser body + parity TestClaim +
   parse_generated.rs:138 deletion + census shrink). Cannot land
   .dag parser body before Rust deletion.
2. §12 Q5 phasing clarified:
   - Phase 3a (separate PR): grammar table extension; P5 receipt
     = ROADMAP deferral row naming Step 3b/4 as future-receipt
   - Phase 3b/4 COMBINED (single PR): atomic substrate substitution
3. §9 + §15 update notes: sequence collapses steps 12-17 into
   single dispatch+merge for combined phase

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 high-level BLOCKING findings 1 + 2

Codex high-level BLOCKING (sha 140eb6bb) had 4 findings:
- Findings 3 + 4 already addressed in PR #3138 commits 015977347
  + a1607a88d
- Findings 1 + 2 addressed in this commit

**Finding 1 (Diagnostic kind/record conflated)**:
§4.2 ParseDiagnostic was modeled as variants-directly with span
+ kind-fields mixed. Live Diagnostic at diagnostics.dag:150 uses
record-wraps-kind pattern (`{ kind, span, message, correction }`).
Need consistent shape.

Fix: refactored to record-wraps-kind:
- `type ParseDiagnostic { kind: ParseDiagnosticKind, span: SourceSpan }`
- `type ParseDiagnosticKind = UnexpectedToken | UnterminatedConstruct | InvalidLiteral | DuplicateRecordFieldLabel | ...`

Span lives on ParseDiagnostic record (single source of truth);
variant-specific spans (opener_span / prior_span) remain on kind
variants where meaningful.

**Finding 2 (PB-2 §6 obsolete sibling-lane assumption)**:
§6 line 198 said "PB-2 Tokenize | src/v3/std/tokenize.dag (NEW
per PB-2 L2.5)" — but tokenize.dag is ALREADY LIVE per
design-pure-bootstrap.md §"PB-2 — tokenize retire" (substantially
landed).

Fix: §6 prereq table row updated to reflect LIVE state. PB-2's
residual scope is scaffold-retirement (SG-1a + character-level +
codegen-driver per PB-2 L2.5 §1), not carrier authoring. PB-3
consumes the live Token carrier; carrier shape stable across
PB-2 residual-retirement timing.

Findings 3 + 4 already addressed:
- Finding 3 (pipeline.dag target): commit 015977347
- Finding 4 (Step 3b/4 atomic vs sequential): commit a1607a88d

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 APPROVE_WITH_COMMENTS — phantom "module-level metadata" removed

Cursor APPROVE_WITH_COMMENTS line:79: §4.1 said "SurfaceModule
(verified live) carries List<SurfaceItem> + module-level
metadata" — but live parse_surface.dag:29 has ONLY
`{ items: List<SurfaceItem> }`. No metadata fields.

Phantom addition violated INVARIANTS P1 live-state honesty.

Fix: §4.1 reframed to match live carrier exactly. Same
phantom-addition class as tokenize.rs phantom (commit 8ae37b4a5).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3085 high-level BLOCKING findings — AtomPayload 5-variant + diagnostic-table PROPOSED

Codex high-level BLOCKING (sha bf4d3152) — 2 substantive findings:

**Finding 1 (AtomPayload stale)**:
§2 + §5 cited infer.rs:10 top-comment "Atom(Identifier { name,
resolved })" — but live substrate.dag:87 has 5-variant
AtomPayload sum: Literal | UnresolvedIdentifier |
ResolvedByStructure | ResolvedByName | TypeParam.

infer.rs top-comment is STALE vs live substrate. My doc inherited
the drift.

Fix: §2 enumeration corrected to all 5 AtomPayload variants per
live substrate.dag:87.

**Finding 2 (diagnostic-table PROPOSED, not live)**:
§2 + §4.1 + §4.3 said "diagnostics.contains(port_id)
biconditional" as if live — but Dag at substrate.dag:525 has
ONLY { declarations, nodes, ports, clusters }. NO diagnostics
field. Diagnostic-table is PROPOSED substrate extension.

Fix: §2 fail-closed note flagged PROPOSED — Step 2 PR scope
includes `diagnostics: Map<PortId, Diagnostic>` field extension
to Dag, OR PB-Substrate prereq adds it before PB-5 dispatch.

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline (PR #3126 SurfaceModule analogous case); needed to
grep type Dag fields BEFORE claiming the diagnostics field
exists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3085 INLINE BLOCKING line:94 — §4.1 + §4.3 PROPOSED diagnostic-table marking

Cursor INLINE BLOCKING line:94: §4.1 + §4.3 said "diagnostics
coupled INTO InferredDag" without flagging the diagnostic-table
as a substrate extension. Live Dag at substrate.dag:525 has
{ declarations, nodes, ports, clusters } — NO diagnostics field.
Earlier commit c0d96af25 added PROPOSED marking only in §2;
§4.1 + §4.3 needed same treatment.

Fix: §4.1 + §4.3 reframed with explicit PROPOSED substrate-
extension marking + reference to §2 for extension scope.
Construction-time invariant + structural coupling are both
contingent on the substrate-extension landing (Step 2 PR scope
or PB-Substrate prereq).

Same recurring feedback_grep_carrier_field_before_coupling_claim
discipline applied to §4.1 + §4.3 consistently with §2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3077 high-level BLOCKING — Practice 4 classifications for PB-4 coproducts

Codex high-level BLOCKING (sha b812db91): "Diagnostic substrate
shape was repaired around anchoring but not re-audited as new
substrate type declarations" — meaning DiagnosticAnchor +
LowerDiagnostic + SurfaceFormRef + IdentifierRef need
🟢/🟡/🔴 classifications per modeling-discipline Practice 4
(Coproduct dissolution).

Same pattern as PB-5 fix in commit 040681f21 (AlgebraAxis +
InferDiagnostic) applied here.

Fix: added classifications + dissolution triggers:

- SurfaceFormRef 🟢 TERMINAL: closed-axis sum over live Surface*
  carriers; no further dissolution.
- IdentifierRef 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates full identifier-kind set against lower.rs
  identifier-resolution sites; promote to 🟢 TERMINAL when
  SurfaceVarRef + TypePathRef + ModulePathRef cover actual axes.
- LowerDiagnostic 🟡 SCAFFOLD: dissolution trigger = Step 2 brief
  enumerates variant set against lower.rs Diagnostic emission
  sites + Q6.5 anti-bridge preserved + PR #3077 §12 Q7
  ratification path.
- DiagnosticAnchor 🟢 TERMINAL: closed-axis covering all
  lowering-stage anchor kinds; no further dissolution.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3077 INLINE BLOCKING line:161 — DiagnosticSource Practice 4 classification

Cursor inline finding adds DiagnosticSource to Practice 4
classification scope. Earlier commit aba79a142 classified
SurfaceFormRef + IdentifierRef + LowerDiagnostic + DiagnosticAnchor
but missed DiagnosticSource.

Fix: DiagnosticSource 🟢 TERMINAL at pipeline-stage
discrimination scope. Closed-axis sum (Parse | Lower | Infer |
Emit); adding new pipeline stage requires explicit substrate-
extension audit per Practice 4 + stop-signal discipline (same
shape as PB-5 §3.2 TypeConnective stop-signal).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 INLINE BLOCKING line:151 — §11 SurfaceItem::Let + SurfaceLiteral span correction

Cursor INLINE BLOCKING line:151: §11 claimed "every Surface
variant carries SourceSpan" but live parse_surface.dag has:
- SurfaceItem::Let { name, type_ann, expr } — no direct span
- SurfaceLiteral = Int(String) | Bool(Bool) | String(String) —
  plain-tuple variants with no direct span

Source-span provenance for these cases is via enclosing carrier:
SurfaceLiteral wraps within `Literal { value, span }` at
parse_surface.dag:150. Let-item inherits container span.

INVARIANTS P2/P3 source provenance is structurally guaranteed
via direct-OR-enclosing carrier, but my "every variant"
overstatement obscured this.

Fix: §11 corrected to "most Surface variants carry SourceSpan
directly" + explicit Let + SurfaceLiteral exceptions noted +
Step 2 PR scope audits whether exceptions are structural-honest
(enclosing-carrier-provides-span) OR require substrate extension.

Same recurring overstatement class as earlier "every SurfaceItem
maps 1:1 to Declaration" (commit 61e2b67cd) — need to grep live
substrate variant fields BEFORE claiming uniform shape.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3126 INLINE BLOCKING line:32 — GrammarSpec is concept, not carrier

Earlier draft (now-merged PR #3126) framed `GrammarSpec` as a second declared
`.dag` input type alongside `List<Token>`. Verified via grep that no
`type GrammarSpec` exists anywhere under src/v3/ or dsl/ — the repo carries
only `parse_tables.dag`'s 6 SG-2c table-families. Per cursor 2026-05-14T23:30:04Z
inline finding, this violates INVARIANTS P2 (the Step-2 signature names a
carrier the substrate doesn't declare).

Reframed §3 (and downstream mentions in §3.2, §3 preamble line:33, §4.2
codex-correction recap line:165, §5.1 line:181, §12 Q6 line:344) so that:

- Parse has ONE input at the API boundary: `List<Token>`.
- "GrammarSpec" is a concept-level grouping for the 6 compile-time table-families
  in `parse_tables.dag`, not a substrate carrier and not a runtime parameter.
- Step-2 signature stays `fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
  per cfe842b2d (already applied in #3126); these edits remove the lingering
  "two input types" framing that contradicted that signature.

Per Decision 3.B (b) compile-time parser tables: substrate authority is
parse_tables.dag (6 table-families) consumed via direct table lookups inside
the parser body — no runtime grammar value, no `GrammarSpec` carrier needed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING — §4.3 + §16 internal contradictions resolved

Codex review id 12391 on sha cd6e8d15 flagged two live statements of the
already-rejected typed-state-with-coupled-diagnostics model still sitting
inside the parse L2.5, contradicting the corrected Result-sum framing in §4.2
and §6 Step 2:

1. §4.3 still proposed a `SurfaceModule { items, diagnostics }` extension
   modeled on PB-4/PB-5 patterns. Reframed: §4.3 now states explicitly that
   parse-stage uses Result-sum (no diagnostics field on SurfaceModule),
   restates the cross-stage discriminator (Result-sum for fail-fast output
   domains: parse + emit; typed-state for structural output domains: lower +
   infer), and cites `parse_generated.rs:138` as the live shape.

2. §16 "Memory disciplines applied" bullet read
   "feedback_state_space_vs_behavioral_invariants (typed-state SurfaceModule
   at output)". Rewritten to "parse output is `Result<SurfaceModule,
   ParseDiagnostic>` — the type rules out partial-parse states by
   construction; SurfaceModule itself carries no diagnostic field per §4.3".

Per `feedback_discipline_change_audit_all_contract_mentions`: when a
contract changes (here: SurfaceModule extension dropped in favor of
Result-sum), all §-internal restatements must be swept in the same diff
or they leak through as authoritative parallel claims.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3127 BLOCKING — drop TokenizedSource extension + cite concrete P5 receipts

Codex review id 12370 on sha d15e1f29 raised two load-bearing planning-shape
findings on the tokenize L2.5. Addressed in this fix-forward branch (PR #3138)
since PR #3127 is still open but accumulating cycles.

Finding 1 — parallel boundary carriers (INVARIANTS P2 / Modeling Practices 3+5):
§4.3 + §9 Step 2 + §12 Q6 named `TokenizedSource { tokens, diagnostics }` as the
output, while §2/§7.2/§7.3 named `List<Token>`. Same misclassification I just
removed from the parse L2.5 in PR #3138 §4.3: tokenize sits in the fail-fast
output domain alongside PB-3 parse + PB-6 emit (a partial token list with a
corrupt token in the middle is not a valid downstream input for parse), so the
failure couples via `Result`, not into the structural carrier.

Resolution:
- §4.3 rewritten to ratify `Result<List<Token>, TokenizeDiagnostic>` — the live
  `tokenize_generated.rs:96` shape — with no `TokenizedSource` extension.
- §9 Step 2 row signature updated to match; explicit "single canonical boundary
  carrier: List<Token> on the Ok branch" framing.
- §12 Q6 resolved REJECTED in-doc (no operator ratification needed; disposition
  follows from the cross-stage discriminator that's also load-bearing in PR
  #3138 parse L2.5).
- §16 memory-disciplines bullets rewritten parallel to PR #3138 parse §16:
  Result-sum, no diagnostics field on `List<Token>`.
- §14 "Surfaces awaiting" trimmed Q6 from the operator-ratification list.

Finding 2 — soft deferral of `regen_tokenize` retirement (INVARIANTS P5):
deferral previously named "PB-Bootstrap-Process lane scope" without a concrete
ROADMAP.md row. Updated §9 Step 2 + Step 4 rows to cite the named receipts:
- `docs/design-pure-bootstrap-zero.md:116` (PB-Bootstrap-Process lane: author
  bootstrap.dag + generated trampoline; sized M).
- `docs/design-pure-bootstrap-zero.md:118-123` (N=0 runtime verification gates).
- ROADMAP.md:467 (Character-level under-consumption in tokenize + syntax
  authorities — phase-2 char-class retype owns the codegen-driver path).
- ROADMAP.md:416 (Class 5 Gap 3 — top-level `ValueBody` boundary; gating
  substrate-capability for the phase-2 retype).
- ROADMAP.md:53 (T-PB-A — non-test census → 0 floor).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3126 BLOCKING line:119 — §6/§7.2/§15 Q7 sequencing resolved by PR #3077 merge

Cursor inline BLOCKING at docs/design-parse-stage-l25-model.md:119 (sha at
merge time of PR #3126) flagged a real internal contradiction:
- §6 (lines 191, 207, 308) claimed "Step 2 (pipeline-slot declaration) is
  unblocked"
- §7.2 (line 227) claimed "PR #3077 §12 Q7 must ratify before any Step 2
  worker brief authoring"

A worker reading the doc could land Step 2 (pipeline boundary) before the
diagnostic carrier's P2/P3 failure shape was fixed.

Resolution: PR #3077 (PB-4 lower L2.5) merged at 2026-05-15T00:21:19Z,
carrying the §12 Q7 ratification of the Decision 2.B per-stage diagnostic
extension path. The gate IS now satisfied at HEAD, so the resolution is
fact-update (annotate Q7 as DONE with the merge timestamp) rather than
retracting either §6 or §7.2.

Edits:
- §15 step 4: annotated "DONE 2026-05-15T00:21:19Z when PR #3077 merged"
  and added the explicit "Step 2 is now genuinely unblocked, not just
  procedurally next" framing so workers reading the sequence don't bypass
  the gate.
- §7.2 line 227: rewritten from "Q7 must ratify before Step 2 brief
  authoring" (future tense, the contradiction surface) to "Gate satisfied
  2026-05-15T00:21:19Z when PR #3077 merged; Step 2 worker brief authoring
  is unblocked at HEAD per §15 step 4." Cites the cursor finding as the
  resolution path.

§6 unblocking statements stay as-is — they were correct at HEAD; the
contradiction lived in §7.2's pre-merge framing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 16d21f4a) — Q7 cited at every Step 2 unblocking claim

Codex BLOCKING finding 2 (sha 16d21f4a, 216s thinking): the Q7 dependency
was recorded in §7.2/§15 but not in the other Step-2-unblocking sites at
§6 lines 191, 207, 308 — risking a worker reading "Step 2 unblocked" without
also reading the Q7 prerequisite.

Codex framing: "make Q7 a hard precondition wherever Step 2 is called
unblocked, or split Step 2 into pre-Q7 and post-Q7 scopes with separate
receipts."

Chose the first option since PR #3077 has already merged (2026-05-15T00:21:19Z)
and splitting into pre/post-Q7 scopes is no longer load-bearing. Annotated
all three §6 sites:
- Line 191 (Implication for PB-3 migration): cites gate + merge timestamp
  + explicit "must NOT be brief-authored before that merge timestamp."
- Line 207 (Critical observation): cites the Step 2 gate as PR #3077 §12 Q7
  + merge timestamp + P3 failure-shape consequence if violated.
- Line 308 (Director-recommend phase list): cites gate + §7.2/§15 step 4
  cross-refs + merge timestamp.

Codex BLOCKING finding 1 (GrammarSpec carrier non-existence) verified
already resolved at HEAD via commit c97dc15ae — every GrammarSpec mention
now explicitly marks it as a concept-not-carrier; Step 2 signature is
`fn parse(tokens: List<Token>) -> Result<SurfaceModule, ParseDiagnostic>`
with no GrammarSpec parameter.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3126 BLOCKING (sha 16d21f4a) — Step 3b/4 atomicity + Q1 ratification-pending status

openai-pro REQUEST_CHANGES on PR #3126 sha 16d21f4a flagged three findings.
Findings 1 + 2 (signature vs realization mismatch under LAYER MODEL + P2/P3
boundary discipline) were already resolved at HEAD by earlier fix-forward
commits — Step 2 signature is now `fn parse(tokens: List<Token>) ->
Result<SurfaceModule, ParseDiagnostic>` matching live `parse_generated.rs:138`
with no GrammarSpec parameter and no diagnostics-coupled SurfaceModule
extension (§4.3 Result-sum disposition).

Finding 6 (TRACKED vs UNTRACKED DEBT — Step 3b/4 same-PR vs two-PR):
§9 had both a "Step 3b/4 COMBINED" row (line 250) and a leftover separate
"Step 4: Parity test" row (line 251) — internally contradictory. §15 also
still sequenced Steps 3b + 4 as four separate authoring/dispatch/ratify
beats (steps 12-17), contradicting §12 Q5's "same PR" decision and the
explicit "Update to §15" note at §12 line 325.

Resolution: collapsed §9 to one COMBINED row absorbing the parity-TestClaim
mechanics + P5 dissolution receipt from the deleted Step 4 row; collapsed
§15 steps 12-17 into single COMBINED authoring + dispatch + ratify (steps
12-14). Added explicit `feedback_paper_shrink_variants` reasoning in both
sections.

Finding 2.5 (PM intent — substrate-capability bundled vs separate):
§12 Q1 line 288 said "Director-recommend: (b) bundled" while §13 line 340
listed substrate-capability as a non-goal of PB-3 and §15 step 11 had
"WAIT for substrate-capability landing" — three sections, two different
execution paths.

Resolution: annotated Q1 as "PENDING operator/PM ratification" with
explicit default-execution clause: until operator ratifies, the doc
treats substrate-capability as path (a) separate lane (matching §13 + §15
+ §9 row). If/when ratified to (b), §13 drops the non-goal and §15 step 11
collapses into the COMBINED Step 3b/4 brief.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3126 BLOCKING (sha 5619afac) — parse_tables.dag as single enumerated authority

Codex caught me copying the prose summary at parse_tables.dag:23-29 (which
enumerates only SG-2c-numbered families) instead of grepping the live
`^type [A-Z]` declarations. Result: SoftKeywordIdentRow (line 334) was
missing from §3.2 / §5.1 / §6 / §12 because it lacks an SG-2c-N number
in the prose summary.

Per `feedback_parallel_representation_debt`: structural fix is to stop
hand-enumerating in the doc — cite parse_tables.dag itself as the single
enumerated authority and use `type`-declaration line-anchors for the
worked example, not a hand-maintained count.

Edits:
- §3.2 §"Live substrate authority": replaced the SG-2c-numbered bullet
  list with `type`-declaration line-anchor enumeration including
  SoftKeywordIdentRow at parse_tables.dag:334 + the supporting enum
  BinaryOpLevel at line 133. Added codex-finding callout explaining the
  miss + the discipline shift.
- §3 preamble line:33, §3.2 line:45 callout, §3.2 line:57 framing,
  §3.2 §"Substrate authority" line:71, §5.1 line:171-180, §12 Q2 line:298,
  §12 Q6 line:334: all hardcoded "6 table-families" counts dropped; doc
  now points readers to §3.2 enumeration / `parse_tables.dag` directly.
- §5.1 sub-enumeration list (the parallel 6-item list at lines 173-178)
  deleted; replaced with redirect to §3.2 + restated 3a-vs-3b/4 scope split.

Code-level check before commit: `grep -nE '^type [A-Z]' src/v3/compiler/parse_tables.dag`
returns 7 types: BinaryOpLevel (133), BinaryOpRow (167), TopLevelItemKwRow (289),
SoftKeywordIdentRow (334), BracketRow (385), PrimaryPrefixRow (449),
PrimaryAtomRow (486). Doc enumeration matches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING (sha f08b9525) — bind TokenizeDiagnostic to live CharClass authority

Codex Finding 1 (sha f08b9525, 216s thinking): the §4.2 TokenizeDiagnostic
draft and the §5.1/§5.2 headings used `ScannerCharClass` / `ScannerClassRef`
— names that exist only as the *generated Rust enum spelling*, not as a
declared .dag substrate type. Verified via grep:

  grep -rn '^type CharClass\|^type ScannerC' dsl/ src/v3/

returns ONE authority: `dsl/std/unicode.dag:62`
  type CharClass = Whitespace | Digit | IdentStart | IdentContinue

consumed at `src/v3/compiler/tokenize.dag:103`
  data ascii_scan_order: List<CharClass> = [Whitespace, Digit, IdentStart, IdentContinue]

There is no `ScannerCharClass` declaration anywhere — that name was copied
from generated Rust without grep-verification, the same failure mode as
`feedback_grep_substrate_before_naming_ratification` (carrier-name
collision discipline).

Resolution:
- §4.2 TokenizeDiagnostic carrier: `expected_class: ScannerClassRef` →
  `expected_class: CharClass`, dropped the `type ScannerClassRef =
  ScannerCharClass` alias entirely; added a codex-finding callout citing
  the substrate authority + naming the failure mode.
- §5.1 heading "Byte → ScannerCharClass dispatch" → "Byte → CharClass
  dispatch"; bullets unchanged; added line-anchor cites for the substrate
  authority + explicit "NOT ScannerCharClass" disclaimer.
- §5.2 heading "ScannerCharClass → token-recognition state machine" →
  "CharClass → token-recognition state machine".

Finding 2 (TokenizedSource not reconciled with parse input contract): no
new fix required — already resolved by commit 6adb99227 (TokenizedSource
extension dropped entirely; tokenize uses Result<List<Token>,
TokenizeDiagnostic>; List<Token> is the single canonical boundary carrier
consumed by parse). Codex was reviewing sha f08b9525, which predated the
TokenizedSource drop.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3138 INLINE BLOCKING line:188 — clarify §7.3 chain is Ok-branch propagation

Cursor INLINE at sha f08b9525 worried that §7.3's cross-stage chain
"tokenize → List<Token> → parse → ..." was inconsistent with §4.3's
TokenizedSource carrier (diagnostics not flowing forward). At HEAD the
TokenizedSource extension is dropped (commit 6adb99227); tokenize uses
Result<List<Token>, TokenizeDiagnostic>, so List<Token> IS the canonical
Ok-branch payload that flows forward and Err branches terminate the
pipeline fail-fast.

To make this explicit at §7.3 (instead of leaving readers to infer it
from §4.3), annotated the chain with:
- "Ok-branch propagation; Err branches are stage-terminal fail-fast per
  §4.3 Result-sum discriminator" framing prefix.
- Per-stage Result/typed-state annotations: tokenize/parse show
  Result<Ok, Err>; lower/infer show typed-state structural-output;
  emit shows Result<EmittedArtifact, EmissionDiagnostic>.
- Explicit "on any stage's Err branch the pipeline aborts at that stage
  (no partial-output propagation across boundaries)" trailer.

This makes the chain self-consistent vis-a-vis §4.3 without requiring
the reader to walk back-and-forth, and prevents future readers from
re-introducing a TokenizedSource-shaped extension to "make diagnostics
flow forward" — they already do, just via the Err branch terminating
the pipeline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): openai-pro PR #3127 BLOCKING (sha d15e1f29) — character-level scaffold dissolution trigger

openai-pro REQUEST_CHANGES on sha d15e1f29: §1 line:26 character-level
under-consumption scaffold named the problem but lacked a checkable
dissolution trigger. SG-1a scaffold above had the right shape — "once
those bodies lower structurally under compile_to_dag, delete the raw-text
extractor + derive directly from lowered Dag in same PR." Character-level
scaffold just said "PB-2 Step 4 carries this scope" — a lane assignment,
not a trigger.

Resolution: rewrote §1 item 2 with the same SG-1a-shape trigger structure:
- Substrate-consumption condition (a): scanner classes / string escape /
  local punctuation retype to `dsl/std/unicode.dag` `CharClass` /
  `char_in_class` (concrete field retypes named:
  `StringEscapeSpec.suffix: Char`, `LocalPunctSpec.pattern: List<Char>`,
  `string_literal_delimiter: Char`).
- Codegen-driver condition (b): `tokenize_generated.rs` no longer emits
  hidden `byte.is_ascii_*` predicates because the driver reads class
  facts structurally from lowered `tokenize.dag`.
- Same-PR dissolution: delete the parallel character-predicate scaffold
  in the same PR that flips substrate consumption — no Rust-and-`.dag`
  coexistence per `feedback_paper_shrink_variants`.
- Cross-ref to §9 Step 4 gating prereqs: ROADMAP.md:467 + ROADMAP.md:416
  Class 5 Gap 3 + std.unicode bootstrap/load-set decision (already cited
  in §9 from earlier commit 6adb99227).

Per openai-pro's framing: "small fix — mirror the SG-1a scaffold wording
by naming the exact substrate-consumption condition and same-PR deletion
receipt for the hidden Rust character predicates."

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* WIP: gunbc Director

* docs(r3): cursor PR #3127 APPROVE_WITH_COMMENTS — §14 + §15 Q6-already-rejected sweep

Cursor review id 12405 (APPROVE_WITH_COMMENTS) caught the same
`feedback_discipline_change_audit_all_contract_mentions` failure mode
recurring: §12 Q6 was resolved REJECTED in commit 6adb99227, and §14
"Surfaces awaiting" + §12 Q6 heading + §9 Step 2 row were updated, but
two §-internal contract restatements were missed:

- §14 acceptance criterion 11: "Operator/PM ratification on §12 Q1-Q6"
- §15 step 1: "Operator / PM-delegate ratifies §12 Q1-Q6"

Both contradicted §12 Q6 + §14 "Surfaces awaiting" (which already said
"Q1-Q5 only"). A worker reading §14/§15 could schedule sign-offs on Q6
after it was already resolved-rejected elsewhere.

Resolution: both sites now say "Q1–Q5" with the explicit Q6-rejected
crossref + "see §14/§15 for same scoping" pointer at the §14 criterion
so the three sections agree internally.

Cursor verdict was APPROVE_WITH_COMMENTS (substantive APPROVE — "fix the
checklist/sequence so every section agrees Q6 is closed"); the
exploratory volatile-line-anchor note is harmless and out-of-scope for
this PR.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3138 BLOCKING (sha 887c69671) — 3 findings swept across 3 L2.5 docs

Codex review id (sha 887c69671, 333s thinking) flagged three findings,
all `feedback_discipline_change_audit_all_contract_mentions` recurrences
where partial sweeps left §-internal contradictions.

Finding 1 — infer doc cross-stage trigger leak:
§4.2 InferDiagnostic dissolution trigger said "when Step 2 worker brief
enumerates the full variant set against `parse_generated.rs`
Diagnostic::ParseError, lower.rs Diagnostic construction sites, and
infer.rs Dag::mark_unresolved emission sites." That cross-stage trigger
surface is wrong for the infer-specific scaffold. Narrowed to infer-only:
"against `src/v3/compiler/src/infer.rs` Diagnostic construction sites +
`Dag::mark_unresolved` emission sites (infer-stage only)." Parse and lower
have their own per-stage carriers + own Q7 mapping; this doc no longer
reaches into their dissolution-trigger surface.

Finding 2 — parse + tokenize coproduct receipts:
`ParseDiagnosticKind` (parse §4.2) and `TokenizeDiagnostic` (tokenize §4.2)
sums were declared without the 🟡 SCAFFOLD / 🟢 TERMINAL classification +
dissolution trigger that the infer doc carries (per modeling-discipline
Practice 4 + `feedback_coproduct_dissolution`). Added matching scaffold
receipts to both: 🟡 SCAFFOLD at PROPOSED stage with stage-local dissolution
trigger (Step 2 brief enumerates against parse_generated.rs / tokenize_generated.rs
Diagnostic::* construction sites — stage-only, not cross-stage); promote
to 🟢 TERMINAL when full variant set lands. Both also carry the anti-bridge
note per Q6.5.

Finding 3 — tokenize Q7 reconciliation:
The parse doc has the "Q7 DONE 2026-05-15T00:21:19Z" annotation at §15 step 4
(commit f85fa1f6b) but the tokenize doc still framed Q7 as a pending lane
dependency at §4.2 line:113 ("Lane dependency"), §15 step 3, and §14
"Surfaces awaiting". Annotated all three:
- §4.2 "Lane dependency": Q7 DONE timestamp + "TokenizeDiagnostic per-stage
  variant authoring is the remaining lane work."
- §15 step 3: full DONE annotation matching parse doc §15 step 4 shape +
  cross-ref to the parse doc.
- §14 "Surfaces awaiting": strikethrough + DONE annotation.

PR #3127 also carries the tokenize doc; same edits will port there in a
companion commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3138 INLINE BLOCKING line:372 — §16 Surfaces awaiting Q7 contradiction

Cursor inline at sha 887c69671 caught the symmetric finding to the codex
5-finding BLOCKING #3: parse-doc §15 step 4 already annotated Q7 as DONE
(commit f85fa1f6b) but parse-doc §16 "Surfaces awaiting" still listed Q7
as pending. Same `feedback_discipline_change_audit_all_contract_mentions`
sweep failure — the tokenize doc had three sites carrying the pending
framing (commit e9739ea4f fixed those) but the parse doc's §16 site was
missed in the original Q7-DONE sweep.

Resolution: §16 bullet now strikes through + "DONE 2026-05-15T00:21:19Z
(PR #3077 merged carrying Q7 ratification; see §15 step 4)" — same shape
as the tokenize doc's §14 "Surfaces awaiting" Q7-DONE annotation landed
in e9739ea4f.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): codex PR #3127 BLOCKING (sha 8e682eef) — §3.1 source identity + §4.2 live-vs-proposed clarity

PR #3127 was merged at 02:04:22Z but codex's later review on sha 8e682eef
caught two findings that still live in main. Both addressed in this PR
(PR #3140) since the tokenize doc lives on main and this is the open
post-merge cleanup branch.

Finding 1 — §3.1 source-identity drop:
Earlier draft framed tokenize input as bare `String → List<Token>`. That
contradicts live `tokenize_generated.rs:96`:
  pub fn tokenize(source: &str, file: &str) -> Result<Vec<Token>, Diagnostic>
The `file` parameter is load-bearing because `SourceSpan` requires a
file/source-id field for byte ranges to be attributable; without it the
Token + Diagnostic span fields would have to fabricate source-id at the
pipeline boundary (P3 fail-closed violation). §3.1 now ratifies the live
two-input shape:
- `source: String` — UTF-8 source text (primitive).
- `file: SourceFileId` — source-identity carrier; Step 2 brief ratifies
  the appropriate `.dag` shape (NonEmptyStr newtype or richer sum if
  multiple source-class kinds).

Finding 2 — §4.2 live-vs-proposed blur:
Prior framing said "TokenizeDiagnostic (substrate extension per Decision
2.B)" without making clear that this is a PROPOSED per-stage carrier,
not the live one. The live carrier at `diagnostics.dag:150` is generic
`Diagnostic { kind: AnyDiagnosticKind, ... }` with kind sum at `:139-142`
discriminating CompilerKind vs LensInstanceKind; tokenize emits today
via `Diagnostic::TokenizerError`-shaped sites carrying
`kind: CompilerKind(...)`. §4.2 heading now reads "PROPOSED — NOT yet
live" + a codex-finding callout explicitly distinguishing the live
carrier from the proposed per-stage refinement + Q7 status DONE.

Both fixes preserve the e9739ea4f 🟡 SCAFFOLD coproduct receipt that
was already addressing codex sha 887c69671 Finding 2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3127 INLINE BLOCKING line:121 — §4.3 + §9 Step 2 signature sweep

Cursor inline at sha 8e682eef line:121 (the §4.3 Step 2 signature site)
caught the same source-identity drop my §3.1 fix already addressed at the
input-type framing — but two companion sites still had the single-input
signature:
- §4.3 line:151 Step 2 signature recap
- §9 Step 2 row (4-step migration table)

Both now read:
  fn tokenize(source: String, file: SourceFileId) -> Result<List<Token>, TokenizeDiagnostic>

matching the §3.1 source-identity discipline + live `tokenize_generated.rs:96`
`pub fn tokenize(source: &str, file: &str)` shape.

Same `feedback_discipline_change_audit_all_contract_mentions` recurrence
— my Finding 1 fix at §3.1 (commit 18af49fad) didn't sweep the §4.3 + §9
contract-signature restatements.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3127 INLINE BLOCKING line:166 — §6 row TokenizeDiagnostic LIVE/PROPOSED clarity

Cursor inline caught ambiguity in the §6 prereq-table row for
"TokenizeDiagnostic substrate extension." Old wording in column 4 read
"Carrier LIVE; per-stage variant authoring NEW" which could be parsed
as "TokenizeDiagnostic carrier itself is LIVE" — but no `type
TokenizeDiagnostic` exists in `dsl/std/` or `src/v3/std/`. Only the
generic `type Diagnostic` at `diagnostics.dag:150` is live.

Cursor's secondary claim that "line 150 is EmissionDiagnostic receipt
prose" is INCORRECT — verified via grep + read: `diagnostics.dag:150`
is the `type Diagnostic { kind: AnyDiagnosticKind, ... }` declaration
header. `type EmissionDiagnostic` lives separately at `diagnostics.dag:201`.
The §6 row's citation of `:150` for the underlying Diagnostic carrier
is correct.

But the ambiguity-in-wording finding stands. Reframed the row:
- Column 1: explicitly "PROPOSED — no `type TokenizeDiagnostic` exists yet"
- Column 2: clarifies the LIVE underlying carrier is `type Diagnostic`
  with `kind: AnyDiagnosticKind` sum at `:139-142`; cites Q7 ratification
  done timestamp.
- Column 4: explicit "Underlying Diagnostic carrier LIVE at HEAD;
  per-stage TokenizeDiagnostic variant authoring is NEW work, NOT live
  substrate" with cursor-finding callout.

Consistent with §4.2's "PROPOSED — NOT yet live" framing from commit
18af49fad.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3141 INLINE BLOCKING line:68 — §2 structural-signature sweep

Cursor inline caught §2 still framing tokenize structural signature as
`String → List<Token>` — single bare-string input, no source-identity,
no Result-sum failure shape — while §3.1 + §4.3 + §9 Step 2 row all
ratify the live two-input + fail-fast Result-sum shape.

Same `feedback_discipline_change_audit_all_contract_mentions` recurrence
that's been compounding across this fix-forward sequence: each finding
fix needs a doc-wide signature-restatement sweep, not just the §-local
sentence.

Resolution: §2 structural-shape sentence now reads
  `(String, SourceFileId) → Result<List<Token>, TokenizeDiagnostic>`
matching §3.1 (source-identity), §4.3 (Result-sum cross-stage
discriminator), §9 Step 2 row, and live `tokenize_generated.rs:96`
`pub fn tokenize(source: &str, file: &str) -> Result<Vec<Token>,
Diagnostic>`. Cross-refs added to §3.1 + §4.3 to make the discipline
visible at the high-level framing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): cursor PR #3141 INLINE BLOCKING line:73 — replace invented SourceFileId with live FilePath

Cursor caught me inventing `SourceFileId` as the Step 2 boundary type
without a live `.dag` declaration. Verified via grep:
  grep -rn '^type SourceFileId' dsl/ src/v3/
returns empty. Same `feedback_grep_substrate_before_naming_ratification`
family error as the earlier `ScannerCharClass` miss — naming a substrate
carrier without grep-verification.

Live carrier already exists at `dsl/std/types.dag:276`:
  type FilePath = String where non_empty

referenced by `type SourceSpan { file: FilePath, ... }` at
`dsl/std/types.dag:293`. So source identity flows through every Token +
Diagnostic span field via the live `FilePath` carrier — no new carrier
authoring needed at Step 2.

Edits:
- §2 structural-signature: `(String, SourceFileId)` → `(String, FilePath)`
- §3.1 source-identity bullet: replaces fictional SourceFileId with the
  live `FilePath` declaration cite (dsl/std/types.dag:276), points out
  that `SourceSpan.file` already references this same carrier, and adds
  a cursor-finding callout naming the failure mode.
- §4.3 + §9 Step 2 signature restatements: `file: SourceFileId` →
  `file: FilePath`.

Single substrate authority (INVARIANTS P1/P2) restored: FilePath is the
one carrier for source identity across SourceSpan + tokenize input +
Token construction.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant