diff --git a/docs/design-infer-stage-l25-model.md b/docs/design-infer-stage-l25-model.md index 3197721b107..ccba337846b 100644 --- a/docs/design-infer-stage-l25-model.md +++ b/docs/design-infer-stage-l25-model.md @@ -36,10 +36,15 @@ This is structurally distinct from PB-4 lower: Per the live `infer.rs:5-15` rules: - `Arrow { inputs, output, body }` → direct signature inference from input/output declarations -- `Atom(Identifier { name, resolved })` → follow resolved link OR look up name in declaration table (§8.9 inhabitance walk) +- `Atom(AtomPayload)` per LIVE substrate.dag:87 5-variant sum (per codex PR #3085 BLOCKING finding 1 — earlier "Atom(Identifier { name, resolved })" was stale; infer.rs:10 top-comment drifted vs live substrate): + - `Literal(LiteralBits)` → literal-typed port; no resolution needed + - `UnresolvedIdentifier(String)` → fail-closed Unresolved + diagnostic at this port + - `ResolvedByStructure(DeclarationId)` → follow chain to declaration + - `ResolvedByName(DeclarationId)` → follow chain to declaration (§8.9 inhabitance walk) + - `TypeParam(String)` → type-parameter (post-instantiation resolved or unresolved) - Other TypeConnective variants → not callable; produce Unresolved + diagnostic -**Fail-closed (INVARIANTS C-8)**: every detectable problem routes through `Dag::mark_unresolved` (current Rust API) or equivalent `.dag` substrate operation. Post-infer invariant: `state != Uninferred for all ports` AND `state == Unresolved iff diagnostics.contains(port_id)`. +**Fail-closed (INVARIANTS C-8)**: every detectable problem routes through `Dag::mark_unresolved` (current Rust API) or equivalent `.dag` substrate operation. Post-infer invariant: `state != Uninferred for all ports` AND `state == Unresolved iff diagnostics.contains(port_id)` — where the diagnostic-table is a PROPOSED substrate extension to `Dag` carrier (live `Dag { declarations, nodes, ports, clusters }` at substrate.dag:525 has NO diagnostics field per codex PR #3085 BLOCKING finding 2). Step 2 PR scope includes `diagnostics: Map` field extension to Dag, OR PB-Substrate prereq adds it before PB-5 dispatch. **Typed-state carrier per Decision 3.A operator-ratified shape** (sum-variant `Dag = PreInferDag | InferredDag` per `feedback_coproduct_dissolution` Practice 4): infer's signature is `fn infer(d: PreInferDag) -> InferredDag` — the variant transition IS the structural invariant of inference completion. Pre-infer state cannot exist in `InferredDag` by construction. @@ -63,7 +68,7 @@ Per `feedback_lenses_not_passes` + observation in §2: infer's "rule book" is em - Lower maps SURFACE FORMS (which are user-authored grammar) → substrate behaviors. The mapping is a design choice (multiple valid mappings possible per grammar/elaboration separation per Decision 2.A 4-param compile). - Infer propagates TYPES through substrate-declared algebraic structure. The propagation rule per TypeConnective variant is determined by the variant's algebraic role (Arrow's direct-signature, Atom's lookup, etc.). The rule is forced by the substrate's structure; no design freedom for re-mapping. -If a future substrate refactor adds new TypeConnective variants, the inference rules extend automatically — they're per-variant structural facts, not pluggable rules. +**If a future substrate refactor adds new TypeConnective variants**, the inference rules do NOT extend automatically. Per cursor INLINE BLOCKING #3085 + thesis stop-signal discipline: a 7th TypeConnective variant requires (a) explicit C1 substrate-extension audit + (b) named infer-rule receipt for the new variant's structural inference behavior. The "per-variant structural facts" framing means new variants need new per-variant facts, NOT silent inheritance. Earlier draft "rules extend automatically" weakened the substrate-extension stop signal; corrected here. --- @@ -71,7 +76,7 @@ If a future substrate refactor adds new TypeConnective variants, the inference r ### §4.1 `InferredDag` (typed-state output carrier per Decision 3.A) -Per Decision 3.A operator-ratified, infer produces `InferredDag` variant explicitly. Construction-time invariant: every `Port.state` is either `Resolved(TypeShape)` or `Unresolved` (with diagnostic in the diagnostic table per the `state == Unresolved iff diagnostics.contains(port_id)` biconditional). `Uninferred` cannot exist in `InferredDag` by construction. +Per Decision 3.A operator-ratified, infer produces `InferredDag` variant explicitly. Construction-time invariant: every `Port.state` is either `Resolved(TypeShape)` or `Unresolved` (with diagnostic in the PROPOSED diagnostic-table substrate extension — live `Dag { declarations, nodes, ports, clusters }` at `src/v3/std/substrate.dag:525` has NO diagnostics field; Step 2 PR scope includes `diagnostics: Map` field extension OR PB-Substrate prereq). The biconditional `state == Unresolved iff diagnostics.contains(port_id)` is the construction-time enforcement target, not a live invariant. `Uninferred` cannot exist in `InferredDag` by construction (independent of the diagnostic-table substrate extension). **Substrate authority — DEPENDS on Decision 3.A landing**: extension of `src/v3/std/substrate.dag` `Dag` declaration to sum-variant; `InferredDag` variant constructor enforces port-state invariant. PB-Substrate work. @@ -109,6 +114,21 @@ type IdentifierRef // String-wrapper masquerading as typed; that contradicts the typed-carrier // discipline + creates a second authority on algebra identity. Closed-axis // enum is the structurally honest form): +// 🟡 SCAFFOLD coproduct at PROPOSED stage. Per cursor PR #3085 INLINE +// BLOCKING + modeling-discipline Practice 4: substrate coproducts require +// 🟢/🟡/🔴 classification + named ledger/trigger for dissolution. +// +// **Dissolution trigger**: when Step 2 worker brief enumerates the full +// algebra-axiom set against infer.rs algebra-inhabitance check sites at +// `src/v3/compiler/src/infer.rs` AND verifies coverage parity with the +// 3-variant subset already live at verification.dag:146 AlgebraicLawKind, +// promote to 🟢 TERMINAL at the per-stage algebra-axis scope. +// +// **Adjacent live**: verification.dag:146 AlgebraicLawKind already declares +// 3-variant subset (Associativity / Commutativity / Identity) at the +// algebraic-law-kind scope. AlgebraAxis is a broader closed-axis enumeration +// covering the algebra-inhabitance failure axes specifically (Closure / +// Inverse / Distributivity / OrderingTotality / etc.). type AlgebraAxis = Closure | Associativity @@ -119,9 +139,26 @@ type AlgebraAxis | OrderingTotality | OrderingTransitivity | OrderingAntisymmetry - // Step 2 brief enumerates the full closed set against infer.rs algebra-inhabitance check sites; - // adjacent to verification.dag:146 `AlgebraicLawKind` which has the 3-variant subset already live - + // Variants enumerated per Step 2 brief grep of infer.rs algebra-inhabitance + // check sites; promote to 🟢 TERMINAL once full set covers actual check axes. + +// 🟡 SCAFFOLD coproduct at PROPOSED stage. Per cursor PR #3085 INLINE +// BLOCKING + modeling-discipline Practice 4 (Coproduct dissolution): +// substrate coproducts require 🟢/🟡/🔴 classification + dissolution trigger. +// +// **Dissolution trigger**: when Step 2 worker brief enumerates the full +// variant set against `parse_generated.rs` Diagnostic::ParseError, lower.rs +// Diagnostic construction sites, and infer.rs Dag::mark_unresolved emission +// sites — promote to 🟢 TERMINAL at the per-stage diagnostic-variant scope. +// PR #3077 §12 Q7 ratification on Decision 2.B extension path determines +// whether this stays per-stage sum (option b) or extends shared Diagnostic +// (option a). +// +// **Anti-bridge**: per Q6.5 anti-bridge invariant at diagnostics.dag:135-141, +// InferDiagnostic does NOT collapse into CompilerDiagnosticKind without +// substrate-extension ratification; the relationship between this proposed +// per-stage sum + the shared Diagnostic carrier is itself the ratification +// scope of PR #3077 §12 Q7. type InferDiagnostic = UnresolvedIdentifier { identifier: IdentifierRef, scope: SectionRef } | NotCallable { type_connective: TypeConnective } @@ -136,9 +173,9 @@ type InferDiagnostic ### §4.3 `InferResult` — NOT a separate sum-variant -Unlike lower's `LowerResult = Either>`, infer's output is plain `InferredDag` — diagnostics are coupled INTO the InferredDag via the `state == Unresolved iff diagnostics.contains(port_id)` biconditional. Per `feedback_state_space_vs_behavioral_invariants`: the structural coupling makes the diagnostic-port relationship a type invariant, not a sum-variant. +Unlike lower's `LowerResult = Either>`, infer's output is plain `InferredDag` — diagnostics couple INTO the InferredDag via the `state == Unresolved iff diagnostics.contains(port_id)` biconditional, where the diagnostic-table is a PROPOSED substrate extension (live Dag at substrate.dag:525 lacks diagnostics field; see §2 fail-closed note for extension scope). Per `feedback_state_space_vs_behavioral_invariants`: the structural coupling — once the substrate extension lands — makes the diagnostic-port relationship a type invariant, not a sum-variant. -`InferredDag` is honest about partial inference success — ports that failed inference are `Unresolved` with explanation in the diagnostic table; ports that succeeded are `Resolved(TypeShape)`. The InferredDag is well-formed even with partial-failure; downstream consumers (emit) handle `Unresolved` ports per their own fail-closed discipline. +`InferredDag` is honest about partial inference success — ports that failed inference are `Unresolved` with explanation in the PROPOSED diagnostic table; ports that succeeded are `Resolved(TypeShape)`. The InferredDag is well-formed even with partial-failure (post substrate extension); downstream consumers (emit) handle `Unresolved` ports per their own fail-closed discipline. --- @@ -230,7 +267,7 @@ infer is target-agnostic structural inference; Shape A/B disambiguation is emit' | Step | Deliverable | Owner | Substrate | |---|---|---|---| | **Step 1: Model review** | THIS DOC | Director (zesty-bear-812) | docs/design-infer-stage-l25-model.md (this doc) | -| **Step 2: Pipeline slot** | `fn infer(d: PreInferDag) -> InferredDag` declared in compiler.dag with `ExternalRealization` body (Rust-backed placeholder pointing to current `infer.rs`). Signature uses Decision 3.A sum-variant typed-state at both boundaries (input = PreInferDag, output = InferredDag); Modeling Practice 6 API-level enforcement. Step 2 worker brief authoring routes through Director after Decision 3.A operator ratification + PB-Substrate carrier extension. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | compiler.dag refinement | +| **Step 2: Pipeline slot** | `fn infer(d: PreInferDag) -> InferredDag` declared in `src/v3/compiler/pipeline.dag` (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag) with `ExternalRealization` body (Rust-backed placeholder pointing to current `infer.rs`). Signature uses Decision 3.A sum-variant typed-state at both boundaries (input = PreInferDag, output = InferredDag); Modeling Practice 6 API-level enforcement. Step 2 worker brief authoring routes through Director after Decision 3.A operator ratification + PB-Substrate carrier extension. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | pipeline.dag refinement | | **Step 3: Implementation** | `src/v3/std/infer.dag` (the .dag implementation; structural TypeConnective dispatch + forward propagation). NO separate "InferenceSpec" file — substrate IS the rule book per §3.2. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 3 brief | src/v3/std/infer.dag (NEW substrate authority) | | **Step 4: Parity test + simultaneous Rust deletion** | Parity verification authored as `.dag` TestClaim — generated test fixture set + `.dag` TestClaim asserting `infer_via_rust(pre_infer_dag) == infer_via_dag(pre_infer_dag)` structural-equality across canonical corpus (Dag comparison up to NodeId renaming + diagnostic table equality). **P5 dissolution receipt**: this TestClaim is transient-by-construction; dissolves when infer.rs deletes in same PR. Any hand-Rust scaffolding bears P5 receipt `parity_infer_dag_vs_rust_scaffolding — transient; dissolves with infer.rs deletion in same PR per Step 4 atomic discipline`. `infer.rs` DELETED in same PR. EXPECTED_HAND_AUTHORED_NON_TEST shrinks by N entries at PR-merge. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 4 brief | tests/parity_infer_dag_vs_rust (TestClaim shape) + `infer.rs` deletion | diff --git a/docs/design-lower-stage-l25-model.md b/docs/design-lower-stage-l25-model.md index 440101dc697..61024a5a7f4 100644 --- a/docs/design-lower-stage-l25-model.md +++ b/docs/design-lower-stage-l25-model.md @@ -54,7 +54,17 @@ The raw surface-form representation produced by parse stage. Tree of `SurfaceIte ### §3.2 `ElaborationSpec` (surface-form → substrate-behavior rules per Decision 3.C) -Per Decision 3.C operator-ratified: `.dag` rules consumed by lower, NOT Rust code reading `.dag`. ElaborationSpec is the declared authority that maps surface forms to substrate behaviors. Each rule is a structural fact mapping a `SurfaceExpr` / `SurfaceItem` / `SurfacePattern` variant to a `Behavior` construction recipe. +Per Decision 3.C operator-ratified: `.dag` rules consumed by lower, NOT Rust code reading `.dag`. ElaborationSpec is the declared authority for **ALL lowering decisions** (per cursor INLINE BLOCKING PR #3077 line:33 — earlier draft "maps surface forms to behaviors" was too narrow; lower constructs Declarations / TypeConnectives / BranchPatterns / Bindings beyond just Behavior recipes). + +**Full ElaborationSpec scope** covers all surface-to-substrate mapping decisions: + +1. **SurfaceItem → Declaration recipes**: per-variant rules (Fn / FnExternalBody / Data / TypeAtom / TypeRecord) for Declaration placeholder shape + Let/Module/Import skip-allocation rules per §5.1 +2. **SurfaceType → TypeConnective recipes**: per-variant rules for Atom / Arrow / Compose / Disj construction +3. **SurfaceExpr → Behavior recipes**: per-variant rules for Value / Transform / Branch / Loop / Bind construction +4. **SurfacePattern → BranchPattern recipes**: per-variant rules for ResolvedVariant / UnresolvedVariant / record-pattern construction (consumed by Branch nodes) +5. **Binding-site rules**: how Bind's params + result_port are constructed from Fn item params + body return-port + +Each rule is a structural fact mapping a surface-form variant to a substrate-construction recipe. ElaborationSpec is single-authority for the surface→substrate mapping across ALL these axes; no axis lives in implementation-tier hand-Rust. Per the live top-comment at `lower.rs:8-21`, the canonical lowering rules currently in hand-Rust: - `SurfaceLiteral::{Int, Bool, String}` → `Value(LiteralBits::*)` @@ -102,6 +112,11 @@ Per Decision 3.A operator-ratified (sum-variant `Dag = PreInferDag | InferredDag ``` // Per Decision 2.B substrate extension (NEW; NOT currently live): +// 🟢 TERMINAL at the pipeline-stage discrimination scope. Closed-axis +// sum over all pipeline stages; no further dissolution. Per Practice 4 +// (Coproduct dissolution) — adding a new pipeline stage requires +// explicit substrate-extension audit + ratification (analogous to +// adding new TypeConnective variants per PB-5 §3.2 stop-signal). type DiagnosticSource = Parse | Lower | Infer | Emit type Diagnostic { kind: AnyDiagnosticKind // existing — kind-layer discrimination @@ -119,7 +134,10 @@ type Diagnostic { // stringifying a closed-axis variant boundary would be a typed-carrier // regression. `reason: String` is intentional human display detail only. // -// Typed surface-form reference for unsupported-form diagnostic: +// 🟢 TERMINAL at the surface-form-reference scope. Closed-axis sum over +// live Surface* carriers (parse_surface.dag); no further dissolution. +// Each variant wraps a typed substrate carrier directly. Per cursor + +// codex PR #3077 PR #3085 Practice 4 discipline (Coproduct dissolution). type SurfaceFormRef = ExprForm(SurfaceExpr) | ItemForm(SurfaceItem) @@ -127,12 +145,21 @@ type SurfaceFormRef | TypeForm(SurfaceType) | LiteralForm(SurfaceLiteral) -// Typed identifier reference for resolve errors (identifier-as-typed-fact, -// not stringified). PR #3077 BLOCKING applies same discipline. +// 🟡 SCAFFOLD at PROPOSED stage. Dissolution trigger: Step 2 brief +// enumerates the full identifier-kind set against lower.rs identifier- +// resolution sites; promote to 🟢 TERMINAL when SurfaceVarRef + +// TypePathRef + ModulePathRef cover the actual axes. Adjacent live +// substrate: none (this is a NEW carrier). type IdentifierRef = SurfaceVarRef { name: NonEmptyStr, span: SourceSpan } // Future: TypePathRef, ModulePathRef per Step 2 brief authoring +// 🟡 SCAFFOLD at PROPOSED stage. Dissolution trigger: Step 2 brief +// enumerates the full variant set against lower.rs Diagnostic emission +// sites; promote to 🟢 TERMINAL at per-stage diagnostic-variant scope. +// Anti-bridge per Q6.5 (diagnostics.dag:135-141): does NOT collapse +// into CompilerDiagnosticKind without PR #3077 §12 Q7 ratification on +// Decision 2.B extension path. type LowerDiagnostic = ResolveError { identifier: IdentifierRef, scope_chain: List } | UnsupportedSurfaceForm { form: SurfaceFormRef, reason: String } // form is typed; reason is human display @@ -156,7 +183,12 @@ Earlier draft proposed `LowerResult = Either> **Corrected diagnostic-anchor model** (typed at the same level as diagnostic variants): ``` -// Typed diagnostic anchor — closed-axis sum over anchor kinds: +// 🟢 TERMINAL at the diagnostic-anchor scope. Closed-axis sum +// covering all lowering-stage diagnostic anchor kinds (port-level / +// declaration-level / field-level / surface-form-level). No further +// dissolution — each variant carries a typed reference to the +// associated substrate object. Per cursor + codex PR #3077 Practice 4 +// discipline (Coproduct dissolution). type DiagnosticAnchor = PortAnchor { port_id: PortId } | DeclarationAnchor { declaration_id: DeclarationId } @@ -189,9 +221,16 @@ lower composes from these substrate facts via 2-pass walk: Per `lower.rs:3-7` top-comment: "Pass 1 walks all top-level items and allocates placeholder Declarations for each named type/fn, populating a symbol table (name → DeclarationId)." -Pass 1 is purely structural: SurfaceItem variants map 1:1 to Declaration placeholders. No resolution happens; identifiers are not yet bound. Output: symbol table (name → DeclarationId) + `PreInferDag` with placeholder declarations. +Pass 1 is purely structural: **NAMED-DECLARATION SurfaceItem variants** map to Declaration placeholders. Per live `lower.rs:2950-2958` (verified per cursor BLOCKING #3077): + +- **Allocate declaration**: `SurfaceItem::Fn` / `FnExternalBody` / `Data` / `TypeAtom` / `TypeRecord` (named type / fn / data declarations get top-level Declaration placeholders) +- **Skip allocation** (continue in collect_symbols): `SurfaceItem::Let { .. }` / `SurfaceItem::Module { .. }` / `SurfaceItem::Import { .. }` — these are parsed facts that flow forward but have NO top-level Declaration at M1(2.7). Let-bodies lower to Bind expressions in Pass 2; Module/Import items are parsed-facts preserved but un-declared. + +No resolution happens in Pass 1; identifiers are not yet bound. Output: symbol table (name → DeclarationId) + `PreInferDag` with placeholder declarations for named-declaration variants only. + +Per Modeling Practice 4 (Coproduct dissolution): SurfaceItem variants split into **DeclarationAllocating** (Fn/FnExternalBody/Data/TypeAtom/TypeRecord/etc.) + **NonDeclarationAllocating** (Let/Module/Import). The mapping is mechanical per variant kind, NOT uniform "1:1 to Declaration" across all SurfaceItem. -Per Modeling Practice 4 (Coproduct dissolution): each `SurfaceItem` variant has exactly one Declaration shape it maps to (per `lower.rs` top-comment + the 4-pattern dissolution receipt at the Declaration sum-variant level). The mapping is mechanical. +Earlier draft "every SurfaceItem variant maps 1:1 to a Declaration placeholder" was wrong — overstated. Per cursor INLINE BLOCKING PR #3077 line:130: Let/Module/Import are skipped in collect_symbols; they're parsed facts that flow forward to Pass 2 (for Let bodies) or are preserved-but-undeclared (Module/Import). ### §5.2 Pass 2 — Connective + behavior body lowering @@ -262,7 +301,7 @@ lower is structural-only; it produces target-agnostic `PreInferDag`. Shape A vs | Step | Deliverable | Owner | Substrate | |---|---|---|---| | **Step 1: Model review** | THIS DOC | Director (zesty-bear-812) | docs/design-lower-stage-l25-model.md (this doc) | -| **Step 2: Pipeline slot** | `fn lower(surface: SurfaceModule, elaboration: ElaborationSpec) -> PreInferDag` declared in compiler.dag with `ExternalRealization` body (Rust-backed placeholder pointing to current `lower.rs`). The signature uses `PreInferDag` directly (NOT a Result sum-variant); diagnostics couple INTO PreInferDag via **anchor-typed** diagnostic table per §4.3 (NOT port-only biconditional — anchors include PortAnchor / DeclarationAnchor / RecordFieldAnchor / SurfaceFormAnchor; biconditional applies only to PortAnchor). Modeling Practice 6 API-level enforcement at the PreInferDag carrier constructor. Step 2 worker brief authoring routes through Director after Decision 3.A operator ratification + PB-Substrate carrier extension. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | compiler.dag refinement | +| **Step 2: Pipeline slot** | `fn lower(surface: SurfaceModule, elaboration: ElaborationSpec) -> PreInferDag` declared in `src/v3/compiler/pipeline.dag` (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag) with `ExternalRealization` body (Rust-backed placeholder pointing to current `lower.rs`). The signature uses `PreInferDag` directly (NOT a Result sum-variant); diagnostics couple INTO PreInferDag via **anchor-typed** diagnostic table per §4.3 (NOT port-only biconditional — anchors include PortAnchor / DeclarationAnchor / RecordFieldAnchor / SurfaceFormAnchor; biconditional applies only to PortAnchor). Modeling Practice 6 API-level enforcement at the PreInferDag carrier constructor. Step 2 worker brief authoring routes through Director after Decision 3.A operator ratification + PB-Substrate carrier extension. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | pipeline.dag refinement | | **Step 3: Implementation** | `src/v3/std/lower.dag` (the .dag implementation of lower; fill the function body using ElaborationSpec rules) + `src/v3/std/elaboration_spec.dag` (NEW substrate carrier for the rules) | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 3 brief | src/v3/std/lower.dag + src/v3/std/elaboration_spec.dag (NEW substrate authorities) | | **Step 4: Parity test + simultaneous Rust deletion** | Parity verification authored as `.dag` TestClaim — generated test fixture set + `.dag` TestClaim asserting `lower_via_rust(surface) == lower_via_dag(surface, elaboration_spec)` structural-equality across canonical corpus (Dag comparison up to NodeId renaming). **P5 dissolution receipt**: this TestClaim is transient-by-construction; dissolves when lower.rs deletes in same PR. Any hand-Rust scaffolding for stage0 lower invocation routing bears P5 receipt `parity_lower_dag_vs_rust_scaffolding — transient; dissolves with lower.rs deletion in same PR per Step 4 atomic discipline`. `lower.rs` + (any per-pass sibling files like `lower_pass1.rs` if extant) DELETED in same PR. EXPECTED_HAND_AUTHORED_NON_TEST shrinks by N entries at PR-merge. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 4 brief | tests/parity_lower_dag_vs_rust (TestClaim shape, not hand-Rust .rs file) + `lower.rs` deletion | @@ -461,7 +500,7 @@ Subsequent L2.5 models (PB-5 infer / PB-3 parse / PB-2 tokenize) follow same Dir **Memory disciplines applied**: - `feedback_lenses_not_passes` (lower = structural elaboration, not decision engine) -- `feedback_fail_closed_discipline` C-8 (diagnostics coupled INTO PreInferDag via biconditional; partial-failure represented structurally, not via Result sum-variant) +- `feedback_fail_closed_discipline` C-8 (diagnostics coupled INTO PreInferDag via **anchor-typed** diagnostic table per §4.3; biconditional applies ONLY to PortAnchor-anchored diagnostics; other anchor kinds (DeclarationAnchor / RecordFieldAnchor / SurfaceFormAnchor) couple without port-state coupling; partial-failure represented structurally across anchor kinds, not via Result sum-variant) - `feedback_state_space_vs_behavioral_invariants` (typed-state PreInferDag at output) - `feedback_coproduct_dissolution` Practice 4 (sum-variant Dag per Decision 3.A) - `feedback_no_textual_enforcement_bridges` (no decision-logic in lower; ElaborationSpec facts) diff --git a/docs/design-parse-stage-l25-model.md b/docs/design-parse-stage-l25-model.md index 5721664f71a..b9cbbbbb969 100644 --- a/docs/design-parse-stage-l25-model.md +++ b/docs/design-parse-stage-l25-model.md @@ -30,7 +30,7 @@ Per the live `src/v3/compiler/src/parse_generated.rs:1-3` top-comment + `parse_t **Parse is a substrate-driven recursive-descent function: `List → SurfaceModule`, dispatching on token kind per compile-time-generated grammar tables. Surface carriers (`SurfaceModule`, `SurfaceItem`, `SurfaceExpr`, `SurfacePattern`, `SurfaceType`) are declared in `.dag` substrate; parser dispatch tables are declared in `.dag` substrate; parser BODY is the residual hand-Rust SG-2c surface still requiring substrate capability to retire.** -Per Decision 3.B operator-ratified (b) compile-time parser tables: GrammarSpec is `.dag` substrate that compiles to parser dispatch tables at build time. The parser engine reads compile-time-generated tables; no runtime grammar interpretation. This is more thesis-accurate than runtime data-driven (option a) — substrate authority all the way down. +Per Decision 3.B operator-ratified (b) compile-time parser tables: the table-families authored in `src/v3/compiler/parse_tables.dag` (conceptually called "GrammarSpec", though there is no `type GrammarSpec` carrier — see §3.2; the complete enumerated set lives in `parse_tables.dag` itself, not in this doc) are codegenned to parser dispatch tables at build time. The parser engine reads compile-time-generated tables; no runtime grammar interpretation, no runtime grammar input. This is more thesis-accurate than runtime data-driven (option a) — substrate authority all the way down. Per `feedback_lenses_not_passes`: parse dispatch decisions are substrate facts (token-to-operator mapping; bracket-role membership; type-rhs-boundary; primary-prefix dispatch). Anything parse "decides" is a grammar table that should be declared in `.dag`, not encoded in parser body logic. @@ -40,33 +40,38 @@ Per `feedback_lenses_not_passes`: parse dispatch decisions are substrate facts ( ## §3 Input types (declared in `.dag` substrate) -Two input types feed parse (`List` from tokenize + `GrammarSpec` per Decision 3.B compile-time tables): +**One** input type feeds parse at the API boundary: `List` from tokenize. Parser-dispatch tables (the conceptual "GrammarSpec" per Decision 3.B (b)) are NOT a runtime input — they are compile-time substrate at `parse_tables.dag`, codegenned into the parser body, and have no `type GrammarSpec` carrier. + +> **Cursor PR #3126 INLINE BLOCKING line:32 correction (2026-05-14T23:30:04Z, fixed in PR #3138)**: earlier draft framed `GrammarSpec` as a second declared `.dag` input type. Verified via `grep -rn "^type GrammarSpec\b" src/v3/ dsl/` — empty. No `type GrammarSpec` carrier exists in the repo. "GrammarSpec" in this doc is a concept-level grouping for the live table-families declared in `parse_tables.dag` (§3.2 enumerates them by `type` declaration with line anchors); it is not a substrate type and never appears in the Step-2 `fn parse` signature. ### §3.1 `List` (output of tokenize stage) The token stream produced by PB-2 tokenize. Each `Token` carries a `TokenKind` discriminant + lexeme + source-span. -**Substrate authority**: `Token` type at `src/v3/compiler/src/tokenize.rs` (current hand-Rust); will migrate to `src/v3/std/tokenize.dag` per PB-2 tokenize L2.5 (sibling doc in flight). The carrier shape is stable across the Rust/`.dag` boundary; PB-3 parse migration is independent of PB-2 tokenize migration status. +**Substrate authority** (corrected per cursor PR #3126 BLOCKING line:29 — earlier draft cited non-existent `tokenize.rs`): `Token` type lives in LIVE `.dag` substrate at `src/v3/std/tokenize.dag:65-67` (the shared taxonomy authority). Tokenizer implementation also lives in `.dag` at `src/v3/compiler/tokenize.dag` (154 lines). Codegen artifact at `src/v3/compiler/src/tokenize_generated.rs` (362 lines, auto-generated from tokenize.dag via `regen_tokenize`). The hand-Rust file `src/v3/compiler/src/tokenize.rs` does NOT exist — that was an earlier-draft phantom reference. Per `design-pure-bootstrap.md` §"PB-2 — tokenize retire" (line ~134), PB-2 has already substantially landed at substrate level (residual scaffold-retirement scope per PB-2 L2.5 PR #3127). The Token carrier shape is stable; PB-3 parse migration is independent of PB-2's residual-retirement timing per §7.1 axis split. **Lane dependency**: PB-2 tokenize (provides Token carrier + List output). -### §3.2 `GrammarSpec` (compile-time parser tables per Decision 3.B (b)) +### §3.2 Compile-time parser tables per Decision 3.B (b) — concept "GrammarSpec", NOT a carrier + +Per Decision 3.B operator-overrode my rec to (b) compile-time parser tables: "GrammarSpec" names the **conceptual** grouping of the live table-families authored in `parse_tables.dag` (enumerated below by `type` declaration line-anchor) and codegenned into the parser body. There is no `type GrammarSpec` declaration; the name is shorthand for those table-families taken collectively, not a substrate carrier or runtime input. + +**Live substrate authority**: `src/v3/compiler/parse_tables.dag` (517 lines at HEAD). The complete enumerated set of table-families lives there — this doc DOES NOT maintain a parallel count. Live `type` declarations at HEAD (verified `grep -nE '^type [A-Z]' src/v3/compiler/parse_tables.dag`): -Per Decision 3.B operator-overrode my rec to (b) compile-time parser tables: GrammarSpec is `.dag` substrate that the build system compiles to parser-dispatch tables. NOT a runtime-interpreted spec. +- `BinaryOpRow` at `parse_tables.dag:167` (SG-2c-1) +- `TopLevelItemKwRow` at `parse_tables.dag:289` (SG-2c-2 + projection SG-2c-3 `is_type_rhs_boundary_keyword`) +- `SoftKeywordIdentRow` at `parse_tables.dag:334` (soft-keyword identifier dispatch; no SG-2c-N number) +- `BracketRow` at `parse_tables.dag:385` (SG-2c-4) +- `PrimaryPrefixRow` at `parse_tables.dag:449` (SG-2c-6) +- `PrimaryAtomRow` at `parse_tables.dag:486` (SG-2c-7) -**Live substrate precedent** at `src/v3/compiler/parse_tables.dag:1-30` (verified): -- Binary operator token-to-semantics mapping (SG-2c-1) -- Top-level item keyword → parse_item dispatch class (SG-2c-2) -- Type-RHS boundary keyword membership (SG-2c-3) -- Bracket opener/closer role membership (SG-2c-4) -- Primary-expression prefix openers (SG-2c-6) -- Primary-expression atomic tail tokens (SG-2c-7) +Plus the supporting enum `BinaryOpLevel` at `parse_tables.dag:133`. The build system reads `parse_tables.dag` + emits `parse_tables_generated.rs` (compile-time codegen); the parser body consumes the generated tables via `binary_op_at_level`, `top_level_item_dispatch`, `is_type_rhs_boundary_keyword`, `bracket_role`, `primary_prefix_dispatch`, `primary_atom_class`, and the soft-keyword-ident lookup. None of these is reached through a `GrammarSpec` carrier; they are direct table lookups inside the parser body. -These 6 table-families ARE the GrammarSpec content. The build system reads `parse_tables.dag` + emits `parse_tables_generated.rs` (compile-time codegen); the parser body consumes the generated tables via `binary_op_at_level`, `top_level_item_dispatch`, `is_type_rhs_boundary_keyword`, `bracket_role`, `primary_prefix_dispatch`, `primary_atom_class`. +> **Codex PR #3126 BLOCKING (sha 5619afac) revision (fixed at HEAD of PR #3138)**: an earlier draft of this enumeration listed only six SG-2c-numbered table-families, copying the prose summary at `parse_tables.dag:23-29` rather than the live `type` declarations — `SoftKeywordIdentRow` (line 334) was missed because it lacks an SG-2c-N number in the prose summary. Per `feedback_parallel_representation_debt`: the doc now treats `parse_tables.dag` as the single enumerated authority. Hand-counts ("6 table-families") are dropped wherever they appeared; readers should grep the file for `^type [A-Z]` to see the live set. -**Per Decision 3.B (b) "harder/more correct"**: compile-time table generation IS substrate authority all-the-way-down. Runtime grammar interpretation (option a my-original-rec) would require a generic parser engine reading GrammarSpec at runtime — more flexible but introduces runtime authority that compile-time tables don't have. (b) preserves the property that GrammarSpec is data-known-at-compile-time, not data-fetched-at-runtime. +**Per Decision 3.B (b) "harder/more correct"**: compile-time table generation IS substrate authority all-the-way-down. Runtime grammar interpretation (option a my-original-rec) would require a generic parser engine reading a runtime spec — more flexible but introduces runtime authority that compile-time tables don't have. (b) preserves the property that the grammar is data-known-at-compile-time, not data-fetched-at-runtime, and therefore needs no runtime carrier. -**Substrate authority**: `src/v3/compiler/parse_tables.dag` is LIVE at HEAD (517 lines; 6 table-families). PB-3 migration extends this with the full SG-2c parser body authority via the substrate capability dependency in §6 (recursive list-body emission per SELF_HOSTING.md §6 Phase 4a). +**Substrate authority**: `src/v3/compiler/parse_tables.dag` is LIVE at HEAD (517 lines). The current table-family set is enumerated in §3.2 above by `type`-declaration line-anchor; this doc does NOT maintain a parallel count. PB-3 migration extends this with the full SG-2c parser body authority via the substrate capability dependency in §6 (recursive list-body emission per SELF_HOSTING.md §6 Phase 4a). **Lane dependency**: PB-Substrate (substrate-capability for recursive list-body emission — REQUIRED for SG-2c full parser body migration per `parse_tables.dag:18-22` STOP-AND-ESCALATE bullet). @@ -76,9 +81,11 @@ These 6 table-families ARE the GrammarSpec content. The build system reads `pars ### §4.1 `SurfaceModule` (typed-state output) -Per `src/v3/std/parse_surface.dag:29` (verified live): `SurfaceModule` is the top-level parse output carrying `List` + module-level metadata. +Per `src/v3/std/parse_surface.dag:29` (verified live): `SurfaceModule { items: List }` — single field only. Per cursor PR #3126 APPROVE_WITH_COMMENTS line:79 + INVARIANTS P1 live-state honesty: earlier draft "+ module-level metadata" was a phantom addition; live carrier has ONLY the `items` field. -**Construction-time invariant**: every Token consumed produces either a SurfaceItem/SurfaceExpr/SurfacePattern/SurfaceType/SurfaceLiteral variant in the output tree (parser advanced) OR a ParseDiagnostic in the diagnostic stream (parser failed-closed at that position). No tokens silently dropped; no surface forms fabricated without source-span provenance. +**Live failure boundary** (per `parse_generated.rs:138`): parse returns `Result` (fail-closed; aborts on first parse error). NO diagnostic-stream coupling in live state. + +**Construction-time invariant** (live + ratified shape per §4.3 Result-sum framing): every Token consumed produces either a SurfaceItem/SurfaceExpr/SurfacePattern/SurfaceType/SurfaceLiteral variant in the Ok-arm output tree (parser advanced) OR triggers Err-arm with a typed `ParseDiagnostic` (parser failed-closed at first error position). No tokens silently dropped; no surface forms fabricated without source-span provenance; no partial-parse SurfaceModule with embedded diagnostics (per codex BLOCKING #3126 fail-closed correction). **Substrate authority**: `src/v3/std/parse_surface.dag:29` (live; closed-axis SurfaceModule + SurfaceItem at :257 + SurfaceExpr at :149 + SurfacePattern at :123 + SurfaceType at :67 + SurfaceLiteral at :143). PB-3 parse output type is stable across the migration. @@ -102,42 +109,56 @@ type SyntaxFormRef | TypeForm // expected in type position | LiteralForm // expected at literal position -type ParseDiagnostic +// Diagnostic = record-wraps-kind pattern per codex BLOCKING PR #3126 finding 1: +// matches live Diagnostic shape at diagnostics.dag:150 (record { kind, span, ... } +// where kind is the closed-axis sum, span on the record). Earlier draft mixed +// kind + span into variant fields directly; corrected to: + +type ParseDiagnostic { + kind: ParseDiagnosticKind + span: SourceSpan + // Optional: additional context fields per Step 2 brief enumeration +} + +type ParseDiagnosticKind = UnexpectedToken { found: TokenKindRef, expected: List, context: SyntaxFormRef } - | UnterminatedConstruct { construct: SyntaxFormRef, opener_span: SourceSpan } + | UnterminatedConstruct { construct: SyntaxFormRef, opener_span: SourceSpan } // additional opener_span here; ParseDiagnostic.span is the unterminated-end position | InvalidLiteral { kind: TokenKindRef, reason: String } // reason is human display - | DuplicateRecordFieldLabel { label: NonEmptyStr, prior_span: SourceSpan } // per PR #3075 ratchet - | (additional variants per Step 2 worker brief authoring against parse_generated.rs) + | DuplicateRecordFieldLabel { label: NonEmptyStr, prior_span: SourceSpan } // per PR #3075 ratchet; ParseDiagnostic.span is current-site + // (additional variants per Step 2 worker brief authoring against parse_generated.rs) + +// Per cursor PR #3126 BLOCKING line:65: span lives on ParseDiagnostic record, +// not on each kind variant. Per codex finding 1: matches live Diagnostic +// shape (kind + span on record carrier). Single source of truth for span; +// variant-specific spans (opener_span / prior_span) live on kind variants +// where they're meaningful. ``` **Lane dependency**: PR #3077 §12 Q7 ratification (cross-stage); Director-tier per-stage variant authoring. -### §4.3 No separate `ParseResult` sum-variant — diagnostics couple via SurfaceModule extension (PROPOSED) +### §4.3 Parse output is `Result` — NO SurfaceModule diagnostics-extension -**Live substrate state at HEAD** (verified via `grep -n "^type SurfaceModule" src/v3/std/parse_surface.dag`): `type SurfaceModule { items: List }` — single field, NO existing diagnostic-table or diagnostic-stream field. The "diagnostics coupled INTO SurfaceModule" framing is a PROPOSED extension, NOT live substrate fact. Per codex INLINE BLOCKING #3126: this needs explicit "proposed-not-live" marking + parallel-substrate-extension authoring scope. +> **Codex PR #3138 BLOCKING (sha cd6e8d15) revision (fixed at HEAD of PR #3138)**: this section previously proposed an `items + diagnostics` extension to `SurfaceModule` modeled on PB-4 lower's PreInferDag / PB-5 infer's InferredDag. That model is REJECTED for parse. The corrected Decision (per codex BLOCKING PR #3126 ratified into §6 Step 2 + §4.2) is that parse uses **Result-sum** (fail-fast first-error abort), like PB-6 emit and unlike PB-4 lower / PB-5 infer. The earlier-draft SurfaceModule extension is superseded; no `diagnostics` field is being added to `parse_surface.dag`. -**Proposed extension** (NEW substrate authoring, parallel to PR #3077 §12 Q7 Decision 2.B extension path): +**Live substrate state at HEAD** (verified via `grep -n "^type SurfaceModule" src/v3/std/parse_surface.dag`): `type SurfaceModule { items: List }` — single field, NO diagnostic field, and no extension is now proposed. Parse-stage diagnostics live OUTSIDE SurfaceModule, on the Err branch of `Result`. -Same pattern as PB-4 lower's PreInferDag + PB-5 infer's InferredDag (per PR #3077 §4.3 + PR #3085 §4.3) — output IS the typed-state carrier with diagnostics coupled structurally. For SurfaceModule, the extension is: +**Cross-stage discriminator** (the rule that places parse with emit, not with lower/infer): +- **Result-sum** (PB-3 parse + PB-6 emit): fail-fast output domain — a partial parse tree or a partial target-byte buffer is not a valid intermediate; the stage either produces the complete artifact or fails with the first diagnostic. +- **Typed-state-with-coupled-diagnostics** (PB-4 lower + PB-5 infer): structural output domain — Unresolved ports / pre-inferred Dag are valid intermediates consumed downstream, so the diagnostics couple structurally into the carrier. -``` -// Proposed extension to src/v3/std/parse_surface.dag SurfaceModule: -type SurfaceModule { - items: List // existing - diagnostics: List // PROPOSED extension per this L2.5 -} -``` +Live parser at `src/v3/compiler/src/parse_generated.rs:138` matches this: `pub fn parse(tokens: &[Token], file: &str) -> Result`. The Step 2 contract refines that to `fn parse(tokens: List) -> Result` per §6. -Where the diagnostic-list is indexed by source-span (the natural key for parse-stage failures since pre-substrate positions don't yet have ports). +Signature: `fn parse(tokens: List) -> Result` per live `parse_generated.rs:138` shape (Result; ParseDiagnostic is per-stage refinement per §4.2). -**Step 2 worker brief must include** the parse_surface.dag SurfaceModule extension as part of the pipeline-slot PR scope — not separately deferrable. Without this, Step 2 has the carrier-mismatch problem codex flagged. +**Per codex BLOCKING PR #3126**: earlier draft of this section proposed `fn parse(tokens, grammar: GrammarSpec) -> SurfaceModule` with embedded diagnostics field. Both load-bearing problems: -**Cross-stage consistency**: -- PB-3 parse + PB-4 lower + PB-5 infer: output IS typed-state carrier with diagnostics coupled structurally (each requires its own carrier extension if not already live; SurfaceModule extension is PB-3-specific) -- PB-6 emit: uses EmissionResult sum because emit produces target-language bytes (different output domain) -- Discriminator: when output is a STRUCTURAL value (Dag-shape or Surface-tree), partial-failure couples structurally; when output is FINAL ARTIFACT (target source bytes), partial-failure couples via Result sum. +1. **GrammarSpec parallel-authority** (P2 violation): §3.2 says the parser tables are compile-time-only, NOT runtime-interpreted, and have no `type GrammarSpec` carrier in the repo. The earlier-draft signature accepted a `GrammarSpec` value as runtime input — that would have invented a parallel authority alongside the compile-time tables AND a substrate type that doesn't exist. The corrected signature consumes the compiled tables via internal parser-table-driven dispatch — no runtime grammar input, no `GrammarSpec` parameter, no carrier needed. -Signature: `fn parse(tokens: List, grammar: GrammarSpec) -> SurfaceModule` (NOT ParseResult). Diagnostics coupled via the PROPOSED `diagnostics: List` field added to SurfaceModule per Step 2 PR scope. +2. **Fail-closed weakening** (P3 + Practice 1+2 violation): live parser at `parse_generated.rs:138` returns `Result` (fail-closed, aborts on first error). Earlier draft proposed `SurfaceModule` with embedded diagnostics — would let partial-parse states be constructible + let downstream observe "success" output after parse failure. The corrected signature preserves fail-closed Result: parse either succeeds with complete SurfaceModule OR fails with first-encountered ParseDiagnostic (no partial states). + +Cross-stage consistency NOTE: parse uses Result-sum (like emit) NOT typed-state-with-coupled-diagnostics (like lower/infer) because parse is fail-fast (single first-error abort). Different discriminator from `feedback_fail_closed_discipline`-applied-to-Dag-output stages: +- **Result-sum** (parse + emit): fail-fast output domain; partial output isn't a valid intermediate +- **Typed-state-with-coupled-diagnostics** (lower + infer): structural output domain where partial-failure (Unresolved ports) IS valid intermediate state consumed by downstream --- @@ -147,16 +168,7 @@ parse's structure is **recursive-descent dispatching on compile-time-generated g ### §5.1 Compile-time table generation -Per Decision 3.B (b) operator-ratified: GrammarSpec is `.dag` substrate compiled to parser dispatch tables at build time. Live precedent at `parse_tables.dag` + `parse_tables_generated.rs`. The 6 table-families: - -1. Binary-operator precedence table (SG-2c-1) -2. Top-level item keyword dispatch (SG-2c-2) -3. Type-RHS boundary keyword membership (SG-2c-3) -4. Bracket opener/closer role (SG-2c-4) -5. Primary-prefix dispatch (SG-2c-6) -6. Primary-atom class (SG-2c-7) - -Step 3 of PB-3 migration extends this with the full parser-body authority once the substrate capability dependency lands (§6). +Per Decision 3.B (b) operator-ratified: the parser-dispatch tables (concept "GrammarSpec"; no `type GrammarSpec` carrier — see §3.2) are `.dag` substrate compiled at build time. Live precedent at `parse_tables.dag` + `parse_tables_generated.rs`. The complete table-family set is enumerated by `type` declaration in §3.2 above (with line anchors); this section does NOT maintain a parallel enumeration. Step 3a of PB-3 migration extends `parse_tables.dag` with additional table-families as new grammar productions get table-driven; Step 3b/4 extends the full parser-body authority once the substrate capability dependency lands (§6). ### §5.2 Recursive-descent parser body @@ -170,7 +182,7 @@ Per `parse_tables.dag:13-22` STOP-AND-ESCALATE bullet: > "SG-2c proper (parser authority proper — retiring `parse_parser_body.txt` as parse logic) is blocked on a named substrate capability: recursive list-body emission over `List` with cursor threading. See `src/v3/std/list.dag:13-15` (...) and SELF_HOSTING.md §6 Phase 4a. Until that lands, any full `.dag` parser port routes through a hidden Rust host layer, which SG-2c's STOP-AND-ESCALATE bullet forbids." -**Implication for PB-3 migration**: Step 3 (`.dag` implementation of parser body) BLOCKED on substrate-capability landing. Step 2 (pipeline-slot declaration) + Step 3a (extend grammar tables in `parse_tables.dag`) are unblocked. +**Implication for PB-3 migration**: Step 3 (`.dag` implementation of parser body) BLOCKED on substrate-capability landing. Step 2 (pipeline-slot declaration) + Step 3a (extend grammar tables in `parse_tables.dag`) are unblocked at HEAD — **gate**: PR #3077 §12 Q7 ratification was a hard precondition on Step 2 brief authoring (§7.2) and merged 2026-05-15T00:21:19Z, so the gate is satisfied. Step 2 must NOT be brief-authored before that merge timestamp; at HEAD it has been. --- @@ -180,13 +192,13 @@ Per `feedback_anchor_mgr_lane_synthesis_on_gap_tier_not_session_id`. | Prereq | Substrate authority | Gap-tier lane | Status at HEAD (as of 2026-05-14) | |---|---|---|---| -| PB-2 Tokenize | `src/v3/std/tokenize.dag` (NEW per PB-2 L2.5) + Token carrier | PB-2 lane (R3 Substrate Mgr post-PB-3) | NOT-STARTED; Director PB-2 L2.5 in flight (sibling doc) | +| PB-2 Tokenize | `src/v3/std/tokenize.dag` (LIVE — 143 lines; Token + TokenKind taxonomy already substantially landed per `design-pure-bootstrap.md` §"PB-2 — tokenize retire") | PB-2 lane (R3 Substrate Mgr) | LIVE substrate at HEAD; PB-2's residual scope is scaffold-retirement (SG-1a + character-level + codegen-driver per PB-2 L2.5 §1), NOT carrier authoring. PB-3 consumes the live Token carrier; carrier shape stable across PB-2's residual-retirement timing. | | Live SurfaceModule + Surface* carriers | `src/v3/std/parse_surface.dag` (live; closed-axis sums) | PB-Substrate | LIVE at HEAD per PR #3077 §3.1 audit | | Live grammar tables substrate | `src/v3/compiler/parse_tables.dag` (517 lines; SG-2c-1/2/3/4/6/7 tables) | PB-Substrate | LIVE at HEAD; Step 3 extends | | Substrate-capability: recursive list-body emission | `src/v3/std/list.dag` capability + SELF_HOSTING.md §6 Phase 4a | PB-Substrate + R3 Grounding Mgr lane | BLOCKER for full Step 3; per parse_tables.dag:13-22 STOP-AND-ESCALATE | | ParseDiagnostic substrate extension | extension of `src/v3/std/diagnostics.dag:150` per Decision 2.B per PR #3077 §12 Q7 | PB-Substrate + Director-tier per-stage authoring | Carrier LIVE; per-stage variant authoring NEW per Q7 ratification | -**Critical observation**: PB-3 parse has a HARD DEPENDENCY on substrate-capability landing (recursive list-body emission) for Step 3 full implementation. Step 2 + grammar-table extensions are unblocked. This makes PB-3 migration HARDER than PB-4/PB-5 — the substrate-capability gap is real, not just an L2.5 ratification question. +**Critical observation**: PB-3 parse has a HARD DEPENDENCY on substrate-capability landing (recursive list-body emission) for Step 3 full implementation. Step 2 + grammar-table extensions are unblocked at HEAD — **the additional Step 2 gate is PR #3077 §12 Q7 ratification** (§7.2 / §15 step 4), which merged 2026-05-15T00:21:19Z; without that merge timestamp in repo history Step 2 brief authoring is structurally blocked (P3 failure shape would land unfixed). This makes PB-3 migration HARDER than PB-4/PB-5 — the substrate-capability gap is real, not just an L2.5 ratification question. --- @@ -194,13 +206,19 @@ Per `feedback_anchor_mgr_lane_synthesis_on_gap_tier_not_session_id`. ### §7.1 Upstream dependencies -parse depends on `List` from tokenize (PB-2). PB-2 tokenize migration is downstream in the bottom-up order. Per `src/v3/SELF_HOSTING.md` §2 migration order, parse migrates AFTER tokenize substrate-side stable, but Token carrier shape is stable regardless of whether tokenize-emitter is hand-Rust or `.dag` — PB-3 parse migration is independent of PB-2 tokenize migration status (same independence pattern as PB-4 lower vs PB-3 parse per PR #3077 §12 Q6). +parse depends on `List` from tokenize (PB-2). **Two distinct axes per cursor PR #3126 APPROVE_WITH_COMMENTS clarification**: + +1. **Substrate-stability ordering** (SELF_HOSTING.md §2 migration order; bottom-up): tokenize's substrate (Token carrier shape) must be stable BEFORE parse migration proceeds. This is already true at HEAD — `src/v3/std/tokenize.dag:65-67` declares the live Token carrier shape; tokenize.dag is the substrate authority. Substrate-side stability ✓. + +2. **Migration-timing independence** (parallel-dispatch axis): PB-3 parse migration can ship in parallel with PB-2 tokenize MIGRATION — i.e., parse migration doesn't WAIT for PB-2's residual hand-Rust retirement (per PB-2 L2.5 §1: SG-1a + character-level scaffold + codegen-driver retirement). What parse needs is the stable Token CARRIER, which already exists; PB-2's migration is about retiring the residual emitter-side hand-Rust, not about changing the carrier. + +So: ordering (substrate-stability) IS satisfied (live); independence (migration-timing) means parse migration is parallel-dispatchable with respect to PB-2's residual-retirement work. Both claims coherent on the same axis split; not contradictory. ### §7.2 Downstream consumers PB-4 lower consumes `SurfaceModule` from parse (per PR #3077 §3.1). The carrier is LIVE at parse_surface.dag; downstream stages don't depend on parse's migration timing. -Per Decision 2.B discriminated-union diagnostics: parse's diagnostics are discriminable by source via whichever substrate-extension path PR #3077 §12 Q7 ratifies. **PR #3077 §12 Q7 must ratify before any Step 2 worker brief authoring** (cross-stage authority for Decision 2.B extension path). +Per Decision 2.B discriminated-union diagnostics: parse's diagnostics are discriminable by source via the substrate-extension path PR #3077 §12 Q7 ratified. **Gate satisfied 2026-05-15T00:21:19Z** when PR #3077 merged; Step 2 worker brief authoring is unblocked at HEAD per §15 step 4. The contradiction cursor PR #3126 BLOCKING line:119 flagged (§6 unblocking claims vs §7.2 gate claim) is now resolved by the Q7 ratification merging rather than by retracting either statement. ### §7.3 Sibling-stage coordination @@ -221,10 +239,9 @@ parse is target-agnostic structural parsing; Shape A/B disambiguation lives at e | Step | Deliverable | Owner | Substrate | |---|---|---|---| | **Step 1: Model review** | THIS DOC | Director (zesty-bear-812) | docs/design-parse-stage-l25-model.md (this doc) | -| **Step 2: Pipeline slot** | `fn parse(tokens: List, grammar: GrammarSpec) -> SurfaceModule` declared in compiler.dag with `ExternalRealization` body (Rust-backed placeholder pointing to current `parse_generated.rs:138`). Signature consistent with PB-4/PB-5 pattern (output IS typed-state carrier; diagnostics coupled INTO SurfaceModule). | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | compiler.dag refinement | +| **Step 2: Pipeline slot** | `fn parse(tokens: List) -> Result` declared in `src/v3/compiler/pipeline.dag` (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag) with `ExternalRealization` body (Rust-backed placeholder pointing to current `parse_generated.rs:138`). Signature matches live parser shape (fail-closed Result; aborts on first parse error). NO runtime GrammarSpec input — compile-time generated parser tables consumed via internal dispatch per Decision 3.B (b). Per codex BLOCKING #3126: distinct from lower/infer typed-state-with-coupled-diagnostics pattern; parse uses Result-sum like emit (fail-fast output domain). | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | pipeline.dag refinement | | **Step 3a: Grammar table extension** | Extend `src/v3/compiler/parse_tables.dag` with remaining grammar productions (unary operators / let-bindings / fn-signatures / pattern-syntax / module-imports / etc.); regenerate `parse_tables_generated.rs` | R3 Substrate Mgr — worker dispatched against Director-authored Step 3a brief; **unblocked NOW** (doesn't depend on substrate capability) | parse_tables.dag extension | -| **Step 3b: Full parser-body `.dag` migration** | `src/v3/std/parse.dag` — full recursive-descent parser body in `.dag`; consumes parse_tables.dag-generated dispatch tables | R3 Substrate Mgr — **BLOCKED on substrate-capability landing** (recursive list-body emission per `list.dag:13-15` + SELF_HOSTING.md §6 Phase 4a); cannot proceed without it | src/v3/std/parse.dag (NEW substrate authority) — DEPENDS on substrate-capability | -| **Step 4: Parity test + simultaneous Rust deletion** | Parity verification authored as `.dag` TestClaim — assert `parse_via_rust(tokens) == parse_via_dag(tokens, grammar_spec)` structural-equality across canonical corpus. **P5 dissolution receipt**: TestClaim is transient-by-construction; dissolves when `parse_generated.rs:138` parse() body deletes in same PR. Any hand-Rust scaffolding bears P5 receipt: `parity_parse_dag_vs_rust_scaffolding — transient; dissolves with parse_generated.rs parse() body deletion in same PR per Step 4 atomic discipline`. `parse_parser_body.txt` + `parse_generated.rs` parse() body DELETED in same PR. EXPECTED_HAND_AUTHORED_NON_TEST shrinks by N entries at PR-merge. | R3 Substrate Mgr — worker dispatched against Director-authored Step 4 brief | tests/parity_parse_dag_vs_rust (TestClaim shape) + parse body deletion | +| **Step 3b/4 COMBINED: Full parser body + parity + simultaneous deletion** | ATOMIC single-PR deliverable bundling what the SELF_HOSTING.md §2.2 4-step template lists separately as Step 3b and Step 4. Contents: (a) `src/v3/std/parse.dag` (full recursive-descent parser body in `.dag` consuming parse_tables.dag-generated dispatch tables); (b) parity verification as a `.dag` `TestClaim` asserting `parse_via_rust(tokens) == parse_via_dag(tokens)` Result-equality across canonical corpus (both return `Result`; equality on the Ok arm is SurfaceModule structural-equality, on the Err arm ParseDiagnostic content-equality); (c) `parse_parser_body.txt` + `parse_generated.rs:138` parse() body DELETION in the same PR. **P5 dissolution receipt**: the parity `TestClaim` is transient-by-construction — it dissolves the moment the Rust parse() body it compares against deletes in this same PR. Any hand-Rust scaffolding bears the same receipt: `parity_parse_dag_vs_rust_scaffolding — transient; dissolves with parse_generated.rs parse() body deletion in same PR per Step 3b/4 atomic discipline`. Per cursor PR #3126 BLOCKING line:181 + `feedback_paper_shrink_variants`: cannot land `.dag` parser body BEFORE Rust deletion (temporary Rust+`.dag` coexistence is the paper-shrink-relocation failure mode). Per openai-pro PR #3126 BLOCKING (sha 16d21f4a) Finding 6: this row REPLACES the previously-leftover "Step 4: Parity test" row that contradicted the COMBINED framing; §15 sequencing also collapsed (see §15 step 12). EXPECTED_HAND_AUTHORED_NON_TEST shrinks by N entries at PR-merge. | R3 Substrate Mgr — **BLOCKED on substrate-capability landing** (recursive list-body emission per `list.dag:13-15` + SELF_HOSTING.md §6 Phase 4a); cannot proceed without it (subject to §12 Q1 ratification — see "Substrate-capability bundling" below) | src/v3/std/parse.dag (NEW substrate authority) + parse_generated.rs deletion + tests/parity_parse_dag_vs_rust TestClaim — DEPENDS on substrate-capability | **Critical: parity test is against PARSE.RS OUTPUT, not against parse.dag-template-of-parse.rs** per `feedback_paper_shrink_variants`. Same discipline as PB-6/PB-5/PB-4 Step 4. @@ -243,7 +260,7 @@ Per Step 3b brief authoring (post-substrate-capability landing): use structural Per `feedback_fail_closed_discipline` C-8 + parser hygiene: - **No silent token-skipping**: every consumed token contributes to a Surface variant OR triggers a ParseDiagnostic. -- **Source-span provenance**: every Surface variant + every ParseDiagnostic carries `SourceSpan`. No fabricated spans. +- **Source-span provenance** (corrected per cursor PR #3126 BLOCKING line:151): **most** Surface variants carry `SourceSpan` directly, but per live `src/v3/std/parse_surface.dag`: SurfaceItem::Let (line ~33) has fields `{ name, type_ann, expr }` with NO direct span; SurfaceLiteral variants `Int(String) | Bool(Bool) | String(String)` are plain-tuple variants with NO direct span. These cases acquire span via their enclosing carrier (Let-item inherits container span; SurfaceLiteral always wraps within `Literal { value, span }` per parse_surface.dag:150). Step 2 PR scope SHOULD audit whether the Let + SurfaceLiteral exceptions are structural-honest (acceptable per enclosing-carrier-provides-span) OR substrate-extension-required (add span directly to those variants). Earlier "every Surface variant carries SourceSpan" claim was incorrect; corrected to "every parse output structurally has source-span provenance via direct field OR enclosing carrier". Every ParseDiagnostic carries `SourceSpan` directly per §4.2. - **Recursive depth bounded**: parser body uses bounded recursion per grammar rules; no fixed-point iteration. - **Anti-bridge invariant** (per `feedback_no_textual_enforcement_bridges`): grammar tables are the SINGLE authority on token-to-syntactic-form mapping; no fallback hand-Rust heuristics. @@ -262,13 +279,17 @@ Per `parse_tables.dag:13-22` + SELF_HOSTING.md §6 Phase 4a: SG-2c full parser-b - **(b) Operator-ratified scope extension**: substrate-capability is part of PB-3 Step 3b scope; warm-wolf-698 authors both in same PR - **(c) Defer PB-3 Step 3b to post-R3**: only PB-3 Step 2 + Step 3a (grammar table extension) land in R3; full parser body migration is R4 scope -**Director-recommend: (b) bundled** — PB-3 Step 3b explicitly depends on substrate-capability; bundling keeps the authority chain clean (one worker, one PR, one P5 receipt). Operator/PM ratification. +**Director-recommend: (b) bundled** — PB-3 Step 3b explicitly depends on substrate-capability; bundling keeps the authority chain clean (one worker, one PR, one P5 receipt). + +**Status: PENDING operator/PM ratification.** Until ratified, the doc treats the substrate-capability dependency as path (a) — a separate PB-Substrate / R3 Grounding Mgr lane, listed as a §13 non-goal of PB-3 and as the §15 step-11 wait-gate. The §9 Step 3b/4 COMBINED row's "BLOCKED on substrate-capability landing" framing matches that default. If/when operator ratifies (b), §13 drops the substrate-capability non-goal and §15 step-11 collapses into the COMBINED Step 3b/4 brief. + +> **Per openai-pro PR #3126 BLOCKING (sha 16d21f4a) Finding 2.5**: the previous draft asserted "Director-recommend: (b) bundled" as if it were the ratified plan while §13 + §15 still framed substrate-capability as a separate lane / wait-step. That made it possible to read the doc and execute either bundled or waited execution. The PENDING-ratification annotation above resolves the ambiguity: default execution is path (a), and bundling only activates on explicit operator ratification. ### Q2: Compile-time table generation discipline (per Decision 3.B (b)) Decision 3.B operator-overrode my rec to (b) compile-time parser tables. Existing precedent at `parse_tables.dag` + `parse_tables_generated.rs`. Step 3a extends this. -**Open question**: does Step 3a extension introduce any axis NOT currently in the 6 table-families? (e.g., custom-operator definitions, module-imports syntax, where-clause refinement syntax). Director-recommend: Step 3a worker brief enumerates the FULL grammar table set by grepping `parse_generated.rs` for remaining open-coded TokenKind dispatch sites; each is a candidate new table-family. +**Open question**: does Step 3a extension introduce any axis NOT currently in the `parse_tables.dag` table-family set (enumerated in §3.2 above by `type` declaration with line anchors)? Examples of likely-new axes: custom-operator definitions, module-imports syntax, where-clause refinement syntax. Director-recommend: Step 3a worker brief enumerates the FULL grammar table set by grepping `parse_generated.rs` for remaining open-coded TokenKind dispatch sites; each is a candidate new table-family. ### Q3: ParseDiagnostic variant exhaustiveness @@ -281,7 +302,7 @@ Decision 3.B operator-overrode my rec to (b) compile-time parser tables. Existin If Director-recommend (b) in Q1 holds: PB-3 Step 3b waits for substrate-capability. What CAN PB-3 do BEFORE substrate-capability lands? **Director-recommend**: -- Step 2 (pipeline-slot in compiler.dag) — unblocked +- Step 2 (pipeline-slot in `src/v3/compiler/pipeline.dag`) — unblocked at HEAD; gate was PR #3077 §12 Q7 (per §7.2 + §15 step 4), satisfied 2026-05-15T00:21:19Z - Step 3a (grammar table extension in parse_tables.dag) — unblocked - Step 3b (full parser body `.dag` migration) — blocked on substrate-capability @@ -291,17 +312,20 @@ This phases naturally with the substrate-capability landing as the trigger for S `parse_generated.rs` is 2018 lines. Per `feedback_paper_shrink_variants` discipline, phased migration with per-phase P5 receipts is acceptable. -**Director-recommend: Step 3a + Step 3b PHASING**: -- 3a: grammar table extension (parse_tables.dag growth + parse_tables_generated.rs regen) — landed as own PR -- 3b: full parser body `.dag` migration + parser body deletion in same PR +**Director-recommend: Step 3a + COMBINED-Step-3b/4 PHASING** (per cursor PR #3126 BLOCKING line:181 — earlier framing conflated Step 3b "migration" vs Step 4 "parity-and-delete" creating P5 receipt ambiguity): + +- **Phase 3a** (separate PR): grammar table extension (parse_tables.dag growth + parse_tables_generated.rs regen). P5 receipt: ROADMAP.md deferral row naming Step 3b/4 as future-receipt scope (no parser-body deletion in 3a; refactor-only phase per `feedback_paper_shrink_variants` deletion-or-deferral discipline). +- **Phase 3b/4 COMBINED** (single PR): full parser body `.dag` migration + parity test (`.dag` TestClaim) + `parse_generated.rs:138` parse() body deletion + EXPECTED_HAND_AUTHORED_NON_TEST shrink — ALL atomic in one PR. P5 receipt: deletion + census shrink. + +**Why combined Step 3b/4** (not sequential phases): if Step 3b lands `.dag` parser body BEFORE Step 4 deletes Rust parse() body, the Rust + `.dag` parser bodies coexist temporarily — paper-shrink-relocation risk per `feedback_paper_shrink_variants`. Combining ensures atomic substrate substitution: Rust parse() body never coexists with `.dag` parser body on main. -Each phase = own PR + own parity test + own P5 receipt. Step 4 (parity-and-delete) is ONLY the 3b phase since 3a doesn't yet delete the parser body. +**Update to §9 + §15**: Step 3b row reframed as "Step 3b/4 COMBINED"; §15 sequence collapses steps 12-17 into single dispatch+merge for combined phase. ### Q6: PB-2 tokenize landing dependency PB-3 parse's input is `List` from tokenize (PB-2). **Does PB-3 parse migration block on PB-2 tokenize migration completing?** -**Director-recommend: NO — PB-3 parse migrates independently of PB-2 tokenize status**, same shape as PB-4 lower per PR #3077 §12 Q6 + PB-5 infer per PR #3085 §12 Q6. Token carrier shape is stable; PB-3 parse migrates when its substrate (GrammarSpec + parse_tables.dag + substrate-capability for Step 3b) is at HEAD. +**Director-recommend: NO — PB-3 parse migrates independently of PB-2 tokenize status**, same shape as PB-4 lower per PR #3077 §12 Q6 + PB-5 infer per PR #3085 §12 Q6. Token carrier shape is stable; PB-3 parse migrates when its substrate (`parse_tables.dag` table-families per §3.2 enumeration + substrate-capability for Step 3b recursive list-body emission) is at HEAD. --- @@ -345,20 +369,17 @@ Post-ratification: this doc becomes substrate authority for Step 2 + Step 3a wor 1. **Operator / PM-delegate ratifies §12 Q1–Q6** (per 2026-05-14 directive) 2. **PM amends close plan** to route through PB-X lanes + cite this doc as PB-3 L2.5 substrate 3. **PM amends §1.8** with PB-3 gate row citing this doc -4. **PR #3077 §12 Q7 ratifies** (cross-stage Decision 2.B extension path) — affects PB-3 ParseDiagnostic shape +4. **PR #3077 §12 Q7 ratifies** (cross-stage Decision 2.B extension path) — affects PB-3 ParseDiagnostic shape. **DONE 2026-05-15T00:21:19Z** when PR #3077 (PB-4 lower L2.5) merged; Q7 ratification carried in that merge. The §6 / §7.2 "Step 2 must wait on Q7" gate is therefore now satisfied — Step 2 brief authoring is genuinely unblocked, not just procedurally listed as the next step. 5. **Director authors PB-3 Step 2 worker brief** (pipeline-slot ExternalRealization PR scope) 6. **R3 Substrate Mgr (warm-wolf-698)** dispatches Step 2 worker 7. **Director ratifies Step 2 PR + admin-merges** when CI clears 8. **Director authors PB-3 Step 3a worker brief** (grammar table extension — unblocked NOW) 9. **R3 Substrate Mgr** dispatches Step 3a worker 10. **Director ratifies Step 3a PR**, admin-merges -11. **WAIT for substrate-capability landing** (recursive list-body emission per Q1) -12. **Director authors PB-3 Step 3b worker brief** (full parser body `.dag` migration) -13. **R3 Substrate Mgr** dispatches Step 3b worker -14. **Director ratifies Step 3b PR**, admin-merges -15. **Director authors PB-3 Step 4 worker brief** (parity test + simultaneous parser body deletion) -16. **R3 Substrate Mgr** dispatches Step 4 worker; parity against `parse_generated.rs` parse() body OUTPUT -17. **Director ratifies Step 4 PR**, admin-merges → parser body DELETED → PB-3 gate row CLOSES +11. **WAIT for substrate-capability landing** (recursive list-body emission per §12 Q1; this wait dissolves only if Q1 ratifies path (b) "bundled" — see Q1 + §9 row note. Default is path (a) wait until Q1 is operator-ratified.) +12. **Director authors PB-3 Step 3b/4 COMBINED worker brief** — single atomic worker scope per §9 row: `src/v3/std/parse.dag` body + parity `TestClaim` + `parse_generated.rs:138` parse() body deletion. Per openai-pro PR #3126 BLOCKING (sha 16d21f4a) Finding 6, this replaces the previously-listed separate Step 3b authoring + Step 3b ratify + Step 4 authoring + Step 4 ratify steps; no inter-PR Rust+`.dag` coexistence is permitted (paper-shrink-relocation risk per `feedback_paper_shrink_variants`). +13. **R3 Substrate Mgr** dispatches the COMBINED Step 3b/4 worker +14. **Director ratifies COMBINED Step 3b/4 PR**, admin-merges → `.dag` parser body lands + parity TestClaim asserts equivalence + Rust parser body DELETED in the same merge → PB-3 gate row CLOSES Subsequent L2.5 model (PB-2 tokenize) follows same sequence with its own substrate-capability dependencies. @@ -384,8 +405,8 @@ Subsequent L2.5 model (PB-2 tokenize) follows same sequence with its own substra **Memory disciplines applied**: - `feedback_lenses_not_passes` (parse = substrate-table-driven dispatch, NOT decision engine) -- `feedback_fail_closed_discipline` C-8 (ParseDiagnostic coupled INTO SurfaceModule) -- `feedback_state_space_vs_behavioral_invariants` (typed-state SurfaceModule at output) +- `feedback_fail_closed_discipline` C-8 (parse returns `Result` per §4.3 + live `parse_generated.rs:138` — fail-closed Result-sum, NOT typed-state-with-coupled-diagnostics; distinct from lower/infer pattern per cross-stage discriminator at §4.3) +- `feedback_state_space_vs_behavioral_invariants` (parse output is `Result` — the type rules out partial-parse states by construction; SurfaceModule itself carries no diagnostic field per §4.3) - `feedback_target_agnostic_ir` (parse output carries no target-specific facts) - `feedback_paper_shrink_variants` (Step 4 parity = genuine deletion, not relocation) - `feedback_anchor_mgr_lane_synthesis_on_gap_tier_not_session_id` (Gap-tier anchors) diff --git a/docs/design-tokenize-stage-l25-model.md b/docs/design-tokenize-stage-l25-model.md new file mode 100644 index 00000000000..e78f015ed18 --- /dev/null +++ b/docs/design-tokenize-stage-l25-model.md @@ -0,0 +1,368 @@ +# Tokenize Pipeline Stage — L2.5 Domain Model (PB-2) + +**Status:** DRAFT — Director-tier authoring per operator ratification 2026-05-14 (Decision 1.A scoping = Option A). + +**Authoring date:** 2026-05-14. +**Authoring tier:** Director (zesty-bear-812). +**Lane:** PB-2 (tokenize) per `docs/design-pure-bootstrap-zero.md` + `src/v3/SELF_HOSTING.md` §2 4-step migration discipline. +**Migration order rank:** 5th + LAST in pipeline-stage sequence (per `docs/substrate-reflection-design.md` §12.6 — emit → lower → infer → parse for the 4 pipeline-stage migrations explicitly tabled there; `docs/design-pure-bootstrap.md` §"PB-2 — tokenize retire" (line ~134) extends with tokenize). +**Routing authority chain:** operator-ratification + PM-delegate (per 2026-05-14 directive) → PM amends close plan + §1.8 PB-2 gate row → Director authors per-step worker briefs → R3 Substrate Mgr (warm-wolf-698) dispatches workers. + +--- + +## §1 Purpose + scope + +This document is the **Step 1 model review** per `src/v3/SELF_HOSTING.md` §2.2 4-step discipline applied to PB-2 tokenize-stage migration. + +**Distinct from sibling pipeline-stage L2.5s**: PB-2 tokenize is the FURTHEST-ALONG pipeline stage. Substrate authority already lives in `.dag`: +- `src/v3/std/tokenize.dag` (143 lines; Token / TokenKind taxonomy as shared std/ vocabulary) +- `src/v3/compiler/tokenize.dag` (154 lines; tokenizer implementation in `.dag`) +- `src/v3/compiler/src/tokenize_generated.rs` (362 lines; AUTO-GENERATED via `regen_tokenize` codegen-driver from `tokenize.dag`) + +PB-2 migration is FURTHER ALONG than other pipeline stages BUT NOT complete (per codex BLOCKING #3127 — corrected understanding). Substrate authority exists in `.dag` AND has documented open scaffolds requiring substantive retirement work: + +1. **SG-1a tracked scaffold** (`src/v3/compiler/tokenize.dag:16-22`): `regen_tokenize` currently parses raw source text for `dag_keyword_set` / `dag_operators` because those shared-syntax bodies lower as `ValueBody::Unparsed`. Named dissolution trigger: once those bodies lower structurally under `compile_to_dag`, delete the raw-text extractor + derive directly from lowered Dag in same PR. PB-2 Step 4 carries this scaffold-retirement scope. + +2. **Character-level under-consumption scaffold** (`tokenize.dag:23-30+`): scan phases slice ASCII/Unicode codepoint space in parallel forms NOT consuming `dsl/std/` character authorities. `StringEscapeSpec` / `LocalPunctSpec.pattern` / `string_literal_delimiter` treated as opaque Strings; hidden Rust character predicates (`byte.is_ascii_digit()` / `byte.is_ascii_lowercase()` etc. at `tokenize_generated.rs:15-22`) emitted into codegen rather than driven by `.dag`. **Named dissolution trigger** (mirror of the SG-1a scaffold's shape, per openai-pro PR #3127 BLOCKING sha d15e1f29 fixed in PR #3138): when (a) scanner classes / string escape / local punctuation consume `dsl/std/unicode.dag` `CharClass` + `char_in_class` authorities structurally — i.e. when `StringEscapeSpec.suffix` retypes to `Char`, `LocalPunctSpec.pattern` to `List`, `string_literal_delimiter` to `Char`, etc. — AND (b) `tokenize_generated.rs` no longer emits hidden `byte.is_ascii_*` predicates because the codegen-driver reads class-predicate facts structurally from lowered `tokenize.dag`, delete the parallel character-predicate scaffold in the same PR (no Rust-and-`.dag` coexistence per `feedback_paper_shrink_variants`). Same-PR receipt: the deletion lands in the same PR that flips the substrate consumption. Gating substrate prereqs are tracked at §9 Step 4: ROADMAP.md:467 (Character-level under-consumption row) + ROADMAP.md:416 (Class 5 Gap 3 top-level `ValueBody` boundary) + `std.unicode` bootstrap/load-set decision. PB-2 Step 4 carries this scaffold-retirement scope. + +3. **Residual hand-Rust**: NOT just the codegen artifact `tokenize_generated.rs`. The actual residual includes (a) `regen_tokenize` codegen-driver logic (per Q1 PB-Bootstrap-Process lane), (b) SG-1a raw-text extractor scaffold, (c) character-predicate scaffold leaking through codegen. + +So PB-2 is "substrate-driven by `.dag` AT scanner-state-machine level but NOT at character-predicate level"; the doc's earlier "mostly verification" framing UNDERSTATED the substantive scaffold-retirement work. + +**This doc does NOT**: +- Re-author the tokenize substrate (already live in `tokenize.dag`) +- Own retirement of the `regen_tokenize` codegen-driver itself (that's PB-Bootstrap-Process lane scope — codegen-driver retirement is cross-cutting) +- Own implementation of bootstrap-runtime-loop concerns (separate lanes) + +**Authority chain**: Director-tier ratification grounds the model; subsequent worker briefs cite this doc; §1.8 PB-2 gate row close-criterion predicate cites this doc as L2.5 authority. + +--- + +## §2 What tokenize IS structurally + +Per the live substrate at `src/v3/std/tokenize.dag:1-15`: + +**Tokenize is a substrate-driven scanner: `String → List`, dispatching on byte character class (Whitespace / Digit / IdentStart / IdentContinue) per `tokenize.dag` declarations. Token taxonomy + scanner logic both live in `.dag` substrate; `tokenize_generated.rs` is the codegen artifact from `regen_tokenize`.** + +The substrate authority is structured as: +1. **Token taxonomy** in `src/v3/std/tokenize.dag` (shared vocabulary for compiler + future user-space tooling) +2. **Scanner implementation** in `src/v3/compiler/tokenize.dag` (declarative scanner-class + token-kind-recognition tables) +3. **Codegen artifact** in `src/v3/compiler/src/tokenize_generated.rs` (auto-regenerated; do NOT hand-edit) + +Per `feedback_lenses_not_passes`: tokenize is a fold over bytes with scanner-class dispatch — substrate-driven, NOT decision engine. The class definitions IN tokenize.dag ARE the rule book; no separate "TokenizationSpec" carrier needed (substrate IS the spec, same pattern as PB-5 infer per PR #3085 §3.2). + +**Failure shape**: fail-closed (per `feedback_fail_closed_discipline` + INVARIANTS C-8). Live state at `tokenize_generated.rs:96`: `fn tokenize(...) -> Result, Diagnostic>` returns generic `Diagnostic::TokenizerError { message, span, correction }` (per `src/v3/std/diagnostics.dag` Diagnostic carrier). PROPOSED extension (NOT currently live): typed `TokenizeDiagnostic` variants (unterminated string literal / invalid character / numeric literal overflow / etc.) per §4.2 — same shape as PB-3/4/5 per-stage diagnostic extension per PR #3077 §12 Q7 ratification path. + +--- + +## §3 Input types (declared in `.dag` substrate) + +### §3.1 `String` (source text) + +Plain UTF-8 source text. No structural pre-processing beyond byte access. + +**Substrate authority**: `String` is primitive per `src/v3/std/substrate.dag`. No new carrier needed. + +**Lane dependency**: none upstream of tokenize — it's the pipeline entry point. + +### §3.2 No GrammarSpec / TokenizationSpec — substrate IS the spec + +Per `feedback_lenses_not_passes`: tokenize.dag declarations (scanner classes + token-kind tables) ARE the tokenization rule-book. No separate "TokenizationSpec" carrier needed (same pattern as PB-5 infer's no-InferenceSpec per PR #3085 §3.2). + +Distinct from PB-3 parse which has explicit GrammarSpec per Decision 3.B (b) compile-time tables: parse's grammar tables are richer (6 SG-2c table-families covering precedence / dispatch / bracket-roles / etc.); tokenize's scanner is simpler (byte → character class → token kind) and the rule book is implicit in the substrate file structure. + +--- + +## §4 Output types (declared in `.dag` substrate) + +### §4.1 `List` (typed-state output) + +Per `src/v3/std/tokenize.dag` (live; Token + TokenKind closed-axis sum): +- Token carries `kind: TokenKind` + `span: SourceSpan` (verified at `src/v3/std/tokenize.dag:65-67`; the live carrier is 2 fields only, no optional lexeme on Token itself) +- TokenKind variants carry their own payloads: `Ident(String)`, `IntLit(String)`, `StringLit(String)`, etc. — lexeme-content lives ON the variant, NOT on Token +- TokenKind closed-axis: keywords (KwLet/KwIf/KwThen/...), identifiers (Ident), literals (IntLit/StringLit), operators (Eq/EqEq/Lt/Gt/...), punctuation (LParen/RParen/...), etc. + +**Construction-time invariant**: every consumed byte advances the scanner position; every emitted Token carries source-span provenance. No silent byte-skipping; no fabricated tokens. + +**Substrate authority**: `src/v3/std/tokenize.dag` (LIVE at HEAD). PB-2 output type is stable. + +### §4.2 `TokenizeDiagnostic` (substrate extension per Decision 2.B / PR #3077 §12 Q7) + +Same cross-stage Decision 2.B framing as PB-3/4/5: per-stage diagnostic variants attach via whichever path PR #3077 §12 Q7 ratifies (carrier-field vs lane-local-sum). + +``` +// Typed reference carrier (cross-stage discipline per openai-pro +// PR #3077 BLOCKING + INVARIANTS P2/P3) — bound to the live .dag +// character-class authority `dsl/std/unicode.dag:62`: +// type CharClass = Whitespace | Digit | IdentStart | IdentContinue +// consumed at `src/v3/compiler/tokenize.dag:103`: +// data ascii_scan_order: List = [Whitespace, Digit, IdentStart, IdentContinue] +// There is NO `ScannerCharClass` declaration in any .dag file; an +// earlier draft of this section invented that name from the generated +// Rust enum spelling (codex PR #3138 BLOCKING sha f08b9525 Finding 1 +// fixed in PR #3138). TokenizeDiagnostic references the live `CharClass` +// directly. + +type TokenizeDiagnostic + = UnterminatedStringLiteral { opener_span: SourceSpan } + | InvalidCharacter { byte: Nat, span: SourceSpan, expected_class: CharClass } + | NumericLiteralOverflow { lexeme: NonEmptyStr, span: SourceSpan } + | (additional variants per Step 2 worker brief authoring against tokenize_generated.rs) +``` + +**Lane dependency**: PR #3077 §12 Q7 ratification; Director-tier per-stage variant authoring. + +### §4.3 Tokenize output is `Result, TokenizeDiagnostic>` — NO `TokenizedSource` extension + +> **Codex PR #3127 BLOCKING (sha d15e1f29) revision**: an earlier draft of this section proposed a `TokenizedSource { tokens, diagnostics }` extension parallel to PR #3126's earlier-draft `SurfaceModule { items, diagnostics }`. Both proposals are REJECTED for the same reason: tokenize and parse sit in the **fail-fast output domain**, not the structural-output domain. The cross-stage discriminator (re-stated in PR #3138's §4.3 reframe of the parse doc) places tokenize alongside PB-3 parse + PB-6 emit: a partial token list with a corrupt token in the middle is not a valid downstream input for parse, so the failure couples via `Result`, not into the structural carrier. No `TokenizedSource` carrier is being authored; §12 Q6 (TokenizedSource carrier shape) is resolved as "rejected — Result-sum per cross-stage discriminator". + +**Live state matches**: `tokenize_generated.rs:96` already returns `pub fn tokenize(source: &str, file: &str) -> Result, Diagnostic>`. The L2.5 ratifies this shape rather than replacing it. + +Step 2 signature: `fn tokenize(source: String) -> Result, TokenizeDiagnostic>` — `List` is the Ok-branch payload, `TokenizeDiagnostic` (§4.2 substrate extension) is the Err-branch variant. This is the **single canonical boundary carrier** between PB-2 tokenize and PB-3 parse: `List` on the Ok branch, no parallel `TokenizedSource`. + +**Cross-stage discriminator** (consistent with PR #3138 parse L2.5 §4.3): +- **Result-sum** (PB-2 tokenize + PB-3 parse + PB-6 emit): fail-fast output domain — partial token/parse/emit output is not a valid intermediate. +- **Typed-state-with-coupled-diagnostics** (PB-4 lower + PB-5 infer): structural output domain — Unresolved ports / pre-inferred Dag are valid intermediates consumed downstream. + +--- + +## §5 Substrate-driven tokenization (the core) + +Per `tokenize_generated.rs:6-25` + `tokenize.dag` declarations: + +### §5.1 Byte → `CharClass` dispatch + +Each byte maps to one of the four `CharClass` variants declared at `dsl/std/unicode.dag:62` (`type CharClass = Whitespace | Digit | IdentStart | IdentContinue`) and consumed at `src/v3/compiler/tokenize.dag:103` (`data ascii_scan_order: List`): +- `Whitespace` (tab/newline/form-feed/carriage-return/space) +- `Digit` (ascii digit) +- `IdentStart` (ascii letter or underscore) +- `IdentContinue` (alphanumeric or underscore) + +This dispatch is purely structural — a byte-class lookup function. Live at `tokenize_generated.rs:13-25` (regenerated from tokenize.dag). The substrate authority is the `CharClass` declaration in `unicode.dag`, NOT a `ScannerCharClass` (no such .dag type exists; that name appears only in the generated Rust enum spelling). + +### §5.2 `CharClass` → token-recognition state machine + +Per `tokenize.dag` declarations: byte sequences matching specific patterns produce specific TokenKind variants. The state machine is small + declarative: +- Whitespace sequences: skipped (no token emitted) +- Digit sequences: collected → `IntLit(decimal_string)` +- IdentStart + IdentContinue sequences: collected → either keyword (per closed-axis keyword table) or `Ident(string)` +- String delimiters: collected with escape handling → `StringLit(string)` +- Operator symbols: matched against closed-axis operator-symbol table → operator TokenKind variants +- Punctuation symbols: matched → punctuation TokenKind variants + +### §5.3 Mechanical dispatch via closed-axis enums + +Walker dispatch is **mechanical**: match on byte class → look up the recognition rule → emit Token or continue scanning. **No conditional logic encoded in tokenize body** beyond the state-machine structural transitions; the recognition tables (keyword set / operator-symbol set / etc.) are closed-axis enums in tokenize.dag. + +--- + +## §6 Substrate prereqs (per-Gap-tier anchored) + +| Prereq | Substrate authority | Gap-tier lane | Status at HEAD (as of 2026-05-14) | +|---|---|---|---| +| Token + TokenKind taxonomy | `src/v3/std/tokenize.dag` (Token + TokenKind closed-axis sum) | PB-Substrate | LIVE at HEAD; complete | +| Tokenizer implementation | `src/v3/compiler/tokenize.dag` (scanner-class + recognition tables) | PB-Substrate | LIVE at HEAD; 154 lines | +| Codegen pipeline | `regen_tokenize` codegen-driver → `tokenize_generated.rs` | PB-Bootstrap-Process lane | LIVE at HEAD; codegen-driver retirement is PB-Bootstrap-Process scope (NOT PB-2 scope) | +| TokenizeDiagnostic substrate extension | extension of `src/v3/std/diagnostics.dag:150` per PR #3077 §12 Q7 | PB-Substrate + Director-tier per-stage authoring | Carrier LIVE; per-stage variant authoring NEW per Q7 ratification | + +**Critical observation**: PB-2 tokenize has LIGHTER prereq surface than other pipeline stages — substrate-driven at scanner-state-machine level. BUT substantive residual work remains (per codex BLOCKING #3127 corrected): SG-1a raw-text-extractor scaffold + character-level under-consumption scaffold + regen_tokenize codegen-driver retirement (per Q1 PB-Bootstrap-Process). Earlier "mostly verification" framing was understated. + +--- + +## §7 Cross-stage coordination + +### §7.1 Upstream dependencies + +None. Tokenize is the pipeline entry point (consumes raw source text). + +### §7.2 Downstream consumers + +PB-3 parse consumes `List` from tokenize (per PR #3126 §3.1). The Token carrier shape is stable (live at tokenize.dag); PB-3 parse migration is independent of PB-2 tokenize migration status. + +Per Decision 2.B / PR #3077 §12 Q7: tokenize diagnostics are discriminable by source per the ratified extension path. + +### §7.3 Sibling-stage coordination + +Cross-stage discipline (Ok-branch propagation; Err branches are stage-terminal fail-fast per §4.3 Result-sum discriminator): **tokenize** → `List` (Ok of `Result, TokenizeDiagnostic>`) → parse → `SurfaceModule` (Ok of `Result`) → lower → `PreInferDag` (typed-state with coupled diagnostics; structural output domain) → infer → `InferredDag` (typed-state) → emit → `Result` (fail-fast like tokenize/parse). The chain shows the success-path data flow; on any stage's Err branch the pipeline aborts at that stage (no partial-output propagation across boundaries). + +Tokenize is the FOUNDATION; its output type stability affects every downstream stage. Since `tokenize.dag` is already the live substrate authority, this stability is preserved across PB-2 migration. + +--- + +## §8 Two shapes of omni-emission — N/A for tokenize + +Tokenize is target-agnostic byte-scanning; Shape A/B disambiguation lives at emit stage. PB-2 has no Shape A/B framing. + +--- + +## §9 SELF_HOSTING.md §2.2 4-step applied to PB-2 tokenize + +**The 4-step discipline applies UNUSUALLY for PB-2 because substrate is already live**: + +| Step | Deliverable | Owner | Substrate | +|---|---|---|---| +| **Step 1: Model review** | THIS DOC | Director (zesty-bear-812) | docs/design-tokenize-stage-l25-model.md (this doc) | +| **Step 2: Pipeline slot** | `fn tokenize(source: String) -> Result, TokenizeDiagnostic>` declared in `src/v3/compiler/pipeline.dag` (per dsl/gunbc/compiler.dag:24 — internal pipeline lives in pipeline.dag, NOT generic compiler.dag) with `ExternalRealization` body. **Single canonical boundary carrier**: `List` on the Ok branch — no parallel `TokenizedSource` wrapper, no `diagnostics` field added to tokenize output (§4.3 Result-sum disposition; §12 Q6 resolved REJECTED). Matches live `tokenize_generated.rs:96` `Result, Diagnostic>` shape; only refinement is per-stage `TokenizeDiagnostic` variant typing (§4.2). Note: `tokenize.dag` exists as scanner-state-machine substrate, BUT residual hand-Rust includes (a) `regen_tokenize` codegen-driver logic (P5 deferral receipt: `docs/design-pure-bootstrap-zero.md:116` PB-Bootstrap-Process lane + ROADMAP.md:467 character-level row), (b) SG-1a raw-text-extractor scaffold for `dag_keyword_set`/`dag_operators` per `tokenize.dag:16-22`, (c) character-predicate scaffold (`byte.is_ascii_digit()` etc. at `tokenize_generated.rs:15-22`) leaking through codegen (P5 receipt: ROADMAP.md:467 phase-2 retype, gated on Class 5 Gap 3 + std.unicode bootstrap/load-set decision). Step 4 carries scaffold-retirement scope, NOT just codegen-artifact retirement. | R3 Substrate Mgr (warm-wolf-698) — worker dispatched against Director-authored Step 2 brief | pipeline.dag refinement | +| **Step 3: Verify substrate completeness** | Audit `src/v3/compiler/tokenize.dag` vs `tokenize_generated.rs` — confirm `.dag` is the complete authority (no hand-Rust logic in `tokenize_generated.rs` beyond mechanical codegen artifacts); identify any residual hand-Rust scaffolding that needs retirement. Per `feedback_paper_shrink_variants` discipline: verify tokenize.dag is NOT V1 template-relocation (hand-Rust scanner logic relocated to `.dag` text without substrate-substance growth). | R3 Substrate Mgr — worker dispatched against Director-authored Step 3 brief | tokenize.dag audit | +| **Step 4: Retire residual hand-Rust + codegen-driver decoupling** | If Step 3 reveals residual hand-Rust scaffolding: retire it. `regen_tokenize` codegen-driver retirement is **NOT** in PB-2 scope — it routes to the PB-Bootstrap-Process lane at `docs/design-pure-bootstrap-zero.md:116` (author `bootstrap.dag` + generated trampoline; sized M; verification gates per `docs/design-pure-bootstrap-zero.md:118-123`). Phase-2 char-class retype (`Char` / `List` / `CharClass`) routes to ROADMAP.md:467 ("Character-level under-consumption in tokenize + syntax authorities"), gated on **Class 5 Gap 3** (top-level `ValueBody` boundary at ROADMAP.md:416) plus the std.unicode bootstrap/load-set decision. Overall non-test census reaches 0 per ROADMAP.md:53 T-PB-A. PB-2 Step 4 retires only what closes in-lane; the named cross-lane receipts above are the P5 deferral. EXPECTED_HAND_AUTHORED_NON_TEST shrinks by whatever residual lands. | R3 Substrate Mgr + coordination with PB-Bootstrap-Process lane | residual hand-Rust deletion | + +**Critical**: PB-2's Step 3 is VERIFY not AUTHOR (substrate already exists). PB-2's Step 4 is HANDOFF/RETIRE not PORT (no porting needed beyond what's already in tokenize.dag). + +Per `feedback_paper_shrink_variants` discipline applied at Step 3 audit: verify tokenize.dag is NOT V1 template-relocation. The discriminator: does tokenize.dag declare scanner-class + recognition tables as SUBSTRATE DATA (declarative tables), OR does it carry hand-Rust scanner code in text form? The former is substantive; the latter is paper-shrink. Per the live tokenize.dag content I grep'd (substrate-style declarations), it appears substantive — but Step 3 audit confirms formally. + +--- + +## §10 Determinism invariant preservation + +Tokenization is inherently deterministic (byte stream → unique token stream per scanner state machine). Already structurally enforced via the `.dag` substrate's declarative scanner-class definitions. No HashMap iteration concerns. + +--- + +## §11 Tokenization invariants (cross-cutting) + +Per `feedback_fail_closed_discipline` C-8 + scanner hygiene: + +- **Every byte consumed**: no silent byte-skipping; whitespace explicitly skipped per scanner state +- **Source-span provenance**: every Token + every TokenizeDiagnostic carries `SourceSpan` +- **Lookahead bounded**: scanner uses constant-bounded lookahead (no unbounded backtracking) +- **Anti-bridge invariant** (per `feedback_no_textual_enforcement_bridges`): tokenize.dag declarations are SINGLE authority; no fallback hand-Rust heuristics in tokenize_generated.rs + +--- + +## §12 Open design questions (operator/PM ratification) + +### Q1: codegen-driver retirement scope + +`regen_tokenize` codegen-driver lives at the build-system layer. Per `docs/design-pure-bootstrap-zero.md` PB-Bootstrap-Process lane: codegen-driver retirement is cross-cutting across all pipeline stages (tokenize / parse / etc.). Question: does PB-2 retire `regen_tokenize` specifically, OR does PB-Bootstrap-Process retire all codegen-drivers in one cross-stage PR? + +**Director-recommend: PB-Bootstrap-Process retires all codegen-drivers** as one coherent lane — keeping codegen-driver retirement to a single substrate-cross-cutting authority avoids per-stage paper-shrink risk (`feedback_paper_shrink_variants`). PB-2 confirms tokenize.dag is the substantive authority + flags `regen_tokenize` as retire-target, but the retirement PR lives in PB-Bootstrap-Process. Operator/PM ratification. + +### Q2: Step 3 audit scope — paper-shrink check + +Per `feedback_paper_shrink_variants` discipline + my Director-side history of missing template-relocation in cycles 3/4/5/6: Step 3 audit must explicitly check whether `tokenize.dag` is substantive substrate (declarative tables + scanner-class definitions) or paper-shrink (hand-Rust scanner code in `.dag` text). + +**Director-recommend**: Step 3 brief explicitly enumerates the audit dimensions: +1. Scanner-class definitions are declarative byte-pattern membership (not hand-coded match-arms in text form) +2. Recognition tables are closed-axis enums (not String-keyed maps) +3. State machine transitions are structural (not embedded code) +4. No `pub mod tokenize { ... }` absorption into adjacent files (V2 module-relocation check) + +If audit fails: tokenize.dag may itself need refactoring before retirement-of-residual-Rust is meaningful. Operator/PM ratification on audit criteria. + +### Q3: TokenizeDiagnostic variant exhaustiveness + +Step 2 worker brief enumerates the full variant set by grepping `tokenize_generated.rs` for Diagnostic construction sites. Same discipline as PB-3 §12 Q3. + +**Director-recommend**: defer to Step 2 worker brief authoring (consistent with PB-3 / PB-4 / PB-5 approach). + +### Q4: Substrate completeness criterion + +What's the formal predicate for "tokenize substrate is complete + residual hand-Rust can retire"? + +**Director-recommend**: predicate spans BOTH the codegen artifact AND the codegen-driver boundary (per codex BLOCKING PR #3127 — earlier scoping to tokenize_generated.rs alone missed the regen_tokenize logic + .dag authority boundary). + +Predicate = `cargo test --release -p v3-compiler --test integration tokenize_substrate_authority` shows: +- (a) `tokenize.dag` content unchanged but codegen regenerates `tokenize_generated.rs` byte-identically (idempotent codegen) +- (b) `tokenize_generated.rs` contains NO hand-edit zones (all body is codegen-driver-emitted) +- (c) No fallback Rust scanner logic outside the codegen artifact +- (d) **`regen_tokenize` codegen-driver itself contains NO scanner-logic decisions** — it reads `tokenize.dag` declaratively and emits Rust mechanically. If `regen_tokenize` carries scanner logic (rather than just template-rendering substrate facts), the substrate isn't actually complete: the driver IS hand-Rust scanner logic in disguise. +- (e) **Or explicit ROADMAP.md deferral row** naming `regen_tokenize` codegen-driver retirement scope as PB-Bootstrap-Process lane (per Q1); deferral receipt is PB-2 → PB-Bootstrap-Process handoff at the codegen-driver authority boundary. + +If (a)(b)(c)(d) hold, substrate is complete + tokenize_generated.rs can retire (replaced by Evaluator-loaded `.dag` at runtime per PB-Runtime). If (d) is violated but (e) is named, partial-completion with explicit deferral is acceptable (per `feedback_paper_shrink_variants` P5 receipt discipline). + +### Q5: PB-3 parse landing dependency + +PB-3 parse's input is List from tokenize. **Does PB-2 tokenize migration block on PB-3 parse migration completing? Reverse?** + +**Director-recommend: NO bidirectional blocking** — Token carrier shape is stable across both migrations. Same independence pattern as PB-4 lower vs PB-3 parse per PR #3077 §12 Q6 + PB-5 infer vs PB-4 lower per PR #3085 §12 Q6. + +### Q6: TokenizedSource carrier shape — RESOLVED REJECTED (codex PR #3127 BLOCKING) + +**Disposition**: REJECTED. No `TokenizedSource` carrier is being authored. + +**Why**: this Q was opened on the (incorrect) premise that tokenize, like PB-4 lower / PB-5 infer, should couple diagnostics structurally into its output carrier. The cross-stage discriminator (§4.3, mirrored in PR #3138 parse-L2.5 §4.3) classifies tokenize alongside PB-3 parse + PB-6 emit in the **fail-fast output domain**: a partial token list with a corrupt token in the middle is not a valid intermediate for parse, so the failure couples via `Result`, not into the structural carrier. Live shape at `tokenize_generated.rs:96` already matches this disposition (`Result, Diagnostic>`); the L2.5 ratifies the live shape rather than proposing an extension. + +**Step 2 boundary carrier** is therefore unique: `Result, TokenizeDiagnostic>`, with `List` on the Ok branch and no parallel `TokenizedSource`. No PM/operator ratification is needed beyond confirming the discriminator, which is also load-bearing in PR #3138. + +--- + +## §13 Non-goals + +- **`.dag` implementation of tokenizer** — already exists at `src/v3/compiler/tokenize.dag` (LIVE) +- **`regen_tokenize` codegen-driver retirement** — PB-Bootstrap-Process lane scope +- **Test corpus design + parity-test harness implementation** — Step 4 work +- **Bootstrap-runtime-loop concerns** — separate lanes +- **PB-3 parse migration** — separate L2.5 doc (PR #3126) +- **Shape A/B emission** — emit's concern, not tokenize + +--- + +## §14 Acceptance criteria for this L2.5 model + +This doc lands on main when: + +1. ✅ Input types declared structurally with substrate paths (§3) +2. ✅ Output types declared with live substrate citations (§4) +3. ✅ Substrate-driven tokenization composed without decision logic (§5 — per `feedback_lenses_not_passes`) +4. ✅ All substrate prereqs named with Gap-tier / Mgr-lane anchors (§6) +5. ✅ Cross-stage dependencies explicit (§7) +6. ✅ N/A — Shape A/B framing irrelevant for tokenize (§8) +7. ✅ SELF_HOSTING.md §2.2 4-step applied (§9) — with PB-2-specific adjustments (Step 3 = verify, Step 4 = handoff/retire) +8. ✅ Determinism preservation discipline (§10) +9. ✅ Tokenization invariants explicit (§11) +10. ✅ Open design questions enumerated for operator/PM ratification (§12) +11. ⏳ Operator/PM ratification on §12 Q1–Q5 only (Q6 resolved REJECTED in §12 per codex PR #3127 BLOCKING + cross-stage Result-sum discriminator — no ratification needed; see §14 "Surfaces awaiting" + §15 step 1 for the same scoping) + +Post-ratification: this doc becomes substrate authority for Step 2/3/4 worker brief authoring + §1.8 PB-2 gate row close-criterion predicate. + +--- + +## §15 Authoring sequence post-ratification + +1. **Operator / PM-delegate ratifies §12 Q1–Q5** only (per 2026-05-14 directive; §12 Q6 was resolved REJECTED in-doc per codex PR #3127 BLOCKING + cross-stage Result-sum discriminator and needs no separate ratification cycle) +2. **PM amends close plan + §1.8** to route through PB-X lanes + cite this doc as PB-2 L2.5 substrate +3. **PR #3077 §12 Q7 ratifies** (cross-stage Decision 2.B extension path; affects TokenizeDiagnostic shape) +4. **Director authors PB-2 Step 2 worker brief** (pipeline-slot ExternalRealization PR scope; trivial since substrate already lives in tokenize.dag) +5. **R3 Substrate Mgr (warm-wolf-698)** dispatches Step 2 worker +6. **Director ratifies Step 2 PR + admin-merges** +7. **Director authors PB-2 Step 3 worker brief** (substrate-completeness audit per §12 Q2) +8. **R3 Substrate Mgr** dispatches Step 3 worker +9. **Director ratifies Step 3 PR**, admin-merges → audit findings inform Step 4 scope +10. **Director authors PB-2 Step 4 worker brief** (residual hand-Rust retirement; coordinates with PB-Bootstrap-Process lane per Q1 if codegen-driver retirement is bundled) +11. **R3 Substrate Mgr** dispatches Step 4 worker +12. **Director ratifies Step 4 PR**, admin-merges → PB-2 gate row CLOSES per §1.8 + +--- + +## §16 Cross-references + +**Primary authority**: +- `src/v3/SELF_HOSTING.md` §2.2 (4-step migration discipline) +- `docs/design-pure-bootstrap-zero.md` (PB-X lane framing; PB-Bootstrap-Process lane for codegen-driver retirement) +- `docs/substrate-reflection-design.md` §12.6 (migration order) +- `docs/design-emit-stage-l25-model.md` (PB-6 PR #3066 — sets L2.5 template) +- `docs/design-lower-stage-l25-model.md` (PB-4 PR #3077 — cross-stage diagnostic pattern) +- `docs/design-infer-stage-l25-model.md` (PB-5 PR #3085 — substrate-driven dispatch precedent) +- `docs/design-parse-stage-l25-model.md` (PB-3 PR #3126 — sibling pipeline-stage L2.5) + +**Live substrate referenced**: +- `src/v3/std/tokenize.dag` (Token + TokenKind taxonomy; LIVE; 143 lines) +- `src/v3/compiler/tokenize.dag` (tokenizer implementation; LIVE; 154 lines) +- `src/v3/compiler/src/tokenize_generated.rs` (AUTO-GENERATED codegen artifact; 362 lines) +- `src/v3/std/diagnostics.dag:150` (Diagnostic carrier — extends with TokenizeDiagnostic per PR #3077 §12 Q7 ratification) + +**Memory disciplines applied**: +- `feedback_lenses_not_passes` (tokenize = substrate-driven byte fold, NOT decision engine) +- `feedback_fail_closed_discipline` C-8 (tokenize returns `Result, TokenizeDiagnostic>` — fail-closed Result-sum like parse + emit; the Err branch is the only carrier for diagnostics, no structural coupling into `List`) +- `feedback_state_space_vs_behavioral_invariants` (tokenize output is `Result, TokenizeDiagnostic>` — the type rules out partial-tokenize states by construction; `List` itself carries no diagnostic field per §4.3) +- `feedback_target_agnostic_ir` (tokenize output carries no target-specific facts) +- `feedback_paper_shrink_variants` (Step 3 audit explicitly checks substrate vs paper-shrink) +- `feedback_anchor_mgr_lane_synthesis_on_gap_tier_not_session_id` (Gap-tier anchors) +- `feedback_no_textual_enforcement_bridges` (anti-bridge: tokenize.dag is single authority) +- `feedback_grep_carrier_semantic_before_ratification` (4-axis grep applied at authoring time) + +**Surfaces awaiting**: +- Operator/PM ratification on §12 Q1–Q5 (Q6 resolved REJECTED in-doc per codex PR #3127 BLOCKING — no ratification needed) +- PR #3077 §12 Q7 ratification (cross-stage Decision 2.B extension path) +- PM Phase 2 close plan + §1.8 amendments citing this doc +- Coordination with PB-Bootstrap-Process lane for codegen-driver retirement per Q1