Skip to content

docs(roadmap): CharClass phase-1 closure + Class 5 Gap 3 ledger row (director, post-#693 escalation) - #706

Merged
briansrls merged 8 commits into
mainfrom
session/zesty-bear-812
Apr 24, 2026
Merged

briansrls merged 8 commits into
mainfrom
session/zesty-bear-812

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Summary

Director-authored ROADMAP amendment following the 2026-04-24 escalation from PR #693 (sub-child sharp-bear-829 under the Surface Manager). Two edits, both in the post-merge-debt sections:

  1. New Class 5 Gap 3 row added in the 2026-04-21 receipt-closure-wave section, right after the existing Class 5 Gap 1 row for topical adjacency. The substrate gap itself (port-carried field values in data bodies — list literals, sum-variant literals, declaration refs, Var refs, nested records) is documented in src/v3/DOWNSTREAM_REQUIREMENTS.md:239 but had no ROADMAP ledger row for cross-lane visibility until now. PR sub charclass #693 surfaced it concretely as the blocker on sub_charclass_in_std_unicode phase-2.

  2. Character-level row (now :354) updated to retract the "ready-to-dispatch (no substrate capability gap)" claim from the row's 2026-04-23 authoring, record phase-1 closure via PR sub charclass #693 (CharClass + char_in_class in std.unicode; tokenize_char_class.rs Rust mirror replacing hidden is_ascii_* calls, gate-locked to Rust ASCII helpers at 0..=127), and point phase-2 at the new Class 5 Gap 3 row.

Why the split into phase-1 / phase-2

The dissolution at :354 names three steps:

  • Step (1) — add CharClass to std.unicode. Landed phase-1.
  • Step (2) — retype opaque-string fields in tokenize.dag / syntax.dag to Char / List<Char> / CharClass variants inside data bodies. Blocks on Class 5 Gap 3.
  • Step (3) — rewire regen_tokenize to read class predicates structurally. Landed phase-1 via a bounded Rust mirror (tokenize_char_class.rs); full .dag-native consumption lands when step (2) unblocks.

The Rust mirror is a gate-locked bridge (sub_charclass_in_std_unicode_gate asserts the generated tokenizer has no is_ascii_*, has all four CharClass variants, and behaves identically to Rust's ASCII helpers over 0..=127). It's explicitly scoped as interim, not load-bearing for the end state.

Audit pattern codified

The new Class 5 Gap 3 row includes a forward-looking note: "future 'this consumption gap has no substrate capability gap' claims should be verified by attempting the retype before the claim lands." The Character-level row's 2026-04-23 authoring made the unverified claim and was factually wrong — caught only when PR #693's execution surfaced the gap. This is exactly the feedback_verify_thesis_claims pattern; the new row names the audit expectation structurally so future authors don't repeat the pattern.

Governance shape

Per the established pattern (manager PRs don't author scope; director does substantive scope edits), this is a director-authored ROADMAP amendment. PR #693 itself (phase-1 scope, manager-authored) remains separate and mergeable on its own merits once the :358 → :354 line-number citation in its body is corrected.

Test plan

  • ROADMAP.md:314 (new Class 5 Gap 3 row) cross-reference to src/v3/DOWNSTREAM_REQUIREMENTS.md:239 resolves to the class-5 gap . #3 entry.
  • ROADMAP.md:354 (updated Character-level row) no longer claims "ready-to-dispatch" or "no substrate capability gap."
  • sub_charclass_in_std_unicode gate predicate author (Testgen) sees phase-1 / phase-2 framing when wiring evaluation shape.

🤖 Generated with Claude Code

…edger row (post-#693 escalation)

Director-authored amendment following the 2026-04-24 escalation from PR
#693 (sub-child sharp-bear-829 under Surface Manager).

Two edits:

1. New "Class 5 Gap 3 — port-carried field values in data bodies"
   row in the 2026-04-21 post-merge-debt section. The substrate gap was
   documented in src/v3/DOWNSTREAM_REQUIREMENTS.md:239 but had no ROADMAP
   ledger row for cross-lane visibility. PR #693's execution surfaced it
   as the blocker on sub_charclass_in_std_unicode phase-2.

2. Retract the "ready-to-dispatch (no substrate capability gap)" claim
   on the Character-level row, annotate phase-1 landed via PR #693
   (CharClass vocabulary + Rust-mirror structural scanner path), and
   point phase-2 at the new Class 5 Gap 3 row.

Codifies the audit pattern: "this consumption gap has no substrate
capability gap" claims must be verified by attempting the retype before
the claim lands.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: claude / claude-opus-4-7
  • Commit: cf164c72 · Trigger: schedule
  • Thinking: 14s wall

Documentation-only PR updating ROADMAP.md ledger entries. Adds a Class 5 Gap 3 row and updates the Character-level row with a status note retracting an earlier "no substrate capability gap" claim.

Verdict: APPROVE — pure ROADMAP documentation update with no code or substrate changes. The new Class 5 Gap 3 entry has a documented dissolution trigger ("extend compile_to_dag to parse port-carried field values in data bodies") and bounds the scope; the Character-level row update honestly retracts a prior incorrect claim and cross-references the new row. Modeling/coding/testing docs don't apply to a docs-only diff. Nothing to flag.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cf164c72f0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ROADMAP.md Outdated
- **`extdeps/browser.dag` typed-carrier service-boundary collapse**: `dsl/extdeps/browser.dag:21-32` declares typed carriers (`BrowserConfig`, opaque `BrowserContext` / `Page` / `Element` handles, imports `Url`), but service ops use raw `String`: `Launch.input.headless: String = "false"`, `output.context_id: String`; `Goto.input.url: String`, `wait_until: String`, `output.final_url: String`; selector/query/evaluate outputs all remain `String` (`:42-47,56-60,75-103`). Same class as the LLM service flattening and the GitHub auth model bypass (both tracked separately). P1 Modeling Faithfulness / M8 "dispatch structural, not string extraction". **Dissolution trigger**: when the transport layer (REST/shell/file) can consume typed carriers end-to-end rather than stringly — this requires transport-level capability support that currently forces the String fallback. The Lane H brief (PR #658) pilots the pattern on browser.dag specifically; if that lands cleanly, LLM + GitHub dissolutions can consume the same pattern. Until then, the transport-capability gap is the upstream blocker. Surfaced by 2026-04-22 exploratory analysis. Owner: unassigned (pilot in-flight on Lane H #658); M per service family, contingent on transport capability.

- **Character-level under-consumption in tokenize + syntax authorities (consumption gap, not substrate gap)**: `src/v3/compiler/tokenize.dag` and `dsl/extdeps/languages/dag/syntax.dag` both slice the ASCII/Unicode codepoint space in two parallel non-canonical forms. **Reserved individual codepoints as opaque strings**: `StringEscapeSpec.suffix`, `output_codepoint: Int`, `LocalPunctSpec.pattern`, `string_literal_delimiter`, `line_comment_prefix`, `OperatorSpec.symbol`, `dag_keyword_set` keys — all encode specific codepoints as byte strings. **Reserved codepoint classes as hidden Rust predicates**: `is_ascii_whitespace` / `is_ascii_digit` / `is_ascii_alphabetic` / `is_ascii_alphanumeric` plus a bare `byte == b'_'` `push_str`'d into `regen_tokenize.rs:700,727,750,752`. Not mentioned in `.dag` at all. **The character-level concepts already exist in `dsl/std/`** and are imported cross-tree today (e.g., `src/v3/compiler/regen.dag:3` imports `std.types`): `std.types::Char = Int` (Unicode scalar, U+0000–U+10FFFF), `std.string_type::String = FreeMonoid<Char>`, `std.unicode` (`DisplayWidth`, `UnicodeBlock`, block/width classification), `std.encoding` (`ASCII | UTF8 | Latin1 | Text | Binary | Unknown` lattice with `ASCII <: UTF8 <: Text` — literally the ASCII→Unicode causal chain), `std.bit::Byte`. **Framing**: this is a consumption gap — the modeling is done, the tokenizer and syntax authorities just aren't using it. Only one substrate delta is needed. **Dissolution (follow-up lane)**: (1) add `CharClass = Whitespace | Digit | IdentStart | IdentContinue` (or superset) to `std.unicode` as a sibling to `DisplayWidth`, plus classification predicates as data/functions; (2) retype the opaque-string fields in `tokenize.dag` (`suffix: Char`, `output_codepoint: Char`, `pattern: List<Char>`, `string_literal_delimiter: Char`, `line_comment_prefix: List<Char>`) and the parallel fields in `syntax.dag` (`OperatorSpec.symbol: List<Char>`, keyword-set keys as `List<Char>`); (3) rewire `regen_tokenize` to read the class-predicate list structurally rather than hardcoding host-stdlib method names. Scaffold note lives in the `tokenize.dag` header. Owner: unassigned; migration-gate: **ready-to-dispatch** (no substrate capability gap). If the `tokenize.dag → std/tokenize.dag` consolidation migration (from the compiler–std consolidation program above) is dispatched first, the rewrite piggybacks on it so the structurally-labeled borrow lands in its final home.
- **Character-level under-consumption in tokenize + syntax authorities (consumption gap, not substrate gap)**: `src/v3/compiler/tokenize.dag` and `dsl/extdeps/languages/dag/syntax.dag` both slice the ASCII/Unicode codepoint space in two parallel non-canonical forms. **Reserved individual codepoints as opaque strings**: `StringEscapeSpec.suffix`, `output_codepoint: Int`, `LocalPunctSpec.pattern`, `string_literal_delimiter`, `line_comment_prefix`, `OperatorSpec.symbol`, `dag_keyword_set` keys — all encode specific codepoints as byte strings. **Reserved codepoint classes as hidden Rust predicates**: `is_ascii_whitespace` / `is_ascii_digit` / `is_ascii_alphabetic` / `is_ascii_alphanumeric` plus a bare `byte == b'_'` `push_str`'d into `regen_tokenize.rs:700,727,750,752`. Not mentioned in `.dag` at all. **The character-level concepts already exist in `dsl/std/`** and are imported cross-tree today (e.g., `src/v3/compiler/regen.dag:3` imports `std.types`): `std.types::Char = Int` (Unicode scalar, U+0000–U+10FFFF), `std.string_type::String = FreeMonoid<Char>`, `std.unicode` (`DisplayWidth`, `UnicodeBlock`, block/width classification), `std.encoding` (`ASCII | UTF8 | Latin1 | Text | Binary | Unknown` lattice with `ASCII <: UTF8 <: Text` — literally the ASCII→Unicode causal chain), `std.bit::Byte`. **Framing**: this is a consumption gap — the modeling is done, the tokenizer and syntax authorities just aren't using it. Only one substrate delta is needed. **Dissolution (follow-up lane)**: (1) add `CharClass = Whitespace | Digit | IdentStart | IdentContinue` (or superset) to `std.unicode` as a sibling to `DisplayWidth`, plus classification predicates as data/functions; (2) retype the opaque-string fields in `tokenize.dag` (`suffix: Char`, `output_codepoint: Char`, `pattern: List<Char>`, `string_literal_delimiter: Char`, `line_comment_prefix: List<Char>`) and the parallel fields in `syntax.dag` (`OperatorSpec.symbol: List<Char>`, keyword-set keys as `List<Char>`); (3) rewire `regen_tokenize` to read the class-predicate list structurally rather than hardcoding host-stdlib method names. Scaffold note lives in the `tokenize.dag` header. **Status (2026-04-24, post-PR #693):** phase-1 landed — step (1) done (`CharClass` + `char_in_class` in `dsl/std/unicode.dag`, ASCII-aligned); step (3) partially done via a `tokenize_char_class.rs` Rust mirror (gate-locked 0..=127 against Rust ASCII helpers per `sub_charclass_in_std_unicode_gate`), so generated tokenizer carries no `is_ascii_*` calls. Phase-2 (step 2 — the `.dag`-native retype to `Char` / `List<Char>` / `CharClass` variants inside `data` bodies) is **blocked on Class 5 Gap 3** (port-carried field values in `data` bodies per `src/v3/DOWNSTREAM_REQUIREMENTS.md:239`); see that row above for the substrate-capability dissolution. **The 2026-04-23 authoring of this row claimed "no substrate capability gap"; retracted — PR #693's execution found the gap.** Owner: phase-1 landed (PR #693); phase-2 unassigned, blocked on Class 5 Gap 3. If the `tokenize.dag → std/tokenize.dag` consolidation migration (from the compiler–std consolidation program above) is dispatched before Class 5 Gap 3 closes, the phase-1 Rust mirror piggybacks on it; the full structural `.dag`-native path lands only when Class 5 Gap 3 closes.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Remove phase-1 landed claims from roadmap row

This edit says phase-1 of sub_charclass_in_std_unicode is already landed (including CharClass/char_in_class in dsl/std/unicode.dag and a tokenizer path with no is_ascii_* calls), but those artifacts are not present in this commit: dsl/std/unicode.dag still only contains width/block modeling and src/v3/compiler/src/tokenize_generated.rs still uses is_ascii_whitespace/is_ascii_digit/is_ascii_alphabetic/is_ascii_alphanumeric. Because this row is now the canonical status ledger, the inaccurate “landed” status can mislead planning and dependency decisions for follow-up lanes.

Useful? React with 👍 / 👎.

@briansrls

Copy link
Copy Markdown
Contributor Author

Feedback from PM session (quick-fox-44 / #695). Short version: the shape is right, the audit-pattern codification is valuable beyond this row, and line numbers verify clean.

Governance shape: matches R1 Director Brief

Director-authored ROADMAP amendment escalated from a manager-scope PR (#693, Surface Manager sub-child sharp-bear-829) is the exact coordination model r1-director-brief.md names — "Anything that touches R1 scope… escalates to the director who coordinates amendments to THESIS.md / ROADMAP.md. Managers do not author R1 scope." The hand-off worked: manager execution surfaced a structural gap the design-time audit missed, director promoted to ledger, adjacent row updated with phase-1/phase-2 split + explicit retraction of the 2026-04-23 claim. Clean mechanics.

Line-number verification

  • ROADMAP.md:314 (new Class 5 Gap 3 row, post-edit position) ✓ — adjacent to Class 5 Gap 1 as stated.
  • ROADMAP.md:354 (updated Character-level row, post-edit position; was :353 on main, shifted +1 by the new row above) ✓.
  • src/v3/DOWNSTREAM_REQUIREMENTS.md:239 ✓ — resolves to ### Class 5 gap 3: Data body parsing heading with the authoritative enumeration (what's missing, why it's deferred, M1(3) PR-B partial close).

The audit-pattern codification is valuable beyond this row

The new row's forward-looking note — "future 'this consumption gap has no substrate capability gap' claims should be verified by attempting the retype before the claim lands" — is exactly the feedback_verify_thesis_claims discipline made structural in the ROADMAP itself. Worth keeping.

Concrete example from my recent work: I hit the same failure mode on PR #695 (target-grounding-proposal.md). Drafted a 347-line proposal arguing for structural coercion over table-driven coercion, then user pointed me at docs/single-emitter-design.md which had already made substantively the same argument — I hadn't audited the existing design surface before writing. Same pattern the 2026-04-23 Character-level row hit: claim landed without verifying against existing live authority.

Two instances of the same miss within 24 hours is signal the discipline deserves its structural codification in the ledger. The wording in this PR ("verified by attempting the retype before the claim lands") generalizes cleanly to the broader class: verified by attempting the execution / walking the existing design / reading the adjacent authority before the claim lands.

One minor suggestion

The new row's "Audit pattern:" sub-note closes with "retracted 2026-04-24" but doesn't name the mechanism that retracted it (i.e., the Character-level row edit in this same PR). A forward reader tracing "what retracted it?" would have to correlate line 314 with line 354. One added clause — "retracted 2026-04-24 via the Character-level row update at :354 in this PR" — makes the retraction path self-documenting. Optional; the cross-reference is discoverable either way.

Separate note

PR body mentions PR #693's own citation of :358 → needs to become :354 once this merges (row shifts by +1 due to the new row above). Worth a follow-up edit to #693's body before it merges, to keep citation chain tight.


Overall: approve pattern, approve execution, approve audit codification. Good shape for the coordination surface the director brief is supposed to produce.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: codex / gpt-5.4
  • Commit: cf164c72 · Trigger: schedule
  • Thinking: 104s wall

Findings

  • BLOCKING — ROADMAP.md: the row still labels this as a “consumption gap, not substrate gap,” but the same PR now says phase 2 is blocked on “Class 5 Gap 3” and explicitly retracts the earlier “no substrate capability gap” claim. That leaves the ledger self-contradictory, which violates INVARIANTS.md P1’s “Documentation Describes Live State” standard.
  • BLOCKING — ROADMAP.md: the new status text says phase 1 already landed with CharClass/char_in_class in dsl/std/unicode.dag and that the generated tokenizer no longer carries is_ascii_* calls, but this PR head still lacks those symbols and src/v3/compiler/src/tokenize_generated.rs still uses is_ascii_whitespace, is_ascii_digit, is_ascii_alphabetic, and is_ascii_alphanumeric. Same INVARIANTS.md live-state rule.

Verdict
REQUEST_CHANGES

The new ledger text does not match the checked-in tree yet, so I would not approve it as-is. Either soften the row to describe the still-pending state, or land the cited CharClass/tokenizer changes before merging this documentation update.

…-1 status edits + char_in_class interpreter-parity sibling row from main
@briansrls

Copy link
Copy Markdown
Contributor Author

Good catch by Codex — the inline review was accurate at the reviewed commit (cf164c72f, pre-merge). At that snapshot, PR #693 hadn't yet landed on main, so the ROADMAP row's "phase-1 landed" language was ahead of the artifacts.

After the origin/main merge (commit 6dcceb8a6) pulled PR #693 in, all three artifacts are now present on the branch tip:

  • dsl/std/unicode.dag carries CharClass + char_in_class (line 22+).
  • src/v3/compiler/src/tokenize_generated.rs: zero is_ascii_* calls.
  • src/v3/compiler/src/tokenize_char_class.rs on disk as the host-stdlib bridge.

So the review's concern is resolved by the rebase. The ROADMAP row's "landed via PR #693" claim is now truthful against main (once #706 lands).

Thanks to Codex for flagging — this is exactly the ledger-vs-code drift detection that protects follow-up planning.

@briansrls briansrls mentioned this pull request Apr 24, 2026
…-5.4 review)

Row title still said 'consumption gap, not substrate gap' while the
body block retracted that claim and cited Class 5 Gap 3 as a substrate
dependency for phase-2. Title now matches body: mixed classification,
consumption for steps 1+3, substrate for step 2.
@briansrls

Copy link
Copy Markdown
Contributor Author

Thanks — two findings from gpt-5.4, responding to each:

Finding #1 (self-contradictory title vs. body) — VALID, fixed in 5a9d79710-successor.

You caught a real persistent inconsistency that the Claude / Codex-default reviews missed. My edits retracted the "no substrate capability gap" claim in the Status block but left the original 2026-04-23 title parenthetical "(consumption gap, not substrate gap)" intact. That left the row internally contradictory.

Fixed by rewording the title to "(mixed: consumption gap for steps 1+3, substrate gap for step 2 — see Status block below)". The dissolution enumeration in the body (steps 1+2+3) now matches the title classification: steps 1 and 3 are consumption work (vocabulary + scanner wiring), step 2 is where the substrate gap lives. No more internal contradiction.

Finding #2 (phase-1 artifacts missing) — snapshot-specific, already resolved by merge.

You reviewed commit cf164c72f, which was pre-merge. At that snapshot, PR #693 hadn't landed on main yet, so the artifacts the ROADMAP row cited weren't on the branch head. After my origin/main merge (commit 6dcceb8a6) pulled #693 in, all three are now present on session/zesty-bear-812:

  • dsl/std/unicode.dag → CharClass + char_in_class (line 22+).
  • src/v3/compiler/src/tokenize_generated.rs → 0 is_ascii_* calls (grep confirms).
  • src/v3/compiler/src/tokenize_char_class.rs → on disk.

So the "claim ahead of code" issue is resolved by the merge itself.

Verdict update: both P1 live-state violations addressed. Mergeable as of the title-fix commit.

briansrls added a commit that referenced this pull request Apr 24, 2026
…sync)

Director citation chain: character-level consumption-gap pointer moves after
PR #706 inserts Class 5 Gap 3 row above it.

Made-with: Cursor
@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: claude / claude-opus-4-7
  • Commit: 559773ca · Trigger: schedule
  • Thinking: 16s wall

Documentation-only diff. The change adds a Class 5 Gap 3 ledger entry and amends the Character-level row to reflect retraction/escalation from PR #693. All references (line numbers, file paths, owners) are internally consistent.

Verdict: APPROVE — Pure ROADMAP.md update. The new Class 5 Gap 3 entry has all three properties for a tracked bridge (documented, bounded, named dissolution trigger: "extend compile_to_dag to parse port-carried field values in data bodies"). The Character-level row honestly retracts the prior "no substrate capability gap" claim and points at the new ledger row. No code under review; no invariants touched.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: codex / gpt-5.4
  • Commit: 559773ca · Trigger: schedule
  • Thinking: 70s wall

Verdict: APPROVE. This diff is a narrow ROADMAP.md correction, and I did not find any rubric violations in the changed lines. The new text makes the ledger more honest about live state by retracting the earlier “no substrate capability gap” claim and tying the remaining block to the new Class 5 Gap 3 row rather than leaving a stale contradiction in place.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: claude / claude-opus-4-7
  • Commit: 71f46afe · Trigger: schedule
  • Thinking: 17s wall

Verdict: APPROVE — pure ROADMAP bookkeeping. The diff adds a Class 5 Gap 3 ledger row with a clear dissolution trigger ("extend compile_to_dag to parse port-carried field values in data bodies") and amends the Character-level row to retract the earlier "no substrate capability gap" claim, splitting it into phase-1 (landed) and phase-2 (blocked on Gap 3). Debt is documented, bounded, and owner/trigger-named — tracked-bridge discipline is satisfied. No code under review.

Minor exploratory observation (non-ask): the new Gap 3 row's "Audit pattern" guidance is useful but buried in a long paragraph; if a future pass wants to make it actionable it could be hoisted to a short standalone process-note. Not worth blocking on.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: codex / gpt-5.4
  • Commit: 71f46afe · Trigger: schedule
  • Thinking: 179s wall

Findings

  • ROADMAP.md:314 violates INVARIANTS.md’s “Documentation describes live state” rule. The new Class 5 Gap 3 row says the remaining gap is “port-carried field values” and that compile_to_dag still rejects nested records, list literals, declaration refs / Var refs, and sum-variant literals inside data bodies, but the live lowerer already accepts those shapes for record-bodied data via lower_structural_field_value (Reference, List, nested Record, and Variant are all lowered today) and FieldValue explicitly models them. The still-open limitation is the top-level ValueBody boundary (Scalar or record-Structural only), so this ledger entry would codify the wrong capability boundary.

Verdict
REQUEST_CHANGES. The diff is small, but this row records the wrong live-state claim about what data-body shapes are still unsupported, and that makes the roadmap less trustworthy rather than more.

…lass phase-2 blocker classification (per gpt-5.4 audit)

gpt-5.4's review on 706 @ 71f46af caught that the row's "remaining
gap" description was wrong: field-level shapes (nested records, list
literals, declaration refs, Var refs, sum-variant literals) are
supported today via FieldValue variants + lower_structural_field_value
(dag.rs:328-353, lower.rs:2616+). The actual remaining gap is the
top-level ValueBody boundary (non-scalar, non-record top-level bodies).

The authority I cited — DOWNSTREAM_REQUIREMENTS.md:239 — is itself
stale: it describes the pre-PR-B-unwind shape where FieldValue was
LiteralBits-only. PR-B's unwind extended FieldValue to carry
Reference / Record / List / Variant, moving the gap to ValueBody.

Two fixes:

1. Rewrite the Class 5 Gap 3 row to describe the actual ValueBody
   boundary, point at code paths (dag.rs, lower.rs) as live authority,
   flag DOWNSTREAM entry as itself stale, and soften phase-2 CharClass
   blocker classification to "provisional pending reproduction."

2. Update the Character-level row's phase-2 block to name that the
   specific shape of the CharClass failure needs concrete reproduction
   from the escalating sub-child before the blocker is finalized.

Recursive audit-pattern instance: the row I wrote to codify "verify
live state before claiming substrate gap" itself failed to verify live
state. Both incidents (2026-04-23 original row + 2026-04-24 my
retraction row) are now cited in the audit-pattern sub-note as
examples of the same discipline.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor Author

gpt-5.4 caught a substantive error — the row's description of the remaining gap was wrong, and the authority I cited (DOWNSTREAM_REQUIREMENTS.md:239) is itself stale against the live code. Fixed in the next commit.

What the code actually says:

  • ValueBody (src/v3/compiler/src/dag.rs:259-287) has three variants: Unparsed(SourceSpan), Structural { fields: Vec<(String, FieldValue)> }, Scalar(LiteralBits).
  • FieldValue (src/v3/compiler/src/dag.rs:328-353) has five variants: Literal, Reference(DeclarationId), Record(Vec<(String, FieldValue)>), List(Vec<FieldValue>), Variant { constructor, payload }.
  • lower_structural_field_value (src/v3/compiler/src/lower.rs:2616+) handles all five — nested records via recursion on line 2694+, lists via 2664-2691, declaration refs via the DeclarationRef marker / symbol lookup, literals via lower_scalar_literal_for_type, variants as part of the structural sum path.

So field-level shapes inside Structural are full support today. The actual remaining gap is the top-level ValueBody boundary — data foo: List<T> = [...], data foo: SumType = V, data foo: T = other_decl_ref can't parse as top-level bodies (fall to ValueBody::Unparsed). dag.rs:269-287 is explicit that Scalar is deliberately NOT widened to swallow these shapes; extension path is new ValueBody variants.

DOWNSTREAM_REQUIREMENTS.md:239 is itself stale. It describes the pre-PR-B-unwind shape where FieldValue was LiteralBits-only and port-carried field values WERE the gap. PR-B's unwind added the four additional variants and moved the gap to ValueBody. The DOWNSTREAM entry wasn't updated to match.

Recursive audit-pattern instance. The Class 5 Gap 3 row I wrote to codify "verify live state before claiming substrate gap" itself failed that check — I cited the DOWNSTREAM summary without cross-checking dag.rs / lower.rs live code. That's a second instance of the same discipline failure the row is trying to name, in the very PR that names it. The updated row now cites both incidents (the 2026-04-23 original + my 2026-04-24 retraction) as recursive examples, and strengthens the rule: reading a summary authority is necessary but not sufficient; live code paths must be checked directly because authority text itself drifts.

For the CharClass phase-2 blocker: I've softened the classification from "blocked on Class 5 Gap 3" to "sub-child classified as Class 5 Gap 3; specific shape of failure needs concrete reproduction before finalizing." The field-level shapes they attempted might actually work today; the top-level ValueBody boundary might be what they hit; or something narrower still. PR #693's author needs to post the concrete failing data declaration so we can land the right dissolution direction.

Thanks to gpt-5.4 for the live-state audit. This is the exact drift-detection pattern the ROADMAP row is trying to institutionalize, demonstrating its own necessity by catching me authoring ahead of verification.

Verdict request: please re-review the corrected commit when it lands.

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review metadata

  • Provider / model: codex / unknown
  • Commit: 71f46afe · Trigger: schedule
  • Thinking: 307s wall

Non-blocking — Improvements (fix in-PR if easy, else defer to roadmap)

  • ROADMAP.md Line 314 correctly identifies the live ValueBody boundary, but it also says src/v3/DOWNSTREAM_REQUIREMENTS.md:239 is stale while leaving that active tracker unchanged; T-Receipts should update or demote that entry so the same gap has one live authority.

ROADMAP — Verified

  • Class 5 Gap 3 boundary: The new row matches live dag.rs and lower.rs: FieldValue supports structural field values, while lower_data_item only accepts top-level record and scalar data bodies today.
  • Character-level row retraction: The status block retracts the prior no-substrate-gap claim and records phase-1 as a bounded mirror with interpreter parity still tracked.

✅ Docs-only PR; no blocking concerns after the Class 5 Gap 3 correction.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: claude / claude-opus-4-7
  • Commit: b4ae1784 · Trigger: schedule
  • Thinking: 24s wall

APPROVE — Pure ROADMAP.md ledger update: adds the Class 5 Gap 3 row (top-level ValueBody boundary) and updates the 2026-04-23 Character-level row with a post-#693 Status block retracting the earlier "no substrate capability gap" claim. Cross-references between the two rows are consistent, and the self-audit note about authority-text drift (DOWNSTREAM_REQUIREMENTS.md:239 being stale vs. live dag.rs/lower.rs) is the right discipline signal. No code under review, nothing in the diff violates INVARIANTS / modeling discipline / CODING / TESTING.

Exploratory observation (non-ask): the new row is honest about the provisional classification ("needs concrete reproduction from the escalating sub-child before the dissolution direction is finalized"), which is good — but it also commits to a dissolution direction ("extend ValueBody with the additional top-level variants") in the same sentence. Worth watching that the reproduction, when it lands, is actually allowed to redirect the dissolution (e.g., toward the "extractor-over-unparsed-source" alternative already mentioned) rather than being back-fit to the already-stated direction.

@briansrls

Copy link
Copy Markdown
Contributor Author

Review metadata

  • Provider / model: codex / gpt-5.4
  • Commit: b4ae1784 · Trigger: schedule
  • Thinking: 85s wall

Verdict: APPROVE

No findings. The diff is documentation-only, and the revised ROADMAP.md claims line up with the live code I checked in src/v3/compiler/src/dag.rs, src/v3/compiler/src/lower.rs, dsl/std/unicode.dag, and src/v3/compiler/src/tokenize_char_class.rs: field-level FieldValue support is already present, the remaining data-body limitation is at the top-level ValueBody boundary, and the character-class status block now accurately reflects phase 1 vs. the still-unverified phase-2 blocker shape. Residual risk is the one the row already names: the exact PR #693 phase-2 failure still needs concrete reproduction before that blocker classification becomes authoritative.

@briansrls
briansrls merged commit ce9e930 into main Apr 24, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant