Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
323fc38
WIP: deep-moth-845
briansrls Apr 27, 2026
26fd26c
WIP: deep-moth-845
briansrls Apr 27, 2026
f691d32
Merge remote-tracking branch 'origin/main' into session/deep-moth-845
briansrls Apr 27, 2026
a7ace91
WIP: deep-moth-845
briansrls Apr 27, 2026
015383f
chore: apply cargo fmt
briansrls Apr 27, 2026
b4e0cf1
Fix Option matching and unused mut in v3 regen token generation
briansrls Apr 27, 2026
5ed943d
Fix tokenize literal generation and parse snapshot manifest baseline
briansrls Apr 27, 2026
f421d7a
Add tokenizer ASCII parity regression test
briansrls Apr 27, 2026
ed23214
Restore tokenizer ASCII parity test and refresh parse manifest tuple
briansrls Apr 27, 2026
1fb1dfe
test(v3): expose tokenize parity helpers for ascii parity tests
briansrls Apr 27, 2026
ad241fe
regen(v3): emit tokenizer ASCII predicate helpers as pub(crate)
briansrls Apr 27, 2026
7e003a5
fmt(v3): wrap generated byte_matches signature emission
briansrls Apr 27, 2026
5a2d296
regen(v3): make tokenizer precedence follow ascii_scan_order
briansrls Apr 27, 2026
c1a50d3
regen(v3): drive tokenizer class dispatch from ascii_scan_order
briansrls Apr 27, 2026
fe17aaa
chore(v3): sync checked-in tokenize_generated with current regen output
briansrls Apr 27, 2026
f473403
Fix tokenize_generated.rs snapshot formatting drift
briansrls Apr 27, 2026
7fdff8b
Update tokenize_generated.rs snapshot spacing after diagnostics import
briansrls Apr 27, 2026
83df878
chore(v3): avoid manual byte range checks in generated predicate
briansrls Apr 27, 2026
401e989
chore(v3): emit ascii checks using is_ascii_* helpers
briansrls Apr 27, 2026
fd9cb71
fix(v3): sync tokenizer scanner snapshot formatting
briansrls Apr 27, 2026
1d2241b
Merge remote-tracking branch 'origin/main' into fix-pr-1002-charclass
briansrls Apr 27, 2026
a6ec96a
test(v3): refresh tokenizer parse manifest hash
briansrls Apr 27, 2026
b7d7916
Merge remote-tracking branch 'origin/main' into fix-pr-1002-charclass
briansrls Apr 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions INVARIANTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -150,6 +150,13 @@ A downstream stage reads a lower layer not through its declared accessors but by
- **E-9: External Realization Lives On Arrow.body** — declares the single authority for external semantics
- **DB-14: Substrate External Primitives Materialize Through Declared Arrow.body Plus Target Bindings** — same single-authority principle for external primitives

### Reflection evidence is not structural proof

`reflect_program_dag_nodes_in_file` in `lens_apply.rs` currently emits a **shallow behavior spine** (`result_port` plus limited tags) and intentionally drops structural fields (`target`, `inputs`, `input`, `paths`, `source`, `init`, `body`, `bound`, etc.). Consumers that rely on this view — in practice `LensOutputEquals` and `AlgebraicLaw` runners in `test_runner.rs` — therefore only gain **regression evidence** from the current reflection path, not structural self-inspection parity with
`src/v3/std/substrate.dag`.

**Confidence tag for these gates:** treat these paths as `StructuralEvidence::Shallow` until full substrate reflection is landed. PRs using them should state that status explicitly in their acceptance notes and route dissolution via the dedicated lossy-reflection closure row (`ROADMAP.md`), not by claiming by-construction proof.

---

## P3: Fail-Closed
Expand Down
8 changes: 8 additions & 0 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -475,6 +475,14 @@ Three distinct reflective/exploratory analyses ran against `main@7f74f09` and `m

- **Filename / sentinel bridges in `test_runner.rs`** (NOVEL — class-of-pattern row): `src/v3/compiler/src/test_runner.rs` carries multiple file-path / fixture-name / sentinel bridges as identity references: fixture filename → bind name mapping; `PROGRAM_INPUT_SENTINEL = "r1_lens_output_input_from_program"`; `include_str!("../lenses/named_function_count.dag")` as canonical lens side-channel; `EXTDEPS_BOOTSTRAP_FIXTURES` manually listing currently-loadable extdeps surface; `fold_site_skips_d1_monomorph_list_fold_path` using a path suffix. Each instance is locally explainable; together they show **the same missing substrate fact: identity and role references not yet carried structurally enough for consumers**. **Dissolution trigger**: replace fixture-filename routing + `PROGRAM_INPUT_SENTINEL` with structural references (`DeclarationRef`; explicit input value carriers; bind/output identity in the claim itself; resource/materialization records). Surfaced by Reflective Analysis 2026-04-25 (Pattern 3 + Risk 4) + Exploratory 2026-04-25. P2 Boundary Discipline (string/name/path as temporary identity). Owner: unassigned; M scope.

- **B4 bridge-retirement queue: file/name bridge class (sequenced)** (NOVEL + consolidation): retire file/name bridge sites in a single identity-carrier queue under `docs/briefs/b4-identity-carrier-substrate-pass.md`:
1. **`SourceSpan.file` participation checks** (`lens_apply.rs:340-344`, `lower.rs:848`, `lower.rs:1637-1738`, `lower.rs:1954+`, `emit.rs:1082-1094`, `dag.rs:2914`): dissolve via structural fold-shape / emit-helper carrier landing and migration.
2. **Canonical lens-name dispatch / fixture-name routing** (`test_runner.rs`): consume `DeclarationRef` and claim-role carriers for lens/input identity; delete `PROGRAM_INPUT_SENTINEL` and fixture-name bridges.
3. **`include_str!` side channels** (`test_runner.rs:23`, `:33`): dissolve in favor of canonical declaration references once `apply_lens_declaration` uses same-id carriers.
4. **Exact-string patching in generated helper code** (`lib.rs::patch_lower_helpers_generated_type_alias_refinement`): remove once generated helper output carries refinement fields natively and bridge-patch is no longer needed.

**Sequencing:** this queue is an owner-scoped B4 sub-brief slice; each item follows `B4.1` (`DeclarationRef`) then the relevant carrier-landing rows in B4 order, and records closure only when the previous row proves its successor consumer is structurally routed. This is not a new program, just a single sequence inside the existing B4 identity-carrier lane. Surfaced by 5c bridge sweep request. P2 Boundary Discipline + P5 Progress Is Dissolution. Owner: unassigned; M scope.

### Reviewer-noise class — a practice, not a debt

- **Integration-reflection cadence**: every ~few days, run a reflective + exploratory analysis pair. The 2026-04-15 and 2026-04-18 passes caught items individual PR reviews missed (authority split across PRs, silent cross-PR name-based lookups, the `CreateComment` drift, the `repeat_string` bug). Worth institutionalizing. **Velocity-tripwire reporting (added 2026-04-25 per [PR #810](https://github.com/gunb-ai/gunbc/pull/810) §4)**: each cadence pass also reports introduction:dissolution PR ratio for the window. Heuristic match: PR titles containing `retire`/`dissolve`/`delete`/`remove`/`collapse` count as dissolution; substantial-scaffold-introducing PRs (new hand-Rust file, new sentinel, new Rust mirror of a `.dag` authority) count as introduction. **Calibration caveat**: this is a lower-bound heuristic — feature PRs that ship dissolution work in the same diff (e.g., a T-LensAPI PR adding a lens and retiring three legacy paths) are undercounted. Before a ≥3:1 reading triggers Director review (per INVARIANTS.md §P5 "Dispatch-Discipline Mechanisms" (c)), the cadence pass MUST do a manual sweep for dissolution-bearing feature PRs in the window to avoid false positives. Current 2026-04-18→2026-04-25 baseline: ~1.6:1 (8 dissolution-shaped + 13 substantial introductions; lower bound).
Expand Down
252 changes: 202 additions & 50 deletions src/v3/compiler/src/bin/regen_tokenize.rs
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,7 @@ fn read_shared_syntax_source(manifest_dir: &std::path::Path) -> String {

fn generate(dag: &Dag, shared_syntax: &SharedSyntaxAuthority) -> String {
let keywords = collect_keyword_rows(dag, shared_syntax);
let ascii_scan_order = collect_ascii_scan_order(dag);
let puncts = collect_punct_rows(dag, shared_syntax);
let line_comment_prefix = string_data_named(dag, "line_comment_prefix");
let string_delim = string_data_named(dag, "string_literal_delimiter");
Expand All @@ -125,7 +126,7 @@ fn generate(dag: &Dag, shared_syntax: &SharedSyntaxAuthority) -> String {

let mut out = String::new();
out.push_str("use crate::diagnostics::{Diagnostic, SourceSpan};\n");
out.push_str("use crate::tokenize_char_class::{byte_matches, TokenizerCharClass};\n\n");
out.push_str(&emit_char_scanner_class_scaffolding(&ascii_scan_order));
out.push_str(&emit_token_kind_enum(dag));
out.push_str(
r#"#[derive(Debug, Clone)]
Expand All @@ -145,11 +146,120 @@ pub struct Token {
&diag_int_pre,
&diag_int_suf,
&escapes,
&ascii_scan_order,
));
out.push_str(&emit_punctuation_token(&puncts));
out
}

fn emit_char_scanner_class_scaffolding(scan_order: &[String]) -> String {
assert!(
scan_order.len() == 4,
"`ascii_scan_order` in `tokenize.dag` must list exactly 4 class names"
);
assert!(
scan_order
.iter()
.collect::<std::collections::BTreeSet<_>>()
.len()
== scan_order.len(),
"`ascii_scan_order` in `tokenize.dag` must not contain duplicates"
);

let mut out = String::new();
out.push_str("\n#[derive(Debug, Clone, Copy, PartialEq, Eq)]\n");
out.push_str("pub(crate) enum ScannerCharClass {\n");
for class in scan_order {
out.push_str(&format!(" {class},\n"));
}
out.push_str("}\n\n");

out.push_str("#[inline]\n");
out.push_str("pub(crate) fn byte_matches(byte: u8, class: ScannerCharClass) -> bool {\n");
out.push_str(" match class {\n");
for class in scan_order {
let expr = ascii_scan_class_predicate(class);
out.push_str(&format!(" ScannerCharClass::{class} => {expr},\n"));
}
out.push_str(" }\n}\n\n");

out
}

fn ascii_scan_class_predicate(class_name: &str) -> &'static str {
// Interim bridge: `ascii_scan_order` supplies structural scanner order, but
// class predicate bodies remain here until `std.unicode::char_in_class` is
// structurally consumable by the tokenizer generator.
match class_name {
"Whitespace" => "matches!(byte, b'\\t' | b'\\n' | b'\\x0c' | b'\\r' | b' ')",
"Digit" => "byte.is_ascii_digit()",
"IdentStart" => "byte.is_ascii_lowercase() || byte.is_ascii_uppercase() || byte == 0x5f",
"IdentContinue" => "byte.is_ascii_alphanumeric() || byte == 0x5f",
_ => panic!("unsupported scanner class `{class_name}` in `ascii_scan_order`"),
}
}

fn collect_ascii_scan_order(dag: &Dag) -> Vec<String> {
let scan_decl = dag
.declarations()
.iter()
.find(|d| d.name.as_deref() == Some("ascii_scan_order"))
.unwrap_or_else(|| panic!("missing `ascii_scan_order` in `{TOKENIZE_AUTHORITY_FILE}`"));

let char_class_decl = dag
.declarations()
.iter()
.find(|d| d.name.as_deref() == Some("CharClass"))
.unwrap_or_else(|| panic!("missing `CharClass` in `{TOKENIZE_AUTHORITY_FILE}`"));

let TypeConnective::Disj {
variants: char_class_variants,
} = &char_class_decl.connective
else {
panic!("`CharClass` should be a disj declaration");
};

let Some(ValueBody::List(values)) = scan_decl.value_body.as_ref() else {
panic!("`ascii_scan_order` in `{TOKENIZE_AUTHORITY_FILE}` must be a list");
};

let mut out = Vec::new();
for value in values {
let FieldValue::Variant {
constructor,
payload,
} = value
else {
panic!("`ascii_scan_order` elements must be constructor values");
};
assert!(
payload.is_empty(),
"`ascii_scan_order` class entries must be nullary constructors"
);
let label = char_class_variants
.iter()
.find(|field| field.ty == *constructor)
.map(|field| field.label.clone())
.unwrap_or_else(|| {
panic!(
"`ascii_scan_order` contains constructor {:?} not owned by `CharClass`",
constructor
)
});
out.push(label);
}

let expected = ["Whitespace", "Digit", "IdentStart", "IdentContinue"];
for class in &expected {
assert!(
out.contains(&class.to_string()),
"`ascii_scan_order` in `{TOKENIZE_AUTHORITY_FILE}` must include `{class}`"
);
}

out
}

fn assert_shared_syntax_raw_source_scaffold_still_required(shared_syntax_source: &str) {
let lowered = match compile_to_dag(shared_syntax_source, SHARED_SYNTAX_FILE) {
Ok(dag) => dag,
Expand Down Expand Up @@ -671,7 +781,18 @@ fn emit_tokenize_fn(
diag_int_pre: &str,
diag_int_suf: &str,
escapes: &[(u8, i64)],
ascii_scan_order: &[String],
) -> String {
let ensure_classes = |required: &[&str]| {
for name in required {
assert!(
ascii_scan_order.iter().any(|class| class == name),
"`ascii_scan_order` in tokenize.dag missing `{name}`"
);
}
};
ensure_classes(&["Whitespace", "Digit", "IdentStart", "IdentContinue"]);

let mut arms = String::new();
for (spelling, kind) in keywords {
arms.push_str(&format!(
Expand Down Expand Up @@ -705,10 +826,86 @@ fn emit_tokenize_fn(
));
s.push_str(" while pos < bytes.len() {\n");
s.push_str(" let byte = bytes[pos];\n\n");
s.push_str(" if byte_matches(byte, TokenizerCharClass::Whitespace) {\n");
s.push_str(" pos += 1;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
s.push_str(" let start = pos;\n\n");
for class in ascii_scan_order {
if class == "Whitespace" {
s.push_str(" if byte_matches(byte, ScannerCharClass::Whitespace) {\n");
s.push_str(" pos += 1;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
} else if class == "Digit" {
s.push_str(" if byte_matches(byte, ScannerCharClass::Digit) {\n");
s.push_str(" let mut end = pos;\n");
s.push_str(
" while end < bytes.len() && byte_matches(bytes[end], ScannerCharClass::Digit) {\n",
);
s.push_str(" end += 1;\n");
s.push_str(" }\n");
s.push_str(" let literal = &source[start..end];\n");
s.push_str(
" let value: i64 = literal.parse().map_err(|_| Diagnostic::TokenizerError {\n",
);
s.push_str(&format!(
" message: format!(\"{{}}{{}}{{}}\", {}, literal, {}),\n",
int_pre, int_suf
));
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" fixes: Vec::new(),\n");
s.push_str(" })?;\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind: TokenKind::IntLit(value),\n");
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" });\n");
s.push_str(" pos = end;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
} else if class == "IdentStart" {
s.push_str(" if byte_matches(byte, ScannerCharClass::IdentStart) {\n");
s.push_str(" let mut end = pos;\n");
s.push_str(
" while end < bytes.len() && byte_matches(bytes[end], ScannerCharClass::IdentContinue) {\n",
);
s.push_str(" end += 1;\n");
s.push_str(" }\n");
s.push_str(" let text = &source[start..end];\n");
s.push_str(" let kind = match text {\n");
s.push_str(&arms);
s.push_str(" _ => TokenKind::Ident(text.to_string()),\n");
s.push_str(" };\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind,\n");
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" });\n");
s.push_str(" pos = end;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
} else if class == "IdentContinue" {
s.push_str(" if byte_matches(byte, ScannerCharClass::IdentContinue) {\n");
s.push_str(" let mut end = pos;\n");
s.push_str(
" while end < bytes.len() && byte_matches(bytes[end], ScannerCharClass::IdentContinue) {\n",
);
s.push_str(" end += 1;\n");
s.push_str(" }\n");
s.push_str(" let text = &source[start..end];\n");
s.push_str(" let kind = match text {\n");
s.push_str(&arms);
s.push_str(" _ => TokenKind::Ident(text.to_string()),\n");
s.push_str(" };\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind,\n");
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" });\n");
s.push_str(" pos = end;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
} else {
panic!(
"unsupported scanner class `{}` in `ascii_scan_order`",
class
);
}
}
s.push_str(" // Line comment prefix from `tokenize.dag` (`line_comment_prefix`).\n");
s.push_str(" if bytes.len() >= pos + LINE_COMMENT_PREFIX.len()\n");
s.push_str(
Expand All @@ -721,7 +918,6 @@ fn emit_tokenize_fn(
s.push_str(" }\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
s.push_str(" let start = pos;\n\n");
s.push_str(" if let Some((kind, width)) = punctuation_token(bytes, pos) {\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind,\n");
Expand All @@ -732,50 +928,6 @@ fn emit_tokenize_fn(
s.push_str(" pos += width;\n");

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BLOCKING: Per modeling-discipline facts-flow-forward (cross-stage authority), this drops the structural authority it just read: tokenize.dag now defines ascii_scan_order as precedence, but emit_tokenize_fn still emits a fixed scanner branch order and does not use that order from the model.

s.push_str(" continue;\n");
s.push_str(" }\n\n");
s.push_str(" if byte_matches(byte, TokenizerCharClass::Digit) {\n");
s.push_str(" let mut end = pos;\n");
s.push_str(
" while end < bytes.len() && byte_matches(bytes[end], TokenizerCharClass::Digit) {\n",
);
s.push_str(" end += 1;\n");
s.push_str(" }\n");
s.push_str(" let literal = &source[start..end];\n");
s.push_str(
" let value: i64 = literal.parse().map_err(|_| Diagnostic::TokenizerError {\n",
);
s.push_str(&format!(
" message: format!(\"{{}}{{}}{{}}\", {}, literal, {}),\n",
int_pre, int_suf
));
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" fixes: Vec::new(),\n");
s.push_str(" })?;\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind: TokenKind::IntLit(value),\n");
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" });\n");
s.push_str(" pos = end;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
s.push_str(" if byte_matches(byte, TokenizerCharClass::IdentStart) {\n");
s.push_str(" let mut end = pos;\n");
s.push_str(
" while end < bytes.len() && byte_matches(bytes[end], TokenizerCharClass::IdentContinue) {\n",
);
s.push_str(" end += 1;\n");
s.push_str(" }\n");
s.push_str(" let text = &source[start..end];\n");
s.push_str(" let kind = match text {\n");
s.push_str(&arms);
s.push_str(" _ => TokenKind::Ident(text.to_string()),\n");
s.push_str(" };\n");
s.push_str(" tokens.push(Token {\n");
s.push_str(" kind,\n");
s.push_str(" span: SourceSpan::new(file, start as u32, end as u32),\n");
s.push_str(" });\n");
s.push_str(" pos = end;\n");
s.push_str(" continue;\n");
s.push_str(" }\n\n");
s.push_str(
" // String literal (`string_literal_delimiter` + `StringEscapeSpec` rows).\n",
);
Expand Down
Loading
Loading