Repository navigation
V2 compiler convergence - #170
Merged
Merged
Conversation
Replace dummy leaf_node(name: "Unit") fallbacks with available contextual information at 12 error sites, enforcing the "no fallbacks that fabricate" invariant: CRITICAL: - Wildcard catch-all in infer_expr now emits an explicit "unhandled expression variant" diagnostic instead of silently returning Unit with fake span HIGH: - Undefined variable: use leaf_node(name: name) instead of Unit - Missing field access: use leaf_node(name: field_name) - Function not found: use leaf_node(name: func_name) + new diagnostic - Sig lookup inner None: use leaf_node(name: func_name) - Empty match arms: use scrutinee type + new "no arms" diagnostic - If-branch type mismatch: use then-branch type (error already emitted) - Lambda body: use body_result.typed.resolved_type - ForEach body: use body_result.typed.resolved_type MEDIUM: - build_func_env missing return_type: use item.name - for_each_element_type_node fallback: propagate normed node Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The parser now builds Nodes directly. TypeExpr was a parallel
representation that existed only as an intermediate step before
conversion to Node. This removes the duplicate representation:
- TypeExpr type definition (8 variants) and ContainerKind enum
- type_expr_to_node and 7 helper functions (field_to_node,
variant_to_node, type_expr_named_to_node, etc.)
- Predicate type (7 variants) and predicate_to_field_init
- empty_node (defined but never called)
The parser's predicate handling now builds FieldInit values directly
instead of constructing intermediate Predicate values. This eliminates
the lossy conversion where Range{min,max} bounds were discarded.
Removes ~200 lines from 00_core.dag.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…text Parser error-recovery dummy values were discarding available names and spans. When parsing fails partway through a definition, the name and start_span are often already extracted -- the dummy nodes should carry that information rather than fabricating empty strings and zero spans. Changes: - parse_item wildcard: use current_span and "<unknown>" instead of zero span/empty name - parse_type_expr: use current_span for inline record and wildcard error paths - finish_type_expr_from_name: use type_name and start_span in dummy_te - parse_single_named_int: preserve arg_name in all error paths after extraction - parse_type_def, parse_type_body_after_eq, parse_fn_def, parse_func_def, parse_service_def, parse_resource_def, parse_data_def, parse_extern_decl: introduce named_dummy after name extraction so later error returns carry the actual name Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…t maps - Kahn topological sort: partition edges each round so processed edges are removed. Total work across all rounds is O(E) instead of O(N×E). - Pipeline gate: typecheck_module halts before inference if build_type_env produces Error-severity diagnostics. Unresolved types never reach inference or emit. - Language extdeps: add emit_type_map, emit_keywords, and emit_container_templates to Rust/Python/Go extdeps for future C3 wiring (emitters importing from extdeps instead of inline data). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When matching on a known coproduct (Disj), verify that all variants are covered by the arms. Wildcard and Bind patterns count as covering all remaining variants. Non-exhaustive matches emit a Warning diagnostic. This enables the compiler to detect missing match arms for enum-like types at inference time rather than silently falling through. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
All 38 list accumulator sites in 02_parse.dag used concat([x], acc) which copies the entire list on each call, making list building O(n²). Replaced each with list_push(acc, x) which is O(1) amortized append. Since list_push appends (vs concat's prepend), the corresponding reverse() calls at each accumulator's termination point are no longer needed and have been removed — list_push already builds in forward order. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…t.dag Three emitter backends (Rust, Python, Go) each duplicated identical classification logic for determining item kind, type structure, and transport kind. This violates the sustainability invariant "No duplicate representations" — changing the classification rules required editing three files in lockstep. Extracted three shared classifiers into v2.compiler.emit: - classify_typed_item: item -> type_def|type_alias|function|data_def| service_def|resource_def|extern_func|unhandled - classify_type_structure: node -> leaf|conj|disj - classify_transport_kind: transport -> rest|shell|file|local Each backend now calls the shared classifier and dispatches on the returned string tag to its own per-language rendering functions. The removed per-backend dispatch functions (emit_*_node_type_dispatch) are replaced by inline if/else on the classifier result. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The function converted TypedExpr back to plain Expr. It was imported by 04_infer.dag but never called. All emit hot paths now operate directly on TypedExpr, making this bridge function dead code. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
P0 fix: generated function closures have enormous stack frames in debug mode (infer_expr: 356KB, resolve_node: 148KB) because Node is 544 bytes and TypedExpr is 1,112 bytes. The previous 32KB threshold was 11x too small for infer_expr, causing stack overflows in generated integration tests. Root cause: debug-mode Rust allocates stack space for ALL match arm locals simultaneously, and the codegen emits pervasive value cloning (n.clone().name.clone() creates 544-byte temporaries). Short-term fix: threshold 32KB → 512KB (exceeds largest frame). Medium-term: box Node.transport/config to shrink Node from 544→~120B. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Introduces src/v2/08_artifact.dag with ArtifactKind, Artifact, BoundaryKind, Boundary, ArtifactPlan types and a default_artifact_plan function that wraps the current single-artifact assumption as an explicit data structure. No pipeline wiring — this is the compatibility layer for future multi-artifact support (E1/E2). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add src/v2/07_complexity.dag with the symbolic type definitions for runtime complexity analysis: SizeExpr, CostExpr, Constraint, Certainty, SemanticsCtx, and ComplexitySummary. These are type definitions only -- no analysis implementation yet (that is D2-D4). Costs are parameterized by SemanticsCtx so the same surface operation can have different faithful costs under different backends and data structure models (e.g. PersistentList vs RcVec). Also registers the new file in the v2 test harness (compile_all_modules, detect_duplicate_fn_names, and phase0 parse audit). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The inline data declarations in 05_emit.dag duplicate the canonical
definitions in dsl/extdeps/languages/{rust,python,go}/types.dag.
Replacing them with imports is blocked by four v1 bootstrap limitations:
name collision in build_data_values, no qualified data references in
eval_ident, no import renaming syntax, and data declarations not being
import-bindable. Updated the section comment to document these blockers
and point to the canonical source.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New track for systematically auditing and fixing type representation sizes. Discovered during 2026-03-18 session: Node is 544 bytes because it inlines TransportBinding (184b, 6% usage) and ServiceConfig (256b, 6% usage), making TypedExpr 1,112 bytes and causing generated tests to OOM. R1: Audit + size report with regression assertions R2: Box rare fields on Node (544b → ~120b) — highest impact R3: Clone reduction in v1 emitter (650+ sites) R4: Interpreter value representation (list_push quadratic) Track A (self-hosting) is now explicitly blocked on R2. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The Rust emitter's TCO pattern clones Rc-wrapped state at each loop iteration, bumping refcount so Rc::try_unwrap always falls back to full Vec clone. This makes list_push O(n²) even in compiled native code. Discovered by Track A investigation: tokenize_loop grows at 300MB/s and OOMs at ~10GB on 1,515 lines of gist sources. Fix: emit move instead of clone for TCO state consumption. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The TCO loop pattern was cloning the loop variable at each iteration start (let state = __tco_p_state.clone()), bumping Rc refcount to 2. This caused Rc::try_unwrap to fail in list_push/map_insert, falling back to full Vec clone on every append — making all TCO functions with list accumulation O(n²). Fix: move the loop variable instead of cloning it. The tail-call path always reassigns __tco_p_state before continue, so Rust allows the move. This keeps Rc refcount at 1, enabling in-place mutation. Expected impact: tokenize_loop on 1,515 lines should drop from OOM (300MB/s growth) to O(n) with constant memory. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…Rust Comprehensive audit tracing the complete pipeline from .dag source through v1 codegen to generated Rust runtime. Identifies the root cause of generated crate OOM: char_at(source.clone(), pos) in v2_rt.rs clones the entire source string per character, making tokenization O(n²) in source length. For 50K characters, this produces ~1.25 billion character operations. Prior optimizations (Track S, R5) fixed real issues at the DAG algorithm and Rc refcounting layers, but never examined the hardcoded runtime intrinsics or the codegen's clone insertion strategy. Five smoking guns identified with exact file locations across all layers (DAG source, v1 codegen, generated Rust, runtime intrinsics). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 594b0d373d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
… plan - Complete function-by-function audit of v2_rt.rs (18 functions) - Document three structural failures that enabled the O(n²) issue - Three prevention mechanisms: 1. Performance gate tests (wall-clock assertions in generated crate) 2. Eliminate hardcoded runtime (real file or .dag-defined intrinsics) 3. Codegen complexity contract (no O(n) ops in loop bodies) - Document the codegen clone strategy interaction with runtime - Identify remaining unaudited layers (fn_codegen, type_codegen) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… R6/R7 PERF_AUDIT: Replace wall-clock gate tests with structural prevention: - Kernel primitives with declared complexity contracts - Track D cost algebra proving compiler's own bounds - Self-hosting making the proof load-bearing (fixed point = adequate perf) ROADMAP: Add R6 (string intrinsics fix — immediate Track A blocker), R7 (primitive complexity contracts — prerequisite for Track D on self). Mark R5 as done. Update Track A blocker to reference R6. The long-term answer: self-hosting eliminates v2_rt.rs, fn_codegen.rs, type_codegen.rs, and render_rust.rs. Every layer that caused the perf failure is v1 scaffolding that disappears when the compiler compiles itself. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ing, scan_* char_at: byte-indexed access (as_bytes()[pos]) instead of .chars().nth(pos) string_length: .len() instead of .chars().count() substring: byte slicing instead of skip+take+collect scan_while: direct byte access instead of rebuilding Vec<char> per call skip_horizontal_ws: byte comparison instead of Vec<char> scan_to_eol: byte scan instead of Vec<char> scan_string_end: byte scan instead of Vec<char> All functions retain non-ASCII fallback paths. All .dag source is ASCII so the fast paths always apply in practice. Expected: tokenization drops from O(n²) to O(n), gist resolve should complete without OOM. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Declares every kernel primitive (string ops, collection ops, scanner ops, map ops, record ops, I/O, serialization) with its symbolic complexity contract in dsl/std/primitives.dag. These are the operations built into the v2 runtime shim and evaluator -- the file is the single authority for their cost semantics so the Track D cost algebra can reason about bounds. Uses PrimitiveContract type with symbolic work/output_size expressions, Certainty enum (Proven/Amortized/Expected/Conservative), and per-primitive notes documenting the underlying algorithm and assumptions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Rewrite R3 from vague "reduce 650 clones" to precise fix: compile_ident in fn_codegen.rs emits &variable instead of variable.clone() for String-typed multi-use variables. The CompileContext already has ir_scope/param_types for type lookup. v2_rt functions already accept impl AsRef<str>. This is the v1 investment that unblocks self-hosting. The generated tokenizer currently clones ~4GB of strings for 1,515 lines of input. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
String function parameters in the generated v2 crate are now rendered as &str instead of String. This eliminates O(n) clones on every multi-use String variable, replacing them with O(1) pointer copies — the critical fix for the OOM in the generated compiler. Key changes: - type_codegen: String params render as &str in function signatures - fn_codegen: compile_ident emits .to_string() for &str params (stripped at call sites that also take &str) - compile_call: borrow_for_str_param converts owned String args to &ref when passing to &str params; strips .to_string() on param/literal refs - TCO: loop variables use owned String (.to_string() on init) since they must outlive the borrow scope - v2_crate_emit: fn_str_params set tracks which (fn, param_index) pairs are truly String (not Option<String>) for precise call-site optimization - v2_runtime_shim: scan_while pred changed to Fn(&str) to match generated function signatures Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wrap Value::List and Value::Set inner Vec in Arc for copy-on-write semantics. The interpreter's list_push now uses Arc::try_unwrap to push in-place when refcount=1, avoiding O(n) clone on every append. Arc (not Rc) is required because Value appears in OnceLock<TypeRegistry> which requires Sync. The serde "rc" workspace feature enables Arc serialization. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…e gist tests - Generated test template: remove .to_string() from tokenize() calls since R3 changed String params to &str - Gist resolve/compile tests: add --release to inner cargo test invocation — debug mode is 20x+ slower due to stacker wrapping, making 1,515 lines infeasible - Fix pre-existing clippy warning (useless vec! in type_codegen test) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Major restructure based on 2026-03-18 session results: - Track R (all 7 items) consolidated as DONE - Track S substantially done - New R8 (struct clone reduction) identified as next A1 blocker: gist resolve no longer OOMs (10GB→2GB) but still >20min in release due to non-String type cloning in compile_ident - ASCII dependency chart showing critical path: R8→A1→A4→A5→A6→A7 - Completed work consolidated into tables - Design decisions pending section preserved (unchanged, still need decisions) - Stale detail removed; active blockers emphasized Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Decisions made: - Blocker 1 (emitter triplication): A — shared walk with callbacks, deferred to v2 - Blocker 2 (fabrication-on-error): A — add Result<T,E> to DSL, before A6 - Blocker 3 (TypeExpr deletion): scheduled in parallel with R8 - Blocker 4 (validate_no_unresolved): marked for deletion in B3, violates Invariant 9 R8 reframed as structural principle: DAG value semantics = shared ownership = Rc. Complete callsite migration catalog: 7 type_codegen sites, 4 fn_codegen sites, 2 render_rust sites, 1 runtime site. Key simplification: Rc makes Box-wrapping infrastructure unnecessary (Rc already breaks cycles). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Removed from 00_core.dag: - TypeExpr sum type (8 variants) and ContainerKind enum - Predicate sum type (7 variants) - type_expr_to_node and 7 helper functions (type_expr_named_to_node, type_expr_product_to_node, type_expr_coproduct_to_node, field_to_node, variant_to_node, predicate_to_field_init, opt_str_or_empty, empty_node) Migrated 02_parse.dag: parse_single_predicate now returns FieldInit directly instead of constructing intermediate Predicate values that were immediately converted via predicate_to_field_init. The Predicate type existed solely as a bridge between parsing and Node construction — with the bridge inlined, the type is unnecessary. Updated v2_recursive_types_detected test to remove TypeExpr assertion. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
All non-Copy generated types (structs and non-unit-variant enums) are now Rc-wrapped at usage sites, giving O(1) clone for the value semantics that the DAG language requires. This extends the existing Rc wrapping of List<T> → Rc<Vec<T>> to all generated types. Key changes: - build_rc_wrapped_types: computes the set from type defs (excludes Copy enums, aliases, and hardcoded types like SourceSpan/BindingPower) - type_expr_to_rust: emits Rc<T> for Named types in the rc_wrapped set - compute_recursive_fields: skips Rc-wrapped types (already heap-allocated) - compile_resolved_record_expr: wraps struct/variant construction in Rc::new() - compile_match_typed: uses .as_ref() for Rc scrutinee, .map() for Optional - Field access through Rc: strips receiver .clone(), adds .clone() to field - with() intrinsic: deref+clone for struct update base - Data maps: strip Rc from static LazyLock types (Rc is not Sync) - compile_ident: always clone in Rc context (match bindings are references) Reduces generated crate errors from 1534 → 50. Remaining errors are nested Rc pattern matching (6 LiteralValue, 4 TokenKind) and a few edge cases that need sequential pattern decomposition. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Five codegen fixes to make the generated v2 crate (v2_crate_cargo_check)
compile cleanly with Rc-wrapped non-Copy types:
1. Optional Rc re-binding: When matching Option<Rc<T>> via
.as_ref().map(|__rc| __rc.as_ref()), bindings in Some(x) are &T.
Add `let x = Rc::new(x.clone());` so downstream x.clone() yields
Rc<T> instead of bare T.
2. Data-table lookup wrapping: Static data tables store bare T (Rc is
not Sync). Add .map(Rc::new) to lookup results when the value type
is Rc-wrapped, converting Option<T> to Option<Rc<T>>.
3. Nested Rc pattern destructuring: Patterns like
Expr::Literal { value: LiteralValue::LitStr { value: s } } fail
because `value` is Rc<LiteralValue>. Emit `ref value` binding with
a let-else destructure in the match body via value.as_ref().
4. Fold accumulator type annotation: render_ir_type renders Named("T")
as "T", but fold accumulators need "Rc<T>". Apply rc_wrap_ir_type
to transform the IrType before annotation rendering.
5. Unwrap field access: .unwrap() on Option<Rc<T>> yields Rc<T>.
Add is_unwrap_of_rc_wrapped check so field access on the result
adds .clone() instead of trying to move out of the Rc.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
R2 added Box-wrapping for Node.transport and Node.config as a size optimization. R8 Rc-wraps TransportBinding and ServiceConfig, making them already heap-allocated. Box on top of Rc is redundant and produces incorrect deref code (*field tries to move out of Rc). Removes the size-based override in compute_recursive_fields. All 32 generated crate "cannot move out of Rc" errors resolved. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- matches!(x.kind, ...) → matches!(&*x.kind, ...) for Rc<TokenKind>
- SourceFile { ... } → Rc::new(SourceFile { ... }) in test templates
- Vec<SourceFile> → Vec<Rc<SourceFile>> in pipeline test calls
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…tion The v2 tokenizer uses byte-indexed char_at which assumes byte position = character position. Multi-byte UTF-8 characters (em-dashes, set-theoretic symbols, box-drawing chars) in .dag source files break this assumption, causing the tokenizer to return wrong characters and hang. Changes: - Replace all Unicode mathematical symbols in comments with ASCII equivalents (e.g. [T] for set denotation, <= for subset, _|_ for bottom, -> for arrow) - Replace em-dashes (--) and box-drawing section dividers (---) in comments - Replace em-dashes in string literals that produce generated code headers - Collapse symbols.dag emoji/unicode tiers to ASCII (restore when tokenizer handles UTF-8 properly) All 49 .dag files under dsl/ and src/v2/ are now pure ASCII. Tests: 363 daglang-emit + 92 v2-compiler-tests pass. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- PERF_AUDIT.md: add SG-6 postmortem — byte/char position mismatch in R6 char_at caused silent tokenizer corruption on non-ASCII .dag files - v2_crate_emit.rs: add profile_gist_pipeline generated test (per-file and per-stage timing for tokenize/parse/resolve) - lib.rs: add v2_crate_profile_gist host test (runs profiler in release mode) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Why tests didn't catch the byte/char corruption: - All tokenizer test inputs are ASCII - v2_rt intrinsics have zero test coverage (string constant, invisible) - Gist test is #[ignore] (circular: slow because of this bug) - No token content/position assertions, only "has tokens" checks Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When multiple match arms share the same outer variant but differ in
Rc-wrapped nested sub-patterns (e.g., Literal{value: LitInt{...}} vs
Literal{value: LitStr{...}}), the R8 let-else codegen produced duplicate
outer patterns. Rust always matched the first arm, making subsequent
arms dead code that hit unreachable!() at runtime.
Fix: detect consecutive arms with the same Rc-nested outer variant key
and merge them into a single arm with an if-let / else-if-let chain.
This preserves the original match semantics while working within Rust's
inability to destructure through Rc in patterns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
render_operator_operand did not parenthesize BinOp sub-expressions, so `(value - digit) / 10` rendered as `value - digit / 10` — causing an infinite loop in int_to_string_acc (value never decreased). Root cause: needs_grouping_in_operator only checked for if/match/block/ closure/struct but not BinOp. Adding BinOp to the check ensures all nested binary expressions are parenthesized in the generated code. This was the parser hang on gcp.dag — status code 200 triggered int_to_string which looped forever due to wrong division precedence. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three layered bugs causing the >20 minute gist pipeline hang: - SG-6: char_at byte/char mismatch on Unicode source (tokenizer hung) - SG-7: R8 nested Rc pattern match regression (parser panicked) - SG-8: operator precedence — (a-b)/c rendered as a-b/c (infinite loop) Final result: 24ms total (tokenize 20ms, parse 3ms, resolve 342us). 50,000x improvement from >20 minutes. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…tructs - 04_infer.dag → 04_reconcile.dag - module v2.compiler.infer → v2.compiler.reconcile - typecheck_and_validate() → reconcile() - Updated module header: namespace reconciliation framing - Updated imports in pipeline, emit_rust, emit_python, emit_go - Updated all test references The DAG language's "types" are named structural compositions (namespaces). The pipeline stage verifies name consistency — it reconciles how names are used with how they're defined. This is not type inference in the PL sense. Full rename of internal types (TypedGraph→ReconciledGraph, TypedExpr→etc.) is deferred as a non-blocking parallel task. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
R8 done: gist resolve pipeline runs in 24ms (was >20 minutes). Critical path updated: A1 → A4 → A5 → A6 → A7. Next step: run full gist compile (reconcile + emit). Added namespace rename (Typed*→Reconciled*, ~655 refs) as non-blocking parallel task. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
resolve_node walked type structures recursively (Product -> Coproduct -> Container -> leaf -> expand -> Product -> ...) without bounds. The cycle detection (is_recursive_type) only caught named cycles at leaf lookups, not structural cycles during expansion. Added depth parameter with limit of 50 -- in practice type nesting rarely exceeds 10. Result: gist compile pipeline completes in ~30ms (was infinite loop). 50 reconciler diagnostics remain -- these are domain completeness issues (missing imports for Network, split, map, etc.), not compiler bugs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
May 7, 2026
briansrls
added a commit
that referenced
this pull request
May 7, 2026
briansrls
added a commit
that referenced
this pull request
May 7, 2026
…1702 re-dispatch enabler (#2164) * WIP: Substrate S5: Variant-aware projection metadata carrier — Anthropic #170 * fix(s5): correct DeclarationRef import + refresh parse manifest Mgr review BLOCKING (PR #2164): - Import was `v3.spec.v3_l1 { DeclarationRef }` (empty record `{}`) but brief + Director's path-(a) ratification reference the `dsl/std/serialization.dag::DeclarationRef = String` alias. - Switch to `import std.serialization { DeclarationRef }` to match the disposition'd shape. - Inline doc-comment surfaces the two-co-existing-DeclarationRefs P2 concern as a separate substrate gap (likely folds into audit-row #14 module-convergence dissolution trigger). Refresh `parse_corpus_manifest.txt` for the new `src/v3/std/coproduct_projection.dag` file row (regenerated via `cargo test -p v3-compiler refresh_handwritten_parse_snapshot_manifest -- --ignored`). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(s5): regen bootstrap snapshot + name WireTagValue dissolution trigger Two BLOCKING inline review findings on PR #2164: 1. **Bootstrap snapshot missing carrier** (line 41 finding) — P2 facts- flow-forward violation. Ran `cargo run -p v3-compiler --bin regen_bootstrap --features bootstrap-regen-fresh`; refreshes `bootstrap_generated.rs` + `bootstrap_generated_without_parse_surface.rs` + `bootstrap_std_generated.rs` so `CoproductProjection` is present for downstream consumers. 2. **WireTagValue 🟡 SCAFFOLD with "none" trigger** (line 67 finding) — INVARIANTS.md P5 violation. Replace "Dissolution trigger: none" with a named two-clause checkable trigger: (a) §1.8 gates #29-#30 close (closure-predicate consumers land), AND (b) at least one non-StringTag arm added in response to an observed wire surface, OR explicit Practice-4 closure receipt naming the substrate observation that all REST/LLM wire boundaries are string-shaped. Plus re-escalate-to-Mgr clause if a consumer surfaces a wire-tag shape StringTag cannot express. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(s5): add CoproductProjection to bootstrap_authority + refresh manifest CI failures on PR #2164 sha 4952323: - `parse_stage4_prep::handwritten_parse_snapshot_matches_manifest`: stale hash for `coproduct_projection.dag` — manifest carried the pre-`e65efb82b` hash (before WireTagValue trigger expansion). - `pb1_bootstrap_full_snapshot_test::bootstrap_authority_rows_match_full_bootstrap_source_files`: the new `src/v3/std/coproduct_projection.dag` was bundled by the build.rs std staging but missing from `bootstrap_authority.dag`'s authority map → P2 single-authority gap. Fixes: - Add `"src/v3/std/coproduct_projection.dag": V3StdAuthority` row (alphabetical between computation_model and cross_target_coverage). - Re-run `regen_bootstrap` + `refresh_handwritten_parse_snapshot_manifest` so `bootstrap_generated*.rs` + `parse_corpus_manifest.txt` reflect the current carrier content + authority. Both tests verified green locally. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * WIP: Substrate S5: Variant-aware projection metadata carrier — Anthropic #170 * fix(s5): resolve codex BLOCKING + lane2 stack overflow via typed DeclarationRef Three concurrent findings on PR #2164 sha 0e18508: 1. **Codex BLOCKING** (sha 0e18508): `bootstrap_authority.dag` omits `dsl/std/serialization.dag`, so `import std.serialization { DeclarationRef }` cannot materialize the String-alias authority in generated snapshots → silent mis-resolution at bootstrap time. 2. **Codex non-blocking** (sha c9ebf46): `v3.spec.v3_l1::DeclarationRef` "resolves to a typed declaration reference" (not String); the debt-paydown row prose treating it as a String alias is incorrect. 3. **CI v3 lane2 stack overflow** (sha 0e18508): `lane2_stage_2d_ symbolic_cost_test::branch_reports_constant_when_both_arms_constant` stack-overflows on the 2MB test-thread default after my carrier adds traversal pressure to the bootstrap; reproduces locally. Resolution path (codex's option (b)): switch carrier to import the structural typed `DeclarationRef` from `v3.spec.v3_l1` (already in `V3SpecAuthority`, no authority gap), reverting Mgr's prior `std.serialization` BLOCKING — the tri-way tension resolves cleanly: - No bootstrap-authority gap (v3_l1 is already authoritative) - No stack-budget regression (no new files added to bootstrap) - Debt-paydown row retitled per codex non-blocking guidance: "unrefined-any-declaration handle" — same #1175 substrate gap classification as MethodRef (`methods.dag:31`) + CallableRef (`services.dag:86-100`); dissolution trigger is the shared refinement-typing-on-DeclarationRef landing. - Inline doc-comment in carrier surfaces the tri-way tension and selection rationale for future readers. Additional carrier reduction (separate from BLOCKINGs but absorbed in this commit since it's the lane2 stack-overflow root cause): - `FieldShape {}` removed; `FieldProjection.Fields { fields: Map<String, FieldShape> }` collapsed to `FieldProjection.Populated` placeholder. Defers populated-case detail to first paydown PR per Mgr-disposed sliced-follow-up; same Practice-4 SCAFFOLD discipline. Gate 3 (i) trivial in-PR demo (Mgr-disposed at #2154 c#4400893282) attempted via `data anthropic_chat_message_projection: CoproductProjection` but blocked by DSL parser limitation: `Map<K, Record>` literals are not supported in `src/v3/std/` layer (zero precedent; tried both inline records and ident-reference values, both surfaced `expected field label, got StringLit("UserMessage")` parse errors). Surfaced as substrate-tooling gap; gate 3 reverts to (ii) defer-to- first-paydown disposition. Inline doc-comment in carrier captures the parser limitation for the lane. All gates green locally: - `cargo run regen_bootstrap` ✓ - `cargo test refresh_handwritten_parse_snapshot_manifest --ignored` ✓ - `parse_stage4_prep::handwritten_parse_snapshot_matches_manifest` ✓ - `pb1_bootstrap_full_snapshot_test` 8/8 ✓ - `lane2_stage_2d_symbolic_cost_test::branch_reports_constant_when_both_arms_constant` ✓ (no overflow) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(s5): single-authority debt prose — remove stale String-alias references gpt-5-5-pro REQUEST_CHANGES on PR #2164 sha 6de6a4d (BLOCKING): substrate scaffolds need a single live authority for what is debt and what dissolves it. After the import-switch to typed v3.spec.v3_l1 DeclarationRef, several comments still described the prior String-alias disposition, and `VariantId`'s dissolution trigger named an audit-row #14 OR string-bridge OR-trigger that conflicted with the ledger's new shared-with-#1175 trigger. Fixes: 1. **Top-of-file "DeclarationRef alias debt" block** (was: "DeclarationRef = String alias accepted for this slice; ... 2-clause OR-trigger audit-row #14 OR string-bridge"): retitled "DeclarationRef debt — unrefined-any-declaration handle"; describes the typed-import path and shared #1175 dissolution. 2. **`VariantId` SCAFFOLD trigger** (was: "(a) audit-row #14 closes module-convergence OR (b) string-identity-bridge surfaces"): rewritten to single-authority shared-#1175 trigger paired with `DeclarationRef` ledger row — no separate OR-trigger applies; the carrier-level dissolution is single-keyed on #1175. 3. **`CoproductProjection.declaration` doc** (was: "carries the `DeclarationRef = String` alias soft-typed handle per Director observation #1 disposition (b)"): rewritten to describe the structural typed reference + shared #1175 trigger. Remaining mention of "`DeclarationRef = String` alias" at the import- site comment is intentional historical context — it explains what was rejected (and why) in the selection rationale; not a current-state characterization of the carrier. All local gates green post-fix: - `regen_bootstrap` ✓ + manifest refresh ✓ - `parse_stage4_prep::handwritten_parse_snapshot_matches_manifest` ✓ - `lane2_stage_2d_symbolic_cost_test::branch_reports_constant_when_both_arms_constant` ✓ Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * WIP: Substrate S5: Variant-aware projection metadata carrier — Anthropic #170 * fix(s5): bump lane2 symbolic-cost test stack to 8MB; restore FieldProjection Post-merge of main (commit f29f34c brought in 36 lines of dsl/std/integer.dag + extdeps changes) re-triggered lane2_stage_2d_symbolic_cost_test:: branch_reports_constant_when_both_arms_constant stack overflow that had been resolved at sha 0998705. Carrier-shape reduction alone could not keep this PR under the 2MB cliff after the merge. **Substrate-tooling fix**: apply the existing 8MB-stack-thread precedent (m2_substrate_inhabitance_test.rs:23-35's `with_full_bootstrap_stack` pattern) to the failing test. The test's existing doc-comment explicitly documents it as a ratchet exception that pays a cold bootstrap+pipeline compile and is on the budget edge; the named dissolution trigger ("cache bootstrap Dag state as input to compile_to_dag") remains the load-bearing fix. Stack bump is the cliff-edge workaround until that lands. With the test-thread budget fixed, restore FieldProjection (`Empty | Populated`) on `CoproductVariantProjection.field_projection` — brings the carrier back to brief's path-(a)-ratified shape (per- variant single-keyed fact: payload projection + wire-tag value). Populated-case detail (typed `Map<String, FieldShape>` payload field projection) still deferred to first paydown PR consumer per Mgr- disposed sliced-follow-up. Verified locally: - regen_bootstrap ✓ - refresh_handwritten_parse_snapshot_manifest ✓ - parse_stage4_prep::handwritten_parse_snapshot_matches_manifest ✓ - pb1_bootstrap_full_snapshot_test 8/8 ✓ - lane2_stage_2d_symbolic_cost_test::branch_reports_constant_when_both_arms_constant ✓ (no overflow) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(s5): disambiguate debt-row PR reference per cursor exploratory cursor/composer-2 sha f29f34c exploratory note flagged that the debt-paydown row's 'PR introducing #1947' phrasing was ambiguous; clarified to 'PR #2164, closing #1947' since #1947 is the work-item issue and #2164 is the PR ID. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.