Repository navigation
Implement testgen infrastructure with node I/O examples - #21
Merged
Merged
Conversation
briansrls
force-pushed
the
claude/review-todo-docs-BpmEA
branch
from
February 3, 2026 18:55
ca7558b to
cc8ed8b
Compare
…tate Add MetaTarget.additional_deps field to support extra prerequisites beyond prep_level. The test target now depends on testgen-check, so `make test` fails early if generated tests are stale rather than running against outdated test code. Update testgen-improvements.md: check off 9 of 11 TODOs that are already implemented (MockSpec enforcement, NodeExample types, per-node test generation, Makefile targets, staleness check). Resolve open questions (commit with staleness checks, builder pattern syntax). Document remaining work: auto-discover DAGs (TODO 1.3) and adding node_examples to existing MockSpecs. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Enforce that pure nodes have I/O examples (or explicit skip) by panicking
at testgen time when coverage is missing. Add node_examples to all 7
testgen targets: bootstrap (6 nodes), CI (13 nodes), makegen (3 nodes),
and LLM (2 nodes × 4 variants).
Fix three codegen bugs found during coverage:
- to_check_code() Exact: deref assert_eq! to compare *&Value with Value
- to_check_code() Contains: add {:?} format placeholder for value arg
- to_check_code() all: add trailing semicolons to generated assertions
- value_to_rust_literal(): add Value::Skipped support
- generate_node_example_tests(): sort HashMap iteration for determinism
https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Captures hacks/fallbacks found during node_example coverage work: 1. Satisfies matcher emits comment instead of assertion (~8 CI examples) 2. Parse nodes only test skip path, not actual parsing (7 nodes) 3. NonEmpty matcher vacuously passes on non-string Values 4. value_to_rust_literal catch-all silently degrades to "<MOCK>" All are consequences of codegen not being able to serialize closures or complex Value variants to Rust source. Not blocking — tests pass — but they reduce the depth of what's actually verified. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Three codegen fixes: - Collapse runs of underscores in test name sanitization (non-snake-case) - Omit `mut` on inputs HashMap when no inputs are inserted (unused-mut) - Prefix output variables with `_` when matcher emits only a comment, i.e. Satisfies and Any variants (unused-variables) Added OutputMatcher::generates_assertion() to distinguish matchers that emit executable assertions from those that emit comments. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
briansrls
force-pushed
the
claude/review-todo-docs-BpmEA
branch
from
February 3, 2026 21:10
38af264 to
e11a501
Compare
…tion.md Survey 885 manually written tests, extract 5 consolidation patterns (function unit tests, graph structure, signature validation, mode-based, execution mode). Document integration test gaps with Fermi sizing (XS/S/M/L/XL) and propose make test-integration/test-external targets. Regenerate test files after main merge added cardinality coverage (B.3). https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Replace ad-hoc XS/S/M/L/XL sizing with structural categorization derived from the transport type system: integration = hermetic transport ops (File, Shell), external = non-hermetic (Rest, Http, Tcp). Edge case: Shell ops that hit the network (e.g. gh gist create) belong in external despite the Shell variant. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
TransportRequest::Shell erases hermeticity: GitRequest (hermetic) and GistRequest (non-hermetic) both produce identical ShellRequest values. Hermeticity is a producer-level property, not a transport-level one. Reframe section 6 to lead with this as a design constraint, classify by producer type instead of variant, and document 4 options to fix the type system gap. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
…fragility TODO_hacks.md: - Hack #5: cardinality Empty tests use concrete empty values (false/0/"") not actual absence. BoundaryMocks has no way to represent "absent." Blocks meaningful B.3 boundary testing for scalar types. - Hack #4 amendment: Value::List filter_map silently drops non-string elements (separate from catch-all issue). - Hack #3 note: runtime check() already handles lists correctly; codegen to_check_code() just needs to mirror it. - Notes: typed matchers as finite logic language; Hack #5 priority. consolidation.md: - §5: type_id == "List" dual encoding across 4+ locations. Canonical model is element type + cardinality. Migration strategy documented. - §6: Codebase fragility — builder functions as strings (rename-unsafe), buck-out/gen in 16 locations (no single constant), CODEGEN_SOURCES hardcoded directory list (staleness gap). - Tasks updated for all new items. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Captures longer-term directions from review feedback that aren't blocking but inform design: - Ports referencing TypeContract instead of raw type_id strings - Value::is_empty() centralized emptiness semantics - Typed matcher variants as a finite logic language (replacing Satisfies) - Tape order semantics decision (defer, keep ordered) - Boundary witness generation from TypeContract - Cardinality-driven CLI repeatable detection Also includes minor updates to TODO_hacks.md and consolidation.md from user edits. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Phase 8: 49 hand-written tests across 8 graph_mock.rs files mapped to 5 patterns (A-E). Patterns A (boundary presence) and C (self-chain) are safe to delete now — testgen already generates equivalents. Three new testgen assertions needed: transport-mock coverage (TODO 8.1), signature validation (TODO 8.2), mock-value type compatibility (8.3). Phase 9: DagSpec end-state unifying builder + MockSpec + signature + testgen config into one definition. "DAG definition = test specification" in concrete form. consolidation.md: Pattern 6 added (graph_mock.rs test blocks), notes on no_boundary_tests() caveat and registry-driven targets. https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
briansrls
added a commit
that referenced
this pull request
Apr 7, 2026
- #13/#14: is_valid_proof now validates proof.dimensions against edge evidence lengths (fail-closed on mismatch). proof_has_non_descending_cycle passes the actual TerminationProof instead of fabricating an empty one. - #20: ComplexityViolation.reason now carries the structural root cause from CostUnknown (e.g., "same-argument recursion in X") instead of the asymptotic class string. Added extract_unknown_reason helper. - #21: ComplexityViolation now carries the function's SourceSpan from FuncEntry. complexity_diagnostics propagates v.span instead of no_span(). - CI: DIAG_RATCHET updated 316 → 526 to match honest violation count after CostUnknown restoration. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
May 9, 2026
Remove stale reference to the retired i64-carrier-only test name; point at uint64_upper_half_literal_tokenizes_and_narrows for R3 gate #22. Co-authored-by: Cursor <cursoragent@cursor.com>
This was referenced May 9, 2026
briansrls
added a commit
that referenced
this pull request
May 9, 2026
* WIP: R3 gate #22: int lit full magnitude consumer * WIP: R3 gate #22: int lit full magnitude consumer * WIP: R3 gate #22: int lit full magnitude consumer * WIP: R3 gate #22: int lit full magnitude consumer * WIP: R3 gate #22: int lit full magnitude consumer * WIP: R3 gate #22: int lit full magnitude consumer * refactor(v3-emit): rename Int literal pattern binding to decimal render_value (emit / rust / python) and lens_testgen: LiteralBits::Int payload is a signed decimal string; use `decimal` instead of `n` and clone the string. Co-authored-by: Cursor <cursoragent@cursor.com> * test(integration): refresh parse corpus manifest after Int surface String SG-2 handwritten parse snapshot hashes drift when SurfaceLiteral::Int Debug output changes; regenerate via refresh_handwritten_parse_snapshot_manifest. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: regen bootstrap snapshots after rebase onto main Re-run regen_bootstrap so committed fixtures match merged .dag + compiler emit after resolving rebase conflicts. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(test): align gate #21 doc with UInt64 full-magnitude path Remove stale reference to the retired i64-carrier-only test name; point at uint64_upper_half_literal_tokenizes_and_narrows for R3 gate #22. Co-authored-by: Cursor <cursoragent@cursor.com> * WIP: R3 gate #22: int lit full magnitude consumer * fix(v3): preserve full-magnitude range bounds in where synthesis range(min/max) synthesis no longer requires i64::from_str on bounds, which sent valid ultrawide Int literals down the placeholder path (fail-closed gap vs R3 decimal-string carrier). Validate with BigInt::from_str and thread the authored decimal text into SurfaceLiteral::Int. Also fix evaluator unit test helpers to use literal_bits_int after LiteralBits::Int became String-backed. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
This was referenced May 10, 2026
briansrls
added a commit
that referenced
this pull request
May 10, 2026
briansrls
added a commit
that referenced
this pull request
May 11, 2026
3 tasks
briansrls
added a commit
that referenced
this pull request
May 13, 2026
…ry (post-#2847-merge follow-ons) (#2884) * docs(r3+r4): §1.8 row #28 N-projection expansion + WISHLIST §R4.E full-stack-from-one-.dag entry (post-#2847-merge PM follow-ons) Both follow-ons unblocked by R4 path-b canvas merge (PR #2847 squash 1f88306 2026-05-13T08:05:02Z, Director-ratified Q1-b/Q2-a/Q3-a/Q4-extend/Q5-a/Practice-4 + anti-patterns + scope extension). **§1.8 row #28 ledger update** (Task #20 — Director Q4 ratification msg_7d51b699): - Made NAME layer-count-agnostic per Director rationale (current description's enumeration was incidental, not authoritative) - Cited current PASSING projection set (Rust + canonical route + OpenAPI + Markdown + SQL DDL) - Added R4 extension scope: TS (Shape-A) + React (Shape-A) per ratified canvas; test surface extends to N-target consistency - Encoded Director anti-pattern #6 verbatim: introducing parallel `omni_*_share_one_node_tree` gates is INVARIANTS P1 violation **WISHLIST §R4.E entry** (Task #21 — Director-suggested entry text): - Full-stack-from-one-`.dag` with React framework substrate (R4-Phase-1..5) - All 5 Q-ratifications cited (Q1-b TypingDiscipline / Q2-a Shape-A / Q3-a single-authority / Q4-extend / Q5-a Behavior::Bind) - Composes-with notes: R4.A omni-ingestion + R4.B Introspect-lens + R4.C low-level emission + R4.D faithfulness - Phase 1.5 HookKind Practice-4-promotion canvas requirement noted (pre-Phase-2 dispatch per Director) - Distinct from multi-program-coordination canvas (deferred per msg_3bf3df9c; forward-pointer at §3.8) - Connection to R3 path (a) demo PR #2848 (4-layer cash for Rust + OpenAPI + Markdown + SQL DDL projections) Authority chain (verbatim cites): - Operator directive 2026-05-13 + ratification of paths (a)+(b) - Director msg_7d51b699 (Q1-Q5 + Practice 4 + anti-patterns #7+#8 + 5-phase plan) - Director msg_2c1bfb0e (Q6 negative-degree scope extension + Q7 SymbolicCost preservation + anti-pattern #9) - Director msg_3bf3df9c (option C defer disposition for multi-program-coordination) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(r3): §1.8 row #28 — disambiguate "projections added" → "in scope" per cursor exploratory observation on PR #2884 Cursor review 10956 (non-blocking exploratory): > "the phrase 'TS (Shape-A) + React (Shape-A) projections added' sits under > 'R4 extension scope'; a hurried reader could still read 'added' as > 'already shipped.' If that ambiguity shows up in review chatter, a tiny > edit like 'projections in scope' or 'projections authorized' would > remove the misread without changing meaning." Cursor's verdict was APPROVE; this is the optional polish edit. Tightening: - "TS (Shape-A) + React (Shape-A) projections added" → "TS (Shape-A) + React (Shape-A) projections in scope" - Added explicit framing: "R4-Phase-1..5; NOT shipped at R3-close — authorized for R4 implementation post-R3" Removes the "already-shipped" misread without changing meaning. Aligns with row's CONSUMER_LANDED + PASSING status cell (which refers to current 4-projection set, not the R4-extended N-projection set). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This was referenced May 14, 2026
briansrls
added a commit
that referenced
this pull request
May 14, 2026
- Resolve §1.8 conflict: main gate #19 PASSING + session gate #20 PASSING - §3 T-Numeric: Blocker lists #17/#21–#23/#67; exclude promoted #18/#19/#20/#24 - PM compile note: §9.1 cadence + 2026-05-14 §3 touch; cite gate #19 receipt - r3-close-predicate: gate #19 HARNESS_NAMED; bucket 46 PASSING / 28 DECLARED; verdict 49 HARNESS_NAMED; ledger anchor notes row #19 merge Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls
added a commit
that referenced
this pull request
May 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR implements the testgen improvements roadmap, establishing a complete infrastructure for auto-generating unit tests from DAG node I/O examples. The system enforces MockSpecs for transport nodes, generates per-node tests, and integrates staleness checks into the build pipeline.
Key Changes
Core Infrastructure (Phase 1-2 Complete)
NodeExampletype with builder pattern toMockSpecfor defining I/O examplesOutputMatcherenum supporting exact values, substring matching, non-empty checks, and custom predicatesMockSpecwith.node_example()method to attach examples to mock specificationsexecute_single_node()with real execution mode, validating node behavior in DAG contextTest Generation (Phase 2)
core/testgen/src/codegen.rs:1435-1539TestConfig.example_testsflag (enabled by default)test_example_{node_id}_{index}patternBuild Integration (Phase 3)
MetaTarget.additional_depsfield to support extra dependencies beyond prep leveltestgen-checkas a prerequisite of thetesttarget viaadditional_depsmake testnow enforces testgen freshness before running testsValidation & Enforcement
TestGenerator::generate()fails with detailed diagnostics if DAG has transport nodes but no MockSpectestgen --checkmode verifies generated files match what would be generated (full-regen comparison)Design Decisions
Examples on MockSpec, not Node: Examples live on
MockSpecvia.node_example()rather than onNode<T>directly. This keepsNode<T>generic and free of test concerns, whileMockSpecalready serves as the test specification.Commit generated tests with staleness checks: Generated tests are committed to source control (not .gitignore'd) because:
cargo testworks immediately after clonemake testrunstestgen-check, fails if staleRemaining Work
gunbc-dag/src/bin/testgen.rsnode_example()calls to existing MockSpecs (bootstrap, CI, makegen, LLM) to populate the system with real examplesTesting
The infrastructure is in place and tested. Next step is populating MockSpecs with actual node examples to demonstrate the system in action.
https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC