Skip to content

Implement testgen infrastructure with node I/O examples - #21

Merged
briansrls merged 10 commits into
mainfrom
claude/review-todo-docs-BpmEA
Feb 3, 2026
Merged

briansrls merged 10 commits into
mainfrom
claude/review-todo-docs-BpmEA

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Summary

This PR implements the testgen improvements roadmap, establishing a complete infrastructure for auto-generating unit tests from DAG node I/O examples. The system enforces MockSpecs for transport nodes, generates per-node tests, and integrates staleness checks into the build pipeline.

Key Changes

Core Infrastructure (Phase 1-2 Complete)

  • Added NodeExample type with builder pattern to MockSpec for defining I/O examples
  • Implemented OutputMatcher enum supporting exact values, substring matching, non-empty checks, and custom predicates
  • Extended MockSpec with .node_example() method to attach examples to mock specifications
  • Generated test code now calls execute_single_node() with real execution mode, validating node behavior in DAG context

Test Generation (Phase 2)

  • Implemented per-node test generation in core/testgen/src/codegen.rs:1435-1539
  • Tests are controlled by TestConfig.example_tests flag (enabled by default)
  • Each example generates a separate test function with proper error handling and assertions
  • Test naming follows test_example_{node_id}_{index} pattern

Build Integration (Phase 3)

  • Added MetaTarget.additional_deps field to support extra dependencies beyond prep level
  • Wired testgen-check as a prerequisite of the test target via additional_deps
  • Updated Makefile rendering to compose dependencies from prep level + additional deps
  • make test now enforces testgen freshness before running tests

Validation & Enforcement

  • TestGenerator::generate() fails with detailed diagnostics if DAG has transport nodes but no MockSpec
  • testgen --check mode verifies generated files match what would be generated (full-regen comparison)
  • Staleness check prevents drift between DAG definitions and generated tests

Design Decisions

Examples on MockSpec, not Node: Examples live on MockSpec via .node_example() rather than on Node<T> directly. This keeps Node<T> generic and free of test concerns, while MockSpec already serves as the test specification.

Commit generated tests with staleness checks: Generated tests are committed to source control (not .gitignore'd) because:

  • Visible in PR diffs — reviewers see test changes alongside DAG changes
  • Works without build step — cargo test works immediately after clone
  • CI enforces freshness — make test runs testgen-check, fails if stale

Remaining Work

  • TODO 1.3: Auto-discover DAGs to eliminate hardcoded builder map in gunbc-dag/src/bin/testgen.rs
  • Add node_example() calls to existing MockSpecs (bootstrap, CI, makegen, LLM) to populate the system with real examples

Testing

The infrastructure is in place and tested. Next step is populating MockSpecs with actual node examples to demonstrate the system in action.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC

@briansrls
briansrls force-pushed the claude/review-todo-docs-BpmEA branch from ca7558b to cc8ed8b Compare February 3, 2026 18:55
…tate

Add MetaTarget.additional_deps field to support extra prerequisites beyond
prep_level. The test target now depends on testgen-check, so `make test`
fails early if generated tests are stale rather than running against
outdated test code.

Update testgen-improvements.md: check off 9 of 11 TODOs that are already
implemented (MockSpec enforcement, NodeExample types, per-node test
generation, Makefile targets, staleness check). Resolve open questions
(commit with staleness checks, builder pattern syntax). Document remaining
work: auto-discover DAGs (TODO 1.3) and adding node_examples to existing
MockSpecs.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Enforce that pure nodes have I/O examples (or explicit skip) by panicking
at testgen time when coverage is missing. Add node_examples to all 7
testgen targets: bootstrap (6 nodes), CI (13 nodes), makegen (3 nodes),
and LLM (2 nodes × 4 variants).

Fix three codegen bugs found during coverage:
- to_check_code() Exact: deref assert_eq! to compare *&Value with Value
- to_check_code() Contains: add {:?} format placeholder for value arg
- to_check_code() all: add trailing semicolons to generated assertions
- value_to_rust_literal(): add Value::Skipped support
- generate_node_example_tests(): sort HashMap iteration for determinism

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Captures hacks/fallbacks found during node_example coverage work:
1. Satisfies matcher emits comment instead of assertion (~8 CI examples)
2. Parse nodes only test skip path, not actual parsing (7 nodes)
3. NonEmpty matcher vacuously passes on non-string Values
4. value_to_rust_literal catch-all silently degrades to "<MOCK>"

All are consequences of codegen not being able to serialize closures
or complex Value variants to Rust source. Not blocking — tests pass —
but they reduce the depth of what's actually verified.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Three codegen fixes:
- Collapse runs of underscores in test name sanitization (non-snake-case)
- Omit `mut` on inputs HashMap when no inputs are inserted (unused-mut)
- Prefix output variables with `_` when matcher emits only a comment,
  i.e. Satisfies and Any variants (unused-variables)

Added OutputMatcher::generates_assertion() to distinguish matchers that
emit executable assertions from those that emit comments.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
@briansrls
briansrls force-pushed the claude/review-todo-docs-BpmEA branch from 38af264 to e11a501 Compare February 3, 2026 21:10
…tion.md

Survey 885 manually written tests, extract 5 consolidation patterns
(function unit tests, graph structure, signature validation, mode-based,
execution mode). Document integration test gaps with Fermi sizing
(XS/S/M/L/XL) and propose make test-integration/test-external targets.
Regenerate test files after main merge added cardinality coverage (B.3).

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Replace ad-hoc XS/S/M/L/XL sizing with structural categorization
derived from the transport type system: integration = hermetic
transport ops (File, Shell), external = non-hermetic (Rest, Http, Tcp).
Edge case: Shell ops that hit the network (e.g. gh gist create)
belong in external despite the Shell variant.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
TransportRequest::Shell erases hermeticity: GitRequest (hermetic) and
GistRequest (non-hermetic) both produce identical ShellRequest values.
Hermeticity is a producer-level property, not a transport-level one.
Reframe section 6 to lead with this as a design constraint, classify
by producer type instead of variant, and document 4 options to fix
the type system gap.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
…fragility

TODO_hacks.md:
- Hack #5: cardinality Empty tests use concrete empty values (false/0/"")
  not actual absence. BoundaryMocks has no way to represent "absent."
  Blocks meaningful B.3 boundary testing for scalar types.
- Hack #4 amendment: Value::List filter_map silently drops non-string
  elements (separate from catch-all issue).
- Hack #3 note: runtime check() already handles lists correctly;
  codegen to_check_code() just needs to mirror it.
- Notes: typed matchers as finite logic language; Hack #5 priority.

consolidation.md:
- §5: type_id == "List" dual encoding across 4+ locations. Canonical
  model is element type + cardinality. Migration strategy documented.
- §6: Codebase fragility — builder functions as strings (rename-unsafe),
  buck-out/gen in 16 locations (no single constant), CODEGEN_SOURCES
  hardcoded directory list (staleness gap).
- Tasks updated for all new items.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Captures longer-term directions from review feedback that aren't
blocking but inform design:
- Ports referencing TypeContract instead of raw type_id strings
- Value::is_empty() centralized emptiness semantics
- Typed matcher variants as a finite logic language (replacing Satisfies)
- Tape order semantics decision (defer, keep ordered)
- Boundary witness generation from TypeContract
- Cardinality-driven CLI repeatable detection

Also includes minor updates to TODO_hacks.md and consolidation.md
from user edits.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
Phase 8: 49 hand-written tests across 8 graph_mock.rs files mapped to
5 patterns (A-E). Patterns A (boundary presence) and C (self-chain)
are safe to delete now — testgen already generates equivalents. Three
new testgen assertions needed: transport-mock coverage (TODO 8.1),
signature validation (TODO 8.2), mock-value type compatibility (8.3).

Phase 9: DagSpec end-state unifying builder + MockSpec + signature +
testgen config into one definition. "DAG definition = test specification"
in concrete form.

consolidation.md: Pattern 6 added (graph_mock.rs test blocks),
notes on no_boundary_tests() caveat and registry-driven targets.

https://claude.ai/code/session_01FcPT2VEZdE1W7QdL7SjJgC
@briansrls
briansrls merged commit 9fb94ee into main Feb 3, 2026
1 check passed
briansrls added a commit that referenced this pull request Apr 7, 2026
- #13/#14: is_valid_proof now validates proof.dimensions against edge
  evidence lengths (fail-closed on mismatch). proof_has_non_descending_cycle
  passes the actual TerminationProof instead of fabricating an empty one.

- #20: ComplexityViolation.reason now carries the structural root cause
  from CostUnknown (e.g., "same-argument recursion in X") instead of
  the asymptotic class string. Added extract_unknown_reason helper.

- #21: ComplexityViolation now carries the function's SourceSpan from
  FuncEntry. complexity_diagnostics propagates v.span instead of no_span().

- CI: DIAG_RATCHET updated 316 → 526 to match honest violation count
  after CostUnknown restoration.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 9, 2026
Remove stale reference to the retired i64-carrier-only test name; point at
uint64_upper_half_literal_tokenizes_and_narrows for R3 gate #22.

Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls added a commit that referenced this pull request May 9, 2026
* WIP: R3 gate #22: int lit full magnitude consumer

* WIP: R3 gate #22: int lit full magnitude consumer

* WIP: R3 gate #22: int lit full magnitude consumer

* WIP: R3 gate #22: int lit full magnitude consumer

* WIP: R3 gate #22: int lit full magnitude consumer

* WIP: R3 gate #22: int lit full magnitude consumer

* refactor(v3-emit): rename Int literal pattern binding to decimal

render_value (emit / rust / python) and lens_testgen: LiteralBits::Int payload is a signed decimal string; use `decimal` instead of `n` and clone the string.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(integration): refresh parse corpus manifest after Int surface String

SG-2 handwritten parse snapshot hashes drift when SurfaceLiteral::Int
Debug output changes; regenerate via refresh_handwritten_parse_snapshot_manifest.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: regen bootstrap snapshots after rebase onto main

Re-run regen_bootstrap so committed fixtures match merged .dag + compiler emit after resolving rebase conflicts.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(test): align gate #21 doc with UInt64 full-magnitude path

Remove stale reference to the retired i64-carrier-only test name; point at
uint64_upper_half_literal_tokenizes_and_narrows for R3 gate #22.

Co-authored-by: Cursor <cursoragent@cursor.com>

* WIP: R3 gate #22: int lit full magnitude consumer

* fix(v3): preserve full-magnitude range bounds in where synthesis

range(min/max) synthesis no longer requires i64::from_str on bounds, which
sent valid ultrawide Int literals down the placeholder path (fail-closed
gap vs R3 decimal-string carrier). Validate with BigInt::from_str and thread
the authored decimal text into SurfaceLiteral::Int.

Also fix evaluator unit test helpers to use literal_bits_int after
LiteralBits::Int became String-backed.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls added a commit that referenced this pull request May 11, 2026
* WIP: R3 gate #21: int refinement overflow proven parametric

* Fix gate 21 overflow receipt bounds
briansrls added a commit that referenced this pull request May 13, 2026
…ry (post-#2847-merge follow-ons) (#2884)

* docs(r3+r4): §1.8 row #28 N-projection expansion + WISHLIST §R4.E full-stack-from-one-.dag entry (post-#2847-merge PM follow-ons)

Both follow-ons unblocked by R4 path-b canvas merge (PR #2847 squash 1f88306 2026-05-13T08:05:02Z, Director-ratified Q1-b/Q2-a/Q3-a/Q4-extend/Q5-a/Practice-4 + anti-patterns + scope extension).

**§1.8 row #28 ledger update** (Task #20 — Director Q4 ratification msg_7d51b699):
- Made NAME layer-count-agnostic per Director rationale (current description's enumeration was incidental, not authoritative)
- Cited current PASSING projection set (Rust + canonical route + OpenAPI + Markdown + SQL DDL)
- Added R4 extension scope: TS (Shape-A) + React (Shape-A) per ratified canvas; test surface extends to N-target consistency
- Encoded Director anti-pattern #6 verbatim: introducing parallel `omni_*_share_one_node_tree` gates is INVARIANTS P1 violation

**WISHLIST §R4.E entry** (Task #21 — Director-suggested entry text):
- Full-stack-from-one-`.dag` with React framework substrate (R4-Phase-1..5)
- All 5 Q-ratifications cited (Q1-b TypingDiscipline / Q2-a Shape-A / Q3-a single-authority / Q4-extend / Q5-a Behavior::Bind)
- Composes-with notes: R4.A omni-ingestion + R4.B Introspect-lens + R4.C low-level emission + R4.D faithfulness
- Phase 1.5 HookKind Practice-4-promotion canvas requirement noted (pre-Phase-2 dispatch per Director)
- Distinct from multi-program-coordination canvas (deferred per msg_3bf3df9c; forward-pointer at §3.8)
- Connection to R3 path (a) demo PR #2848 (4-layer cash for Rust + OpenAPI + Markdown + SQL DDL projections)

Authority chain (verbatim cites):
- Operator directive 2026-05-13 + ratification of paths (a)+(b)
- Director msg_7d51b699 (Q1-Q5 + Practice 4 + anti-patterns #7+#8 + 5-phase plan)
- Director msg_2c1bfb0e (Q6 negative-degree scope extension + Q7 SymbolicCost preservation + anti-pattern #9)
- Director msg_3bf3df9c (option C defer disposition for multi-program-coordination)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(r3): §1.8 row #28 — disambiguate "projections added" → "in scope" per cursor exploratory observation on PR #2884

Cursor review 10956 (non-blocking exploratory):
> "the phrase 'TS (Shape-A) + React (Shape-A) projections added' sits under
> 'R4 extension scope'; a hurried reader could still read 'added' as
> 'already shipped.' If that ambiguity shows up in review chatter, a tiny
> edit like 'projections in scope' or 'projections authorized' would
> remove the misread without changing meaning."

Cursor's verdict was APPROVE; this is the optional polish edit.

Tightening:
- "TS (Shape-A) + React (Shape-A) projections added" → "TS (Shape-A) + React (Shape-A) projections in scope"
- Added explicit framing: "R4-Phase-1..5; NOT shipped at R3-close — authorized for R4 implementation post-R3"

Removes the "already-shipped" misread without changing meaning. Aligns with row's
CONSUMER_LANDED + PASSING status cell (which refers to current 4-projection set,
not the R4-extended N-projection set).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request May 14, 2026
The Blocker column listed #19–#23 as open while Current dispatch promotes

Co-authored-by: Cursor <cursoragent@cursor.com>
#20 to CONSUMER_LANDED + PASSING, which overlapped the range. Enumerate
#19, #21–#23, and #67 explicitly (P2 single-authority).
briansrls added a commit that referenced this pull request May 14, 2026
- Resolve §1.8 conflict: main gate #19 PASSING + session gate #20 PASSING
- §3 T-Numeric: Blocker lists #17/#21–#23/#67; exclude promoted #18/#19/#20/#24
- PM compile note: §9.1 cadence + 2026-05-14 §3 touch; cite gate #19 receipt
- r3-close-predicate: gate #19 HARNESS_NAMED; bucket 46 PASSING / 28 DECLARED;
  verdict 49 HARNESS_NAMED; ledger anchor notes row #19 merge

Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls added a commit that referenced this pull request May 14, 2026
The Blocker column listed #19–#23 as open while Current dispatch promotes

Co-authored-by: Cursor <cursoragent@cursor.com>
#20 to CONSUMER_LANDED + PASSING, which overlapped the range. Enumerate
#19, #21–#23, and #67 explicitly (P2 single-authority).
briansrls added a commit that referenced this pull request May 14, 2026
The Blocker column listed #19–#23 as open while Current dispatch promotes

Co-authored-by: Cursor <cursoragent@cursor.com>
#20 to CONSUMER_LANDED + PASSING, which overlapped the range. Enumerate
#19, #21–#23, and #67 explicitly (P2 single-authority).
briansrls added a commit that referenced this pull request May 14, 2026
- §3 T-Numeric Blocker/ETA: exclude promoted #19/#21 from open set (P2)
- PM compile: cite #19 + #21 PASSING receipts
- Predicate index: #19/#21 HARNESS_NAMED; bucket 47 PASSING / 27 DECLARED; verdict 50/56

Co-authored-by: Cursor <cursoragent@cursor.com>
@briansrls
briansrls deleted the claude/review-todo-docs-BpmEA branch June 1, 2026 18:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants