Skip to content

OSAC-2766: Design - Type-Safe Resource References - #121

Merged
openshift-merge-bot[bot] merged 23 commits into
osac-project:mainfrom
htayrie-rh:design/OSAC-1330
Jul 22, 2026
Merged

openshift-merge-bot[bot] merged 23 commits into
osac-project:mainfrom
htayrie-rh:design/OSAC-1330

Conversation

@htayrie-rh

@htayrie-rh htayrie-rh commented Jul 16, 2026 •

Copy link
Copy Markdown
Member

Design: Type-Safe Resource References

Jira: https://redhat.atlassian.net/browse/OSAC-1330
PRD: enhancements/type-safe-resource-references/prd.md (PR #113, merged)

Summary

This design replaces all 34 opaque string reference fields across 15 public API resources with per-type structured protobuf messages (<Type>Reference for cross-tenant and <Type>LocalReference for same-tenant references). A centralized gRPC interceptor using protoreflect validates references before server handlers run, replacing scattered inline validation and standardizing error reporting on InvalidArgument with structured field paths. The interceptor fails closed — unregistered reference types cause a startup failure and runtime Internal error, not a silent pass-through. Delivery is incremental across 5 resource-group chunks (Networking → Compute → IP Management → Clusters+BareMetal → IAM), each updating all layers (proto, server, DB triggers, CLI, UI) atomically.

Requesting Review On

  • Open Question 1: Should status-level references (hub, pool mirrors) use typed messages? Proposed yes for consistency, but adds controller migration work. Feedback requested from controller maintainers.
  • Open Question 2: Should forward reference triggers (Z0002) be kept with updated JSON paths, or removed with FOR SHARE locking moved into the interceptor's DAO lookups? The current triggers use SELECT ... FOR SHARE to serialize concurrent child inserts and parent deletes under READ COMMITTED. To be resolved during Epic 1 implementation.
  • Open Question 3: How should CEL filter path changes be communicated to API consumers? Currently assumes breakage accepted per D1 — alternatives include deprecation warnings.
  • Local vs. full reference assignments (§Implementation Details): The table assigns each of the 34 fields as either local or full reference. Please verify the cross-tenant/platform-scoped categorization matches your understanding.
  • Interceptor chain position: Proposed after Transaction, before Handler — so lookups share the DB transaction.
  • Incremental delivery plan: 5 chunks ordered by dependency. Please review whether the chunking makes sense for your team's workflow.

How to Review

  • Comment inline on specific sections
  • Approve when the design accurately reflects a viable implementation approach

Summary by CodeRabbit

  • Documentation
    • Added a design document for type-safe resource references, including request-time validation, automatic reference completion, and unified BadRequest error reporting.
    • Documented how typed reference fields differ between cross-tenant/project and same-tenant/project scenarios.
    • Added guidance for required updates to REST/JSON, CLI, UI, CEL filters, and database triggers to support the new nested reference shapes.
    • Added a PRD covering scope, user scenarios, exclusions (no backward-compatibility window), and acceptance criteria for rollout readiness.

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@htayrie-rh, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 19 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: fc95303a-d9a4-417d-8253-c2763faa6f9d

📥 Commits

Reviewing files that changed from the base of the PR and between 7ab5da9 and 865d8db.

📒 Files selected for processing (2)
  • enhancements/OSAC-1330-type-safe-resource-references/design.md
  • enhancements/OSAC-1330-type-safe-resource-references/prd.md

Walkthrough

Adds a PRD and design for replacing opaque resource-reference strings with typed protobuf references, centralized request validation, nested consumer wire formats, database/CEL updates, and an incremental rollout and testing strategy.

Changes

Type-safe resource references

Layer / File(s) Summary
Proposal and API contracts
enhancements/OSAC-1330-type-safe-resource-references/prd.md, enhancements/OSAC-1330-type-safe-resource-references/design.md
Defines full and local reference messages, affected APIs, resolution modes, consumer wire formats, scope, and persona requirements.
Reference validation and storage
enhancements/OSAC-1330-type-safe-resource-references/design.md
Specifies reflective gRPC validation, tenant-scoped DAO lookups, request mutation, aggregated field violations, database trigger changes, and CEL path updates.
Consumer migration and rollout
enhancements/OSAC-1330-type-safe-resource-references/design.md
Documents delivery sequencing, security and failure handling, observability, testing, graduation criteria, operational support, and version-skew behavior.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant ReferenceValidator
  participant DAO
  participant ServerHandler
  Client->>ReferenceValidator: Send typed-reference request
  ReferenceValidator->>DAO: Resolve referenced resources
  DAO-->>ReferenceValidator: Return lookup results
  ReferenceValidator->>ServerHandler: Pass validated and populated request
Loading

Possibly related PRs

Suggested labels: lgtm

Suggested reviewers: danmanor

🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately names the main change and the design-doc nature of the pull request.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
No-Hardcoded-Secrets ✅ Passed No hardcoded secrets found; the only token/password mentions are descriptive prose, with no assignments, private keys, embedded creds, or long base64 literals.
No-Weak-Crypto ✅ Passed Only two markdown design docs changed; exact scans found no MD5/SHA1/DES/RC4/3DES/Blowfish/ECB or custom crypto code.
No-Injection-Vectors ✅ Passed Only changed file is a markdown design doc; no dangerous APIs or untrusted-data sinks found. SQL mentions are read-only SELECT ... FOR SHARE notes, not injection.
Container-Privileges ✅ Passed Doc-only PR; no container/K8s manifests or privileged settings (privileged, hostPID, hostNetwork, hostIPC, SYS_ADMIN, allowPrivilegeEscalation) were found.
No-Sensitive-Data-In-Logs ✅ Passed No sensitive-log issue found: the PR only adds docs, and its logging guidance covers resource type/name/field path—not passwords, tokens, PII, or secrets.
Ai-Attribution ✅ Passed Recent commits carry Assisted-by: Claude Code trailers, and no Co-Authored-By trailers were found.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design uses well-understood technologies throughout: per-type protobuf messages, a gRPC unary interceptor with protoreflect for message walking, parameterized DAO lookups, and JSON path updates in PL/pgSQL triggers. The incremental delivery plan (5 chunks, 8-step implementation order per chunk) is realistic and each chunk is independently shippable. The forward-compatibility strategy (including the deferred id field) is pragmatic. The backfill approach (UPDATE table SET data = data) is n
Testability 2/2 The test plan covers three levels: unit tests (Ginkgo) for reference message serialization, interceptor field detection (including nested/repeated/oneof), validation logic, and per-server validation removal; integration tests (kind cluster) for end-to-end create flows with valid/invalid/cross-tenant references, trigger enforcement, and CEL filter paths; and E2E tests (pytest) for full provisioning workflows and error scenarios. The interceptor's clean interface (registry of lookup functions, pro
Scope 2/2 Scope is precisely defined: 34 spec-level reference fields across 15 public API resources, with a complete field-by-field inventory and local-vs-full reference rationale table. Non-goals are explicit and locked (no UUID-to-name migration, no backward compatibility, no RBAC changes). The 5-chunk delivery plan groups resources by functional domain with clear inter-chunk dependencies (Chunk 2 depends on Chunk 1). Each chunk updates all layers atomically (proto, server, interceptor, DB, CLI, UI, tes
Architecture 2/2 The architecture follows sound principles: clean separation between existence validation (interceptor) and business logic validation (servers); defense-in-depth with both interceptor and database triggers enforcing referential integrity; correct interceptor chain positioning (after auth+tx, before handler) ensuring tenant context and transactional consistency; error aggregation for user-friendly multi-error responses; observability with dedicated metrics and structured logging; and a disable fla

Verdict: A thorough, well-structured enhancement proposal with concrete implementation details, sound architectural decisions (interceptor + defense-in-depth triggers), a realistic incremental delivery plan, and comprehensive test strategy across all three levels.

Feedback: The proposal is missing the User Stories section required by the template — the Workflow Description partially compensates but explicit 'As a , I want...' stories would strengthen the motivation and help reviewers validate that all personas are covered. The database backfill strategy ('UPDATE table SET data = data') should include concrete batching detail (batch size, expected runtime, locking implications) rather than a one-line acknowledgment, especially for tables that could grow large. Consider adding a proto lint rule or custom protobuf option to enforce the Reference/LocalReference naming convention programmatically, rather than relying solely on convention — this would prevent accidental false-positive detection by the interceptor if a non-reference message is ever named with a 'Reference' suffix.

Critical (0)

None.

Important (2)

  1. Missing 'User Stories' section from the template. The template instructs 'If a section does not apply to an enhancement, explain why but do not remove the section.' User Stories clearly apply here (tenant users creating resources, admins managing cross-tenant references, CLI users experiencing the flag changes). The Workflow Description covers scenarios but not in the standard user story format that enables product manager review.
  2. The database backfill approach ('UPDATE subnets SET data = data') will acquire row-level locks on every row in the table and fire triggers for each row within a single transaction. For tables with significant row counts, this risks long-running transactions, lock contention, and replication lag. The document mentions batching as a possibility but provides no concrete plan — this should be specified per chunk with estimated row counts and batch sizes.

Suggestions (3)

  1. Add a buf lint rule or custom protobuf option (e.g., '(osac.reference_target) = "VirtualNetwork"') to enforce the reference naming convention at proto compile time, preventing false-positive interceptor detection if a non-reference message accidentally ends with 'Reference'.
  2. The data migration SQL that converts string references to nested objects (the jsonb_set migration) should include a verification query that counts any remaining rows with string-typed references after migration, to confirm completeness before proceeding.
  3. Consider documenting an expected latency budget for the interceptor (e.g., 'validation of N references should add no more than Xms over the existing inline checks') beyond the p99 100ms alert threshold, to give implementers a concrete performance target during development.

Review cost

Model: claude-opus-4-6
Cost: $0.5454
Tokens: 6.1k in / 3.7k out
Cache: 86.2k read
Active time: 1m 27s
API calls: 0

@github-actions github-actions Bot added the rfe-creator-auto-reviewed EP was reviewed by AI label Jul 16, 2026
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Technically sound and implementable. The protoreflect-based interceptor for centralized validation is a proven Go pattern. Proto schema changes are mechanical and well-documented with concrete before/after examples. Database migration strategy (trigger updates + backfill) is realistic. The incremental 5-chunk delivery plan with clear ordering constraints (Chunk 1 before Chunk 2) makes the large change manageable. The interceptor placement in the gRPC chain (after auth + transaction) is correct.
Testability 2/2 Comprehensive three-tier test plan: unit tests (Ginkgo) for interceptor detection/validation logic and serialization, integration tests (kind cluster) for end-to-end reference resolution and DB trigger enforcement, and E2E tests (pytest) for full provisioning workflows. Both happy-path and error scenarios are covered, including cross-tenant references and CEL filter changes. The interceptor's deterministic behavior (message walking + DAO lookup) makes it straightforward to test with mocked DAOs.
Scope 2/2 Well-defined and appropriately sized. The scope is precisely quantified: 34 spec-level fields across 15 resources, organized into 5 delivery chunks by functional domain. Non-goals are explicit (UUID migration, backward compat, RBAC). Each chunk has a clear resource list, field inventory, and implementation order (proto -> buf generate -> interceptor -> server -> DB -> CLI -> UI -> tests). The decision to break backward compatibility (D1) simplifies scope significantly. Cross-chunk dependencies a
Architecture 2/2 Sound architectural design. The dual-layer referential integrity (gRPC interceptor + PL/pgSQL triggers) provides defense-in-depth. The separation between reference existence validation (interceptor) and business logic validation (server handlers) is clean and maintainable. The full-reference vs. local-reference distinction is well-reasoned with a clear rationale table for each field. Observability is thorough (metrics, structured logging, disable flag). The security model correctly inherits exis

Verdict: A thorough, well-architected enhancement proposal with precise scope, concrete implementation details, sound architectural patterns (interceptor + DB triggers for defense-in-depth), and a comprehensive test strategy across three tiers.

Feedback: The Alternatives section is essentially empty, deferring to an external PRD discussion -- reviewers need to see the trade-off reasoning inline, particularly the URI/ARN vs. per-type message decision and why a generic ResourceReference with a type discriminator was rejected. The template's User Stories section is missing entirely; adding 2-3 stories (tenant user creating a cross-reference, platform admin debugging a validation error, developer onboarding a new resource type) would ground the motivation in concrete personas. Consider documenting the naming-convention-based reference detection (*Reference / *LocalReference suffix matching) as a formal proto lint rule or custom option rather than relying on convention alone, since a non-reference message accidentally matching the suffix would be silently treated as a reference.

Critical (0)

None.

Important (2)

  1. Missing User Stories section: the template requires explicit 'As a , I want to so that ' stories. The document has Goals/Non-Goals but no persona-grounded user stories, which weakens the motivation for reviewers unfamiliar with the codebase.
  2. Alternatives section is effectively empty: states 'None' and defers to 'PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion.' The URI/ARN trade-off and generic-reference-message alternative should be summarized inline so the design document is self-contained.

Suggestions (3)

  1. The naming-convention-based reference detection (suffix matching on '*Reference') is fragile. Consider enforcing this as a buf lint rule or using a custom proto option/annotation to explicitly mark reference messages, preventing false positives from non-reference messages that happen to end with 'Reference'.
  2. The backfill strategy ('UPDATE table SET data = data') should include explicit batch-size guidance (e.g., 10k rows per batch with pg_sleep) for tables that may grow large, rather than leaving it as a generic 'can be batched' note in the risks section.
  3. Open Question Bump actions/setup-python from 5 to 6 #1 (status-level typed references) should include the author's recommended position rather than being fully open -- this helps reviewers converge faster.

Review cost

Model: claude-opus-4-6
Cost: $0.4852
Tokens: 789 in / 2.8k out
Cache: 90.5k read
Active time: 1m 5s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Highly feasible. The design provides concrete proto schemas (before/after), Go interceptor signatures with ReferenceLookupFunc, SQL trigger examples, and a clear 5-chunk incremental delivery plan with ordering rationale. The protoreflect-based message walking is a proven Go pattern, and the approach of detecting references by naming convention avoids manual field registration. All 34 fields are inventoried with specific proto file mappings, and consumer updates (CLI, UI, CEL filters, DB triggers
Testability 2/2 Strong test plan covering three layers: unit tests (Ginkgo) for message construction, interceptor detection/validation, and per-server validation removal; integration tests (kind cluster) for end-to-end create flows, invalid references, cross-tenant resolution, DB trigger enforcement, and CEL filters; and E2E tests (pytest) for full provisioning workflows. Observability support (osac_reference_validation_total counter, duration histogram, structured logging at DEBUG/WARN/ERROR) enables runtime v
Scope 2/2 Well-defined and appropriately bounded. Goals are specific (compile-time safety, centralized validation, standardized errors, incremental delivery). Non-goals explicitly exclude the (tenant, project, name) migration, backward compatibility, internal DB schema changes, and quota/RBAC. The 5-chunk delivery plan maps each resource group with rationale for ordering (networking first as foundation, compute next for user impact, IP management for oneof complexity, etc.). Two open questions are identif
Architecture 2/2 Sound architectural principles throughout. Separation of concerns: reference existence validation centralized in the interceptor while business logic stays in servers. Defense in depth: interceptor + PL/pgSQL triggers provide two layers of referential integrity. Forward compatibility: id field included for future (tenant, project, name) migration. Two-tier reference model (LocalReference vs Reference) correctly matches actual scoping requirements with a detailed rationale table. Interceptor posi

Verdict: A thorough, well-architected design document that covers all major dimensions (proto schema, interceptor, consumers, delivery, testing, operations) with concrete examples and sound technical decisions; the only gaps are the missing User Stories section and a thin Alternatives section.

Feedback: Add a User Stories section under Motivation as required by the template — include stories for at least the Tenant User, Tenant Admin, and Developer personas to ground the motivation in concrete user needs. Expand the Alternatives section to briefly summarize the rejected approaches (URI/ARN format, generic reference message, do nothing) inline rather than deferring entirely to the PRD discussion — reviewers should be able to evaluate trade-offs without leaving the document. Consider providing a recommended default for Open Question 1 (status-level typed references) so the chunk implementers have a clear starting point rather than waiting for controller maintainer feedback.

Critical (0)

None.

Important (2)

  1. Missing User Stories section: The template explicitly requires user stories under Motivation in the 'As a role, I want to action so that goal' format. The design has Goals/Non-Goals but no user stories, which makes it harder for product reviewers to validate that the design meets user needs.
  2. Thin Alternatives section: The section says 'None' and defers to an external PRD discussion (PR PRD: Type-Safe Resource References (OSAC-1330) #113). The template expects alternatives to be documented in the enhancement itself so reviewers can evaluate the design without external references. At minimum, briefly summarize why URI/ARN format and a generic reference message were rejected.

Suggestions (3)

  1. Open Question 1 (status-level references) should include a recommended default approach rather than leaving it fully open — this prevents implementation stalls if controller maintainer feedback is delayed.
  2. Consider documenting the interceptor's behavior for Update requests explicitly — the design focuses on Create flows but Update semantics (e.g., whether immutable reference fields are re-validated, how partial updates interact with reference validation) could be ambiguous for implementers.
  3. The UX Alignment section identifies a 'known deviation' in temp-api attachment types but does not propose a fix timeline — clarify whether this deviation is addressed in Chunk 3 (IP Management) or deferred.

Review cost

Model: claude-opus-4-6
Cost: $0.3120
Tokens: 6.1k in / 2.9k out
Cache: 119.0k read
Active time: 1m 9s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Technically sound and implementable. Uses well-established Go/gRPC patterns (protoreflect message walking, unary interceptors, DAO lookups). The incremental 5-chunk delivery plan is realistic with clear ordering rationale. No new infrastructure needed. The per-type reference message approach is straightforward protobuf, and field number reuse is explicitly justified by the no-backward-compat decision (D1). The interceptor replaces existing inline validation rather than adding new work, so net la
Testability 2/2 Three-tier test strategy (unit via Ginkgo, integration on kind cluster, E2E via pytest/osac-test-infra) covers all critical paths: reference serialization, interceptor detection of nested/repeated/oneof fields, error aggregation, end-to-end create with valid and invalid references, cross-tenant resolution, reverse trigger enforcement, and CEL filter path changes. Per-chunk test scoping is practical. Minor gap: test plan details are deferred to implementation time, but the documented strategy is
Scope 2/2 Well-bounded and appropriately sized. Clear inventory (34 fields, 15 resources) drives the work. Explicit non-goals (UUID migration, backward compat, RBAC, quota) prevent scope creep. The 5-chunk delivery plan makes the large change manageable with each chunk independently shippable. Clean separation between interceptor responsibility (existence checks) and server responsibility (business logic) keeps each chunk's surface area contained. One structural gap: the template's required User Stories s
Architecture 2/2 Strong architectural decisions throughout. The two-tier reference pattern (LocalReference vs full Reference) maps cleanly to the cross-tenant/same-tenant distinction with a clear rationale table for all 34 fields. The interceptor design centralizes validation without coupling to individual servers — new resources inherit validation automatically. Interceptor chain positioning (after auth+transaction, before handler) ensures tenant context and DB transaction availability. Forward compatibility wi

Verdict: A thorough, well-structured design that covers all layers (proto, interceptor, DB, CLI, UI) with sound architectural decisions, a practical incremental delivery plan, and comprehensive test strategy — only minor template compliance gaps (missing User Stories) and a thin Alternatives section hold it back from perfection.

Feedback: Add a User Stories section to the Motivation — the template requires it, and it would ground the three problem classes in concrete personas (e.g., 'As a tenant admin, I want to reference a platform-scoped NetworkClass by name so that the system validates the reference exists before I deploy infrastructure'). The Alternatives section should document the evaluated options (URI/ARN format, generic reference message, do-nothing) inline rather than pointing to PR #113 — reviewers shouldn't need to leave the document to understand the design trade-offs. Consider whether naming-convention-based reference detection (*Reference/*LocalReference suffix) should be hardened with a custom proto option or annotation to prevent false positives if a non-reference message name happens to match the pattern.

Critical (0)

None.

Important (2)

  1. Missing User Stories section: The template requires a User Stories section under Motivation with 'As a , I want to so that ' format. The design jumps from Motivation directly to Goals. Adding 3-4 user stories covering tenant user, tenant admin, platform admin, and developer personas would strengthen the motivation and satisfy template requirements.
  2. Thin Alternatives section: The design says 'None' and points to PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion for URI/ARN trade-off details. The template asks for alternatives to be documented inline with reasons for rejection. The evaluated options (URI/ARN strings, a single generic ResourceReference message, do-nothing) should be briefly described here with their rejection rationale so the design is self-contained.

Suggestions (3)

  1. Consider hardening reference detection: The interceptor identifies reference fields by checking if a message type name ends with 'Reference' or 'LocalReference'. While enforced by convention, a custom proto option (e.g., option (osac.reference) = true) would make detection explicit and prevent false positives from unrelated messages that happen to match the naming pattern.
  2. Address performance with high-reference-count requests: A ComputeInstance create with multiple network attachments could trigger 5+ DAO lookups in the interceptor. While the design notes these replace existing inline lookups, it would be worth confirming that the sequential lookup pattern in the interceptor doesn't regress versus the previous per-field inline approach, and whether batching lookups by type could help.
  3. Open Question 2 (CEL filter communication) deserves a recommendation: The design presents three options but doesn't recommend one. Given the breaking-change stance (D1), option (c) is consistent, but briefly stating a recommendation would help reviewers converge faster.

Review cost

Model: claude-opus-4-6
Cost: $0.5925
Tokens: 6.1k in / 3.6k out
Cache: 196.9k read
Active time: 1m 27s
API calls: 0

Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-121

Score: 10/10 | Verdict: PASS

Criterion Score Notes
What 2/2 Clear desired outcome: replace 34 opaque string reference fields across 15 API resources with per-type structured protobuf messages, centralize validation in a gRPC interceptor, and standardize error reporting. Personas explicitly named in workflow descriptions (Tenant User, Tenant Admin, Cloud Infrastructure Admin). Affected services (fulfillment-service) and all impacted API resources are enumerated. User-observable outcomes are specific: structured references, consistent InvalidArgument error
Why 2/2 Compelling business justification with three concrete problem classes: (1) no compile-time safety — a SecurityGroup ID can be passed where a VirtualNetwork ID is expected, (2) no cross-tenant addressability — references carry no tenant/project context, (3) scattered inconsistent validation — some servers return InvalidArgument, others return NotFound for the same condition. Quantified: 34 spec-level fields, 15 resources. Consequences: entangled validation logic, UI requiring extra API calls for
How 2/2 Extremely specific and measurable approach. Per-type reference message schema with concrete proto definitions and before/after examples. gRPC interceptor design with Go interfaces, protoreflect-based message walking, error aggregation. Database trigger changes (forward triggers removed, reverse triggers updated with new JSON paths). CEL filter path changes documented. CLI flag handling. UX alignment table mapping 10+ UI code locations to new wire format. Five delivery chunks with ordering ration
Task 2/2 Clearly an enhancement proposal. Introduces ~25 new protobuf message types, a new gRPC interceptor component, changes to database triggers, and API-wide structural changes. Not a bug fix or simple task. Multi-chunk delivery spanning networking, compute, IP management, clusters, and IAM.
Size 2/2 Well-scoped as one coherent capability: type-safe resource references. The three components (reference messages, validation interceptor, consumer updates) are tightly coupled — messages without validation provide no runtime safety, validation without messages has nothing to validate, consumer updates are required for the feature to function. The 5 delivery chunks organize execution by resource group but are not independent features. Each chunk leaves the system fully functional.

Verdict: A thorough, well-structured design document with clear user-facing outcomes, concrete business justification, an extremely detailed implementation plan, and appropriate scope — one of the stronger enhancement proposals in the OSAC project.

Feedback: The Alternatives section is too thin — it defers all reasoning to an external PR discussion rather than documenting the tradeoffs inline. The review-patterns guide flags this as an anti-pattern. Summarize the URI/ARN format, generic reference message, and do-nothing alternatives with 1-2 sentences each on why they were rejected, so reviewers don't need to hunt through PR comments. The two open questions (status-level reference typing, CEL filter communication strategy) should be resolved before merging to avoid scope ambiguity during implementation.

Critical (0)

None.

Important (1)

  1. Alternatives section defers entirely to external PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion instead of documenting the evaluated approaches (URI/ARN format, generic reference message, do nothing) and rejection rationale inline. This is flagged as an anti-pattern in the OSAC review patterns guide. Reviewers should not need to leave the document to understand why the chosen approach was selected over alternatives.

Suggestions (3)

  1. Resolve the two open questions before merge: (1) whether status-level references (hub, pool mirrors) should use typed messages affects controller migration scope, and (2) the CEL filter breakage communication strategy (option a/b/c) has UX implications that should be decided upfront rather than deferred.
  2. Consider adding a Terminology section (as recommended in the review patterns and modeled by the networking EP) to formally define 'full reference' vs 'local reference' vs 'platform-scoped resource' — these concepts are explained contextually but a dedicated glossary would improve scannability.
  3. The Graduation Criteria section is sparse ('will be defined when targeting a release'). Consider specifying the expected Dev Preview scope (e.g., Chunk 1 networking references) to give reviewers a concrete milestone target.

Review cost

Model: claude-opus-4-6
Cost: $0.7344
Tokens: 5.3k in / 5.3k out
Cache: 181.5k read
Active time: 2m 10s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Exceptionally detailed implementation. Concrete proto schemas with field numbers (before/after examples for SubnetSpec, ExternalIPAttachmentSpec), Go code for the interceptor (ReferenceValidator struct, ReferenceLookupFunc type, Validate method), specific database trigger names and JSON path changes (check_virtual_network_not_in_use, check_subnet_not_in_use, etc.), CLI flag patterns, and REST/JSON wire format diffs. The 5-chunk incremental delivery plan is realistic with clear ordering rationale
Testability 2/2 Test plan specifies concrete scenarios at all three levels. Unit tests: reference message serialization, interceptor field detection (nested messages, repeated fields, oneof), validation logic with field path accuracy, per-server inline validation removal. Integration tests: end-to-end create with valid/invalid references, cross-tenant NetworkClass resolution, DB trigger enforcement on delete (SQLSTATE Z0003), CEL filter with new paths. E2E: full provisioning workflow (NetworkClass -> VirtualNet
Scope 1/2 Goals and non-goals are specific and well-bounded. Non-goals explicitly exclude the (tenant, project, name) migration, backward compatibility, internal DB schema changes, and quota/RBAC. However, the design is missing explicit user stories in the 'As a [role], I want [action] so that [goal]' format required by the template. The Workflow Description section shows Tenant User and Tenant Admin scenarios but these are implementation workflows, not persona-focused user stories. Cloud Provider Admin a
Architecture 2/2 Follows all OSAC patterns. The reference message schema (full vs. local) maps cleanly to the existing resource hierarchy and tenant isolation model. The interceptor uses a registry pattern consistent with OSAC's pluggable architecture preference. Proto conventions are followed: standard message naming, field numbers, field_behavior annotations. The interceptor chain placement (after auth + transaction, before handler) is sound — it reuses tenant context and operates within the existing transacti

Verdict: A strong, deeply technical design that demonstrates thorough understanding of the OSAC architecture and provides implementation-ready detail, held back from a perfect score by missing user stories and a thin Alternatives section.

Feedback: Add explicit user stories in the 'As a [role], I want [action] so that [goal]' format for all affected personas — at minimum Tenant User, Tenant Admin, Cloud Provider Admin (who manages global catalogs/templates referenced cross-tenant), and Cloud Infrastructure Admin (who manages platform-scoped NetworkClass, IP pools). The Alternatives section needs at least one real alternative evaluated inline (e.g., the URI/ARN format or generic Reference approach mentioned in the PRD discussion) with trade-offs and rejection rationale — pointing to an external PR discussion doesn't satisfy the template requirement. Consider explicitly addressing or deferring the Documentation and Installation cross-cutting dimensions from osac-dimensions.md.

Critical (0)

None.

Important (3)

  1. Missing User Stories section: The template requires user stories in 'As a [role], I want [action] so that [goal]' format. The Motivation section describes the problem well but has no user stories. The Workflow Description shows Tenant User and Tenant Admin scenarios as implementation workflows, but these are not persona-focused user stories. Cloud Provider Admin (manages cross-tenant templates/catalogs) and Cloud Infrastructure Admin (manages platform-scoped NetworkClass, IP pools) are not repre
  2. Thin Alternatives section: The section says 'None' and points to PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion. The template and rubric require at least one real alternative with inline trade-off analysis. The URI/ARN format and generic reference message approaches were considered (per the text) but their evaluation is not included in the design document itself.
  3. Cross-cutting dimension gaps: The Documentation dimension is not explicitly addressed — API changelog and release notes are mentioned in Risks but not as a deliberate documentation plan. The Installation dimension is not addressed — the design should state whether the interceptor requires any deployment configuration or explicitly note that no installation changes are needed beyond the code deployment.

Suggestions (3)

  1. Graduation criteria could be more measurable — consider adding specific metric thresholds (e.g., 'osac_reference_validation_duration_seconds p99 < 50ms', 'zero InvalidArgument regressions in existing test suites') rather than the current qualitative conditions.
  2. The Open Question Bump actions/setup-python from 5 to 6 #1 (status-level references) could benefit from a concrete recommendation with a fallback position, rather than leaving it fully open — this would help reviewers assess the design's completeness.
  3. Consider adding a Terminology section (following the networking EP pattern noted in review-patterns.md) to formally define 'full reference', 'local reference', 'platform-scoped', and 'cross-tenant reference' — these terms are used consistently but never formally defined in a dedicated section.

Review cost

Model: claude-opus-4-6
Cost: $0.5086
Tokens: 6 in / 4.4k out
Cache: 212.1k read
Active time: 1m 47s
API calls: 0

@htayrie-rh
htayrie-rh marked this pull request as ready for review July 19, 2026 11:26
@openshift-ci
openshift-ci Bot requested review from CrystalChun and danmanor July 19, 2026 11:26

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/type-safe-resource-references/design.md`:
- Around line 534-538: Update the reference validation design so every detected
reference message type must have a registered lookup function: reject missing
registrations during startup configuration, and return a server error when an
unregistered reference is encountered at runtime instead of passing it through
unvalidated. Apply this consistently to both reference-detection and
runtime-validation flows.
- Around line 309-324: Update the rollout guidance in the attachment migration
section to explicitly require flattening attachment payloads from target: {
case, value } into the selected top-level oneof field, such as public_ip or
external_ip, containing a { name: value } reference. Clarify that
useCreatePublicIPAttachment and useCreateExternalIPAttachment call sites need
this structural change in addition to wrapping string references, and preserve
the documented proto wire format.
- Around line 580-587: Update the forward-reference removal design around
reverse delete checks so database-side race protection remains: retain an
appropriate constraint or locking mechanism, or explicitly serialize validation,
insertion, and deletion. Ensure concurrent child insertion and parent deletion
cannot both commit while leaving a dangling reference, while preserving the
updated JSON paths and reverse-reference trigger behavior.
- Around line 264-278: Resolve the catalog-item reference scope mismatch by
choosing either tenant-local or shared/full semantics for compute instance
catalog items, then consistently update ComputeInstanceCatalogItemReference or
ComputeInstanceCatalogItemLocalReference across the proto definitions, reference
registry, CLI, UI, and wire examples. Ensure the API inventory and reference
matrix use the same symbol and scope model.
- Around line 301-305: Update the subnet creation mapping for
CreateSubnetInput.virtualNetworkId to use the parent virtual network name, not
subnetName, inside spec.virtual_network.name. Rename the input/local variable
from virtualNetworkId to a name-based identifier as needed, while preserving the
existing security-group and virtual-network mappings.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 590f995c-aae6-4f09-b139-2f99204664b0

📥 Commits

Reviewing files that changed from the base of the PR and between cf09184 and fad3fbd.

📒 Files selected for processing (1)
  • enhancements/type-safe-resource-references/design.md

Comment thread enhancements/type-safe-resource-references/design.md Outdated
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/type-safe-resource-references/design.md Outdated
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deeply detailed implementation: full proto before/after examples, Go struct definitions for the interceptor (ReferenceValidator, ReferenceLookupFunc), interceptor chain ordering, database trigger migration plan with specific JSON path changes, 34-field inventory with local vs. full reference rationale per field. Error handling covers race conditions (transaction + reverse triggers), panic recovery, and error aggregation. Risks are specific (proto field number reuse, interceptor latency) with con
Testability 2/2 Test plan specifies concrete scenarios at each level: unit tests for reference serialization, interceptor field detection (nested messages, repeated fields, oneof), validation logic with error aggregation, and per-server validation removal. Integration tests describe 5 specific scenarios (valid/invalid refs, cross-tenant NetworkClass, trigger enforcement, CEL filter with new paths). E2E tests cover full provisioning workflow and error path. Per-chunk graduation criteria are measurable (all field
Scope 1/2 Boundaries are well-defined (focused on reference type changes, explicit non-goals for the tenant/project/name migration and backward compatibility). However: (1) Missing formal User Stories section under Motivation — the template requires 'As a [role], I want [action] so that [goal]' format; workflows describe persona interactions but don't satisfy the template requirement. (2) Alternatives section is weak — says 'None' and defers to PR #113 discussion rather than presenting at least one altern
Architecture 2/2 All OSAC patterns followed: tenant isolation via scoped local references and OPA-governed full references; standard proto object shape for reference messages (id, tenant, project, name); clear spec/status ownership separation; interceptor uses protoreflect for automatic field discovery with fail-closed behavior (unregistered types return Internal). Cross-repo impacts enumerated (fulfillment-service, osac-ux, CLI). Breaking changes explicitly called out with D1 decision and incremental 5-chunk de

Verdict: A thorough, well-architected design with deep implementation detail and concrete test scenarios; the main weakness is scope presentation — missing formal user stories, a thin Alternatives section, and incomplete persona/dimension coverage.

Feedback: Add a User Stories subsection under Motivation with formal 'As a [role]' stories for all four OSAC personas (especially Cloud Provider Admin, who manages global catalogs and cross-tenant template references affected by this change). Bring the Alternatives section inline — the rubric requires at least one real alternative with trade-off analysis in the document itself, not a pointer to PR discussion. Address or explicitly defer the Installation and Documentation cross-cutting dimensions (e.g., does the interceptor require any Helm chart configuration? Which API reference docs need updating per chunk?).

Critical (0)

None.

Important (3)

  1. Missing User Stories section under Motivation: the template requires formal 'As a [role], I want [action] so that [goal]' stories. The Workflow Description covers Tenant User and Tenant Admin scenarios but does not satisfy the template requirement. Cloud Provider Admin (manages global catalogs, cross-tenant resources) is not addressed at all.
  2. Alternatives section says 'None' and defers to 'PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion.' The design should be self-contained with at least one real alternative (e.g., URI/ARN format, generic reference message) and inline rationale for rejection, per template and rubric requirements.
  3. Installation and Documentation cross-cutting dimensions are silent. Does the interceptor require new Helm chart values or kustomize overlay changes? Which API reference docs, user guides, or osac-docs pages need updating per delivery chunk? These should be addressed or explicitly deferred with rationale.

Suggestions (3)

  1. Add a Terminology section defining 'local reference', 'full reference', 'reference message', 'platform-scoped resource', and 'delivery chunk' — past successful EPs (networking, bare-metal-fulfillment) define terms upfront per reviewer expectations in review-patterns.md.
  2. Consider adding a Cloud Provider Admin workflow scenario (e.g., creating a global ClusterTemplate that tenants reference cross-tenant) to demonstrate full reference resolution from the admin side.
  3. The top-level graduation criteria ('will be defined when targeting a release') could be strengthened by stating the expected Dev Preview -> Tech Preview -> GA criteria now, even if approximate, since per-chunk criteria are already well-defined.

Review cost

Model: claude-opus-4-6
Cost: $0.7411
Tokens: 5.3k in / 5.6k out
Cache: 178.5k read
Active time: 2m 4s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep implementation detail throughout. Proto schemas with field numbers and types are provided for both reference patterns (full and local). Go interface for the ReferenceValidator interceptor is defined with concrete types (ReferenceLookupFunc, protoreflect.FullName registry). Before/after examples cover proto, JSON wire format, and database trigger paths. All CRUD lifecycle is addressed (Create via interceptor, Delete via Z0003 triggers, Update implied by same interceptor path). Error handling
Testability 2/2 Test plan specifies concrete scenarios at all three levels. Unit tests cover four distinct areas: reference message serialization, interceptor field detection (including nested/repeated/oneof), validation logic with error aggregation, and per-server validation removal verification. Integration tests describe five specific scenarios on a kind cluster: valid create, invalid reference, cross-tenant reference, trigger enforcement on delete, and CEL filter with new paths. E2E tests describe full prov
Scope 1/2 Summary is concise and clear. Goals are user-visible outcomes (compile-time safety, centralized validation, standardized errors). Non-goals are specific and well-reasoned (UUID-to-name migration deferred, no backward compat per D1, no quota/RBAC changes). However, three gaps pull the score down: (1) No formal 'User Stories' subsection under Motivation -- workflow descriptions serve a similar role but don't follow the 'As a [role]...' format and only cover Tenant User and Tenant Admin, missing Cl
Architecture 2/2 Excellent alignment with OSAC patterns. Proto schemas follow standard object shape conventions. Spec/status ownership is correctly handled -- spec-level references are user-controlled and validated by the interceptor; status-level references are system-managed by controllers and explicitly excluded from interceptor validation. The interceptor chain position is precisely specified (after auth and transaction, before handler) with clear rationale. Tenant isolation is properly addressed: no new res

Verdict: A thorough, well-architected design with deep implementation specificity that is ready for implementation; the main weakness is incomplete persona coverage and a thin Alternatives section that delegates analysis to an external PR discussion.

Feedback: Add a formal 'User Stories' subsection under Motivation covering all four OSAC personas -- Cloud Provider Admin (managing global catalogs/templates with typed references) and Cloud Infrastructure Admin (managing platform-scoped resources like NetworkClass and IP pools) are missing from the current workflows. Bring the alternatives analysis into the document itself rather than pointing to PR #113 -- reviewers should not need to leave the document to understand why URI/ARN format and generic reference messages were rejected. Address the Tenant Onboarding, Documentation, and Installation dimensions from osac-dimensions.md explicitly, even if only to state they are not affected.

Critical (0)

None.

Important (3)

  1. Missing 'User Stories' subsection under Motivation: the template requires formal user stories in 'As a [role], I want to [action] so that I can [goal]' format. The workflow descriptions in the Proposal section are detailed but cover only Tenant User and Tenant Admin personas. Cloud Provider Admin (who manages global catalogs and templates that use full references) and Cloud Infrastructure Admin (who manages platform-scoped resources like NetworkClass, ExternalIPPool, PublicIPPool, HostType) are
  2. Alternatives section is a stub: it names three alternatives (URI/ARN format, generic reference message, do nothing) but says 'See the PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion for details on the URI/ARN trade-off.' The design template and review patterns require at least one real alternative with in-document rationale for rejection. Reviewers should not need to find and read an external PR thread to evaluate the design's choice.
  3. Cross-cutting dimensions gap: Tenant Onboarding (are auto-provisioned resources during tenant creation affected by the new reference format?), Documentation (which user guides, API references, and architecture docs need updating?), and Installation (does the interceptor require any deployment configuration beyond code?) are not addressed or explicitly deferred. Per osac-dimensions.md, silence on a relevant dimension is a gap.

Suggestions (3)

  1. Add a Terminology section defining 'full reference', 'local reference', 'reference message', and 'reference validation interceptor' upfront. The terms are used consistently but a formal section would align with the pattern set by the Networking EP (cited as a quality benchmark in review-patterns.md).
  2. Explicitly describe the Update workflow: the Create workflow is detailed but Update is only implied ('Create/Update methods' mentioned in passing). Walk through what happens when a user updates a ComputeInstance's subnet reference -- does the interceptor re-validate? Are immutable reference fields enforced by the interceptor or the server?
  3. Consider describing how the interceptor interacts with List and Get operations -- the design focuses on Create/Update validation but doesn't state whether List/Get responses include resolved reference details (e.g., display_name from the referenced resource) or just echo back the submitted reference.

Review cost

Model: claude-opus-4-6
Cost: $0.7190
Tokens: 5.3k in / 4.9k out
Cache: 178.5k read
Active time: 1m 56s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/type-safe-resource-references/design.md`:
- Line 264: Update the proto API inventory so InstanceTypeLocalReference is
listed only under instance_type_type.proto, removing it from
compute_instance_type.proto. Keep compute_instance_type.proto’s usage by
importing the owning proto definition, preserving the single-definition rule and
avoiding duplicate generated symbols.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 3c25582f-7e6c-4425-a851-6567a6fb8eed

📥 Commits

Reviewing files that changed from the base of the PR and between fad3fbd and ae39578.

📒 Files selected for processing (1)
  • enhancements/type-safe-resource-references/design.md

Comment thread enhancements/type-safe-resource-references/design.md Outdated
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deeply technical with full proto before/after examples, Go type definitions (ReferenceValidator, ReferenceLookupFunc), standardized error codes, all lifecycle operations covered, and specific failure handling (race conditions, panic recovery, database errors). Open Question 2 on forward triggers vs interceptor locking is significant but properly flagged with both options presented.
Testability 2/2 Strong test plan with concrete scenarios at unit (serialization, interceptor detection for nested/repeated/oneof, validation logic), integration (kind cluster with valid/invalid/cross-tenant/trigger/CEL scenarios), and E2E (full provisioning workflow, error scenario) levels. Per-chunk graduation criteria are measurable. Overall graduation criteria are generic but compensated by per-chunk specificity.
Scope 1/2 Scope boundaries are clear and non-goals are specific. However, no formal User Stories section with 'As a [role]' format — workflow descriptions cover Tenant User and Tenant Admin but miss Cloud Provider Admin and Cloud Infrastructure Admin perspectives. Alternatives section defers to external PR discussion rather than documenting inline. Goals mix user outcomes with implementation details. Several cross-cutting dimensions (documentation, installation, provisioning) not explicitly addressed or d
Architecture 2/2 All core OSAC patterns followed. No new resources so tenant isolation N/A is correctly stated. Interceptor uses pluggable registry pattern with protoreflect. Delivery chunks respect cross-repo dependency ordering. Proto schemas follow conventions. Interceptor chain positioning (after auth + transaction) is well-reasoned. Minor gaps: no formal terminology section, cross-repo impacts could be more explicit about osac-operator and osac-test-infra.

Verdict: A thorough, well-architected design with strong technical depth and testability, held back from a perfect score by missing formal user stories, an effectively empty Alternatives section, and some unaddressed cross-cutting dimensions.

Feedback: Add formal user stories in 'As a [role], I want to [action] so that [goal]' format for all affected personas, especially Cloud Provider Admin (IAM references in Chunk 5) and Cloud Infrastructure Admin (platform-scoped resources like NetworkClass). Document the URI/ARN and generic-reference-message alternatives inline with rejection rationale rather than pointing to the external PR #113 discussion — the design should be self-contained. Explicitly address or defer the documentation, installation, and provisioning dimensions from osac-dimensions.md to close the silence gaps.

Critical (0)

None.

Important (4)

  1. Missing formal User Stories section under Motivation — workflow descriptions cover Tenant User and Tenant Admin but the template requires 'As a [role], I want to [action] so that [goal]' format, and Cloud Provider Admin (IAM) and Cloud Infrastructure Admin (platform-scoped resources) personas lack coverage.
  2. Alternatives section defers entirely to external PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion ('See the PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion for details') rather than documenting at least one real alternative with rationale inline, as required by the template.
  3. Goals mix user-visible outcomes with implementation details — 'Centralize reference existence validation in a single gRPC interceptor' and 'Provide compile-time type safety' describe the approach, not what users gain.
  4. Cross-cutting dimensions from osac-dimensions.md (documentation, installation, provisioning) are not explicitly addressed or deferred — silence on relevant dimensions is a gap per review criteria.

Suggestions (4)

  1. Add a Terminology section (as the Networking EP does) defining 'reference', 'local reference', 'full reference', 'platform-scoped' to set clear vocabulary upfront.
  2. Make cross-repo impacts more explicit in the API Extensions section — mention osac-operator rebuild requirement and osac-test-infra E2E test updates alongside fulfillment-service and osac-ux.
  3. Strengthen overall graduation criteria beyond 'will be defined when targeting a release' — the per-chunk criteria are concrete but the umbrella graduation is generic.
  4. Consider adding a workflow description for the Cloud Infrastructure Admin creating a platform-scoped resource (e.g., NetworkClass) that other tenants reference, to show the full reference resolution path from the provider side.

Review cost

Model: claude-opus-4-6
Cost: $0.4900
Tokens: 5.3k in / 5.5k out
Cache: 221.6k read
Active time: 2m 1s
API calls: 0

Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/type-safe-resource-references/design.md Outdated
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md

@rccrdpccl rccrdpccl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

overall looks good, just a few clarifying questions

Comment thread enhancements/type-safe-resource-references/design.md Outdated
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout: Go interface definitions for the interceptor, complete before/after proto schemas with buf.validate annotations, concrete JSON wire format examples, database trigger path changes, interceptor detection and error aggregation mechanics. All lifecycle operations covered including reverse reference triggers on delete. Risks are specific (proto field reuse, interceptor latency, CEL breakage, cross-chunk deps) with concrete mitigations. Failure handling covers databas
Testability 2/2 Test plan specifies concrete scenarios at all three levels: unit tests for reference message serialization, interceptor field detection (nested, repeated, oneof), validation logic with field path correctness and error aggregation, and per-server validation removal verification. Integration tests with kind cluster cover valid/invalid references, cross-tenant resolution, trigger enforcement (Z0003), and CEL filter path changes. E2E tests describe full provisioning workflow and error scenarios. Per
Scope 1/2 Goals are user-visible outcomes and non-goals are specific with clear rationale. Boundaries are well-defined with no scope creep. However: (1) Missing User Stories section — the template requires formal 'As a [role], I want to...' stories under Motivation covering all relevant personas. Workflow descriptions cover Tenant User and Tenant Admin scenarios but Cloud Provider Admin (manages global catalogs/templates that use cross-tenant references) is not addressed. (2) Alternatives section is hollo
Architecture 2/2 Sound architectural decisions consistent with OSAC patterns. The local vs. full reference distinction is well-motivated with a clear table of 32 fields mapping to reference types. Interceptor placement in the gRPC chain (after Auth and Transaction) is justified with correct reasoning about tenant context and transaction scope. Proto schemas follow standard conventions with buf.validate annotations. Cross-repo impacts are clearly enumerated. Tenant isolation correctly addressed — no new resources

Verdict: A thorough and well-architected design scoring 7/8 — strong in architecture, feasibility, and testability, held back from a perfect score by missing user stories, a hollow alternatives section, and incomplete cross-cutting dimension coverage.

Feedback: Add a User Stories subsection under Motivation with formal 'As a [role], I want to...' stories covering at least Tenant User, Tenant Admin, and Cloud Provider Admin (who manages global catalogs and templates that are cross-tenant referenced). Flesh out the Alternatives section inline — the three alternatives (URI/ARN format, generic reference message, do nothing) are named but their trade-offs and rejection rationale should be explained in the design itself rather than deferring to an external PR discussion. Address or explicitly defer the Installation and Documentation cross-cutting dimensions.

Critical (0)

None.

Important (2)

  1. Missing User Stories section under Motivation: the template requires formal 'As a [role], I want to [action] so that I can [goal]' stories. Workflow descriptions cover Tenant User and Tenant Admin scenarios but are not formatted as user stories, and Cloud Provider Admin (who manages global catalogs/templates that are cross-tenant referenced) is not addressed as a persona.
  2. Hollow Alternatives section: states 'None' and names URI/ARN format, generic reference message, and do-nothing as rejected options but defers all rationale to 'PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion'. The template requires at least one real alternative with inline rationale for rejection.

Suggestions (3)

  1. Add explicit cross-cutting dimension coverage for Installation (even if N/A, state why no osac-installer changes are needed) and Documentation (beyond the passing mention in chunk workflow step 8).
  2. Overall graduation criteria start with 'Graduation criteria will be defined when targeting a release' which reads as a placeholder — the per-chunk criteria are concrete and sufficient but the overall Dev Preview → Tech Preview → GA progression needs measurable conditions.
  3. Open Question 2 (whether to keep or remove forward reference triggers Z0002) affects the data integrity guarantees of the system under concurrent operations — consider resolving before merge rather than deferring to implementation, since the choice between options (a) and (b) has architectural implications for the interceptor's DAO lookup interface.

Review cost

Model: claude-opus-4-6
Cost: $0.7330
Tokens: 5.3k in / 5.3k out
Cache: 178.9k read
Active time: 1m 59s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/type-safe-resource-references/design.md`:
- Around line 595-598: Update the trigger migration and JSON-path matching
design to require tenant and project predicates alongside resource names,
matching the interceptor’s scoping rules. Apply these constraints consistently
to forward triggers, reverse triggers, and their related indexes, and add an
integration test covering same-name resources across tenants to verify they
remain distinct.
- Around line 571-577: Update the interceptor ordering in the documented chain
so Auth runs before Transaction, ensuring authentication completes before
opening a database transaction. Revise the surrounding explanation to state that
the interceptor runs after Auth for tenant context while DAO lookups still share
the request transaction.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: b27b1267-14d7-47ff-8cf0-c7d2d7e70567

📥 Commits

Reviewing files that changed from the base of the PR and between ae39578 and 71c3929.

📒 Files selected for processing (1)
  • enhancements/type-safe-resource-references/design.md

Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
Haim Tayrie added 5 commits July 22, 2026 22:19
…hared

- Add blank line before fenced JSON block (MD031)
- Add concurrent create/delete integration test scenario
- Fix mutation example to show public (shared) vs private (tenant) forms
- Move id-only/both-match tests to full reference field (CatalogItem)

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
Match the required OSAC-<jira-key>-<slug> naming convention.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
Local references now include an id field alongside name, matching the
same resolution modes as full references (name-only, id-only, both with
consistency check). This allows clients currently using resource IDs to
continue working during the transition to name-based references.

Updated: proto examples, CLI section (--field-id flags), interceptor
description, before/after example, and test plan.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
- Add project field to private API mutation example
- Add blank lines before CLI code fences (MD031)
- Add scope context (caller vs shared tenant) to CatalogItem test cases

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
Address CrystalChun's review feedback:
- Add `string project` field to public full references so tenant users
  can scope lookups within a project (empty = tenant-global)
- Add `bool shared` field to private full references per API Guidelines
  (private must be a superset of public)
- Add project-level access disclaimer in Security section
- Update CLI flags, mutation examples, test plan, and resolution logic

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
enhancements/OSAC-1330-type-safe-resource-references/design.md (3)

662-684: 🚀 Performance & Scalability | 🟠 Major | 🏗️ Heavy lift

Bound and batch interceptor validation work.

The interceptor recursively validates every repeated/nested reference and aggregates every failure, but the design specifies no maximum reference count, deduplication, batching, or error cap. Large repeated fields can create N+1 DAO queries and oversized error responses. Add per-request bounds, deduplicate lookups, batch where possible, and honor request deadlines.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-1330-type-safe-resource-references/design.md` around lines
662 - 684, Update the interceptor’s recursive message-walking and
error-aggregation design to enforce a per-request maximum reference count and
error cap, deduplicate identical lookups, batch DAO queries where supported, and
propagate request deadlines through validation. Preserve complete FieldViolation
paths for reported failures while preventing unbounded queries and oversized
responses.

95-103: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Align the overview with the canonical reference schemas.

This section says full references contain id, tenant, project, name and local references contain only name, but the schemas below define public full references as id, name, project, shared, private full references as those fields plus tenant, and local references as id, name. This contradiction can produce incompatible proto implementations.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-1330-type-safe-resource-references/design.md` around lines
95 - 103, Update the overview in the “Per-type reference messages in proto”
section to match the canonical schemas below: describe public full references as
id, name, project, and shared; private full references as those fields plus
tenant; and local references as id and name. Ensure the field-replacement
guidance uses these corrected definitions consistently.

717-721: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Keep resource names immutable when triggers reference them.

The name-based reverse triggers mean deleting a referenced resource by name can either leak stale references or become blocked after the name is reused. Either treat metadata.name as immutable for referenced resources, or add atomic rename propagation plus a delete/reuse policy that preserves referential integrity.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-1330-type-safe-resource-references/design.md` around lines
717 - 721, Define and enforce an immutability rule for resource metadata.name
whenever the resource is referenced by triggers, preventing renames that would
invalidate name-based reverse references. Ensure trigger validation and update
paths preserve the existing name for referenced resources, while leaving
unreferenced resource behavior unchanged.
♻️ Duplicate comments (1)
enhancements/OSAC-1330-type-safe-resource-references/design.md (1)

723-741: 🗄️ Data Integrity & Integration | 🟠 Major

Define the complete uniqueness scope for trigger lookups.

The design now treats project as a lookup dimension and includes project-scoped tests, while same-tenant trigger predicates and indexes include only tenant. If names can repeat across projects, triggers can match the wrong resource. Either include project in predicates/indexes or explicitly guarantee tenant-wide name uniqueness, with a cross-project same-name test.

Based on learnings, trigger indexes must use the complete name-uniqueness scope.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-1330-type-safe-resource-references/design.md` around lines
723 - 741, Clarify the name-uniqueness scope for same-tenant trigger lookups in
the design. If names may repeat across projects, update the predicates and
associated indexes to include project alongside tenant for forward and reverse
triggers; otherwise explicitly guarantee tenant-wide uniqueness and add a
cross-project same-name test. Ensure the documented trigger index scope matches
the complete uniqueness rule.

Source: Learnings

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/OSAC-1330-type-safe-resource-references/design.md`:
- Around line 615-623: Update the interceptor’s reference-resolution flow to
derive local-reference tenant/project from the parent resource’s scope rather
than caller auth context. Preserve explicitly resolved tenant/project values for
full references, and ensure the lookup filter always enforces both scope fields
instead of selecting by id or metadata.name alone.

---

Outside diff comments:
In `@enhancements/OSAC-1330-type-safe-resource-references/design.md`:
- Around line 662-684: Update the interceptor’s recursive message-walking and
error-aggregation design to enforce a per-request maximum reference count and
error cap, deduplicate identical lookups, batch DAO queries where supported, and
propagate request deadlines through validation. Preserve complete FieldViolation
paths for reported failures while preventing unbounded queries and oversized
responses.
- Around line 95-103: Update the overview in the “Per-type reference messages in
proto” section to match the canonical schemas below: describe public full
references as id, name, project, and shared; private full references as those
fields plus tenant; and local references as id and name. Ensure the
field-replacement guidance uses these corrected definitions consistently.
- Around line 717-721: Define and enforce an immutability rule for resource
metadata.name whenever the resource is referenced by triggers, preventing
renames that would invalidate name-based reverse references. Ensure trigger
validation and update paths preserve the existing name for referenced resources,
while leaving unreferenced resource behavior unchanged.

---

Duplicate comments:
In `@enhancements/OSAC-1330-type-safe-resource-references/design.md`:
- Around line 723-741: Clarify the name-uniqueness scope for same-tenant trigger
lookups in the design. If names may repeat across projects, update the
predicates and associated indexes to include project alongside tenant for
forward and reverse triggers; otherwise explicitly guarantee tenant-wide
uniqueness and add a cross-project same-name test. Ensure the documented trigger
index scope matches the complete uniqueness rule.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 4f7db594-da6b-4386-90b8-1d6f78ed6b50

📥 Commits

Reviewing files that changed from the base of the PR and between 83ebcc0 and 7ab5da9.

📒 Files selected for processing (1)
  • enhancements/OSAC-1330-type-safe-resource-references/design.md

Comment thread enhancements/OSAC-1330-type-safe-resource-references/design.md
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-121

Score: 10/10 | Verdict: PASS

Criterion Score Notes
What 2/2 Clear desired outcome: replace 34 opaque string reference fields across 15 resources with typed protobuf messages. All four OSAC personas (Tenant User, Tenant Admin, Cloud Provider Admin, Cloud Infrastructure Admin) have concrete workflow descriptions with JSON examples. Complete reference field inventory maps every affected field with local vs. full reference rationale.
Why 2/2 Three concrete problem classes: no compile-time safety (wrong ID types pass silently), no cross-tenant addressability, scattered/inconsistent validation. Scope quantified (34 fields, 15 resources). User-facing consequences identified: inconsistent error messages, extra UI API calls, inability to reference cross-tenant resources.
How 2/2 Exhaustively detailed: proto schemas with code, Go interceptor interfaces, before/after wire format, interceptor chain placement, DB trigger migration matrix, CEL filter paths, CLI flags, observability metrics with alert thresholds, security analysis, failure handling, and a 5-chunk incremental delivery plan with per-chunk implementation ordering.
Task 2/2 Proper platform feature enhancement introducing new protobuf message types, a gRPC validation interceptor, observability metrics, API wire format changes, and CLI/UI updates. Delivers new platform capabilities, not documentation or content.
Size 2/2 Coherent scope — all changes serve type-safe references and are interdependent. The 5 delivery chunks are dependency-ordered (networking, compute, IP, clusters, IAM), each leaving the system functional. One capability consistently applied across the API surface.

Verdict: A thorough, well-structured design document that clearly articulates the problem (type-unsafe string references), provides a detailed and measurable solution (per-type proto messages + centralized interceptor), and demonstrates strong engineering rigor across security, observability, delivery planning, and testing.

Feedback: The Alternatives section is thin — it defers to a PR discussion rather than summarizing the trade-offs inline, which makes the document less self-contained. The three open questions (status-level references, forward trigger retention, CEL filter communication) should be resolved before implementation begins to avoid mid-stream design pivots. Consider explicitly listing which OSAC services (BMaaS, CaaS, VMaaS) are affected in the Summary for quicker reader orientation.

Critical (0)

None.

Important (2)

  1. The Alternatives section says 'None' and defers to PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion. A design document should summarize the key alternatives considered (URI/ARN format, generic reference message, do nothing) and their rejection rationale inline, not require readers to find a separate PR thread.
  2. Three open questions remain unresolved: (1) whether status-level references should use typed messages, (2) whether forward reference triggers should be kept or removed, (3) how CEL filter changes should be communicated. These affect implementation scope and should be resolved before chunk 1 begins.

Suggestions (3)

  1. The Summary could explicitly name which OSAC services are in scope (all services using the fulfillment API) for quicker reader orientation.
  2. The Graduation Criteria section is mostly placeholder ('will be defined when targeting a release'). Consider defining concrete criteria per delivery chunk since the document already specifies what each chunk delivers.
  3. The document could address the Documentation cross-cutting dimension more explicitly — how API docs, user guides, and the API changelog will be updated alongside each delivery chunk.

Review cost

Model: claude-opus-4-6
Cost: $0.5699
Tokens: 6 in / 6.1k out
Cache: 217.3k read
Active time: 2m 12s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Exceptionally detailed implementation. Full proto schemas with before/after examples, concrete interceptor Go types (ReferenceValidator, ReferenceLookupFunc, ResolvedRef), three resolution modes (name-only, id-only, both) with precise behavior, database trigger path changes with tenant scoping rules, CLI flag patterns, REST/JSON wire format examples. Error codes are specific (InvalidArgument with google.rpc.BadRequest FieldViolation). The race condition between validation and persistence is iden
Testability 2/2 Strong test plan with specific scenarios at all three levels. Unit tests cover interceptor reference detection (nested messages, repeated fields, oneof), resolution modes (name-only auto-populates id, id-only auto-populates name, both-match, both-mismatch), request mutation, and per-server validation removal. Integration tests specify 15 concrete scenarios including cross-tenant references, concurrent create/delete race conditions, project-scoped lookups, and oneof target validation. E2E tests c
Scope 1/2 Summary is clear and well-sized (3 sentences covering what, why, key capabilities). Goals are user-visible outcomes (compile-time safety, centralized validation, standardized errors). Non-goals are specific: explicitly excludes the (tenant, project, name) migration, backward compatibility, internal DB schema changes, and quota/RBAC changes. No scope creep signals. PRD is referenced via frontmatter prd: prd.md. Alternatives section is weak — it says 'None' and defers to 'PRD PR #113 discussion'
Architecture 2/2 Excellent alignment with OSAC patterns. No new resources are introduced, so tenant isolation metadata is correctly noted as not needed. Proto schemas follow the standard object shape conventions (reference messages with id, name, project, shared/tenant fields). The public/private API split follows the API Guidelines (private is superset of public with tenant field). The interceptor chain placement is well-reasoned (after Auth and Transaction, before Handler). The design correctly identifies that

Verdict: A thorough, well-architected design with deep implementation detail and concrete test plans, held back slightly by a missing inline alternatives analysis and thin documentation dimension coverage.

Feedback: The Alternatives section needs at least one real alternative evaluated inline with trade-offs — deferring to 'PRD PR #113 discussion' forces reviewers to leave the document to understand why alternatives were rejected. Present the URI/ARN approach and generic reference message briefly with rejection rationale. Additionally, address the Documentation dimension from osac-dimensions.md: state whether API reference docs in fulfillment-service and osac-docs architecture guides will be updated, or explicitly defer documentation to a later milestone.

Critical (0)

None.

Important (2)

  1. Alternatives section says 'None' and defers to external PR discussion. The rubric requires at least one real alternative with inline rationale for rejection. The URI/ARN format and generic reference message approaches should be briefly described with trade-offs, even if rejected.
  2. Documentation dimension from osac-dimensions.md is not addressed. The design mentions updating API changelog and release notes but does not state whether osac-docs architecture guides, API reference docs in fulfillment-service/docs/, or user-facing documentation will be updated or explicitly deferred.

Suggestions (3)

  1. Open Question 2 (keep vs remove forward triggers) is a significant architectural decision that affects data integrity guarantees. Consider resolving this in the design rather than deferring to implementation — the choice between options (a) and (b) affects the interceptor's DAO lookup contract and testing strategy.
  2. The Graduation Criteria section says 'Graduation criteria will be defined when targeting a release' for Dev Preview/Tech Preview/GA staging, which reads as a placeholder. The per-chunk criteria below it are concrete and sufficient — consider removing the placeholder sentence to avoid the impression of incompleteness.
  3. Consider adding a Terminology section (per review-patterns.md feedback themes) to formally define 'full reference', 'local reference', 'platform-scoped', 'resolution mode' — these terms are used consistently but never formally defined.

Review cost

Model: claude-opus-4-6
Cost: $0.3426
Tokens: 9 in / 2.4k out
Cache: 470.7k read
Active time: 1m 14s
API calls: 0

Comment on lines +391 to +395
string id = 1;
string tenant = 2;
string project = 3;
string name = 4;
bool shared = 5;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: I think the field order matters and maybe should match the public order with any additions being after?

Suggested change
string id = 1;
string tenant = 2;
string project = 3;
string name = 4;
bool shared = 5;
string id = 1;
string name = 2;
string project = 3;
bool shared = 4;
string tenant = 5;

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated — private field order now matches public with tenant appended at the end:

string id = 1;
string name = 2;
string project = 3;
bool shared = 4;
string tenant = 5;

- Private full reference field order now matches public with tenant
  appended: id=1, name=2, project=3, shared=4, tenant=5
- Remove cross-tenant information disclosure paragraph (no longer
  relevant since public API has no tenant field)

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Exceptionally detailed implementation. Concrete Go type definitions for the interceptor (ReferenceValidator, ReferenceLookupFunc, ResolvedRef), full proto schemas with buf.validate CEL rules, per-table database trigger JSON path changes, three resolution modes (name-only, id-only, both) thoroughly specified. Race conditions identified with specific mitigations (SELECT FOR SHARE). Risks are specific (proto field number reuse, interceptor latency, CEL filter breakage) with concrete mitigations. Dr
Testability 2/2 Strong test plan with clear unit/integration/e2e separation. Unit tests specify 6 concrete categories (message construction, reference detection, validation logic, resolution modes, request mutation, per-server removal). Integration tests describe 14 specific scenarios including cross-tenant references, concurrent create/delete race conditions, oneof targets, project-scoped resolution, and both-match/both-mismatch cases. E2E tests cover a full provisioning workflow. Per-chunk graduation criteria
Scope 1/2 Boundaries are mostly clear — non-goals are specific (UUID-to-name migration, backward compatibility, quota/RBAC changes) and well-reasoned. PRD is referenced in frontmatter. However, the Alternatives section is inadequate: it names three alternatives (URI/ARN, generic reference message, do nothing) but defers all rationale to 'the PRD PR #113 discussion' instead of including at least one real comparison with trade-offs in the design itself. Goals are a mix of user-visible ('standardize error re
Architecture 2/2 Follows all OSAC patterns. Correctly identifies that no new resources are introduced, so no new tenant isolation metadata is needed. Proto reference messages follow standard naming and structure conventions. The interceptor chain placement (after auth + transaction, before handler) is architecturally sound — DAO lookups share the request's transaction. Public/private API split correctly mirrors the existing superset pattern. Cross-repo impacts clearly enumerated (fulfillment-service proto + serv

Verdict: A thorough and well-architected design that scores 7/8, held back only by a thin Alternatives section that defers rationale to external discussion rather than presenting trade-offs inline.

Feedback: The Alternatives section needs real content: pick at least the URI/ARN approach and the generic-reference-message approach, describe each in 3-5 sentences, and explain why per-type messages were chosen over them — don't send the reader to a PR discussion. Consider explicitly declaring which services (BMaaS, CaaS, VMaaS) are in scope and mapping the delivery chunks to the cross-cutting dimensions from osac-dimensions.md. The three Open Questions (status-level references, forward trigger retention, CEL filter migration) should each have a stated resolution timeline or decision owner to avoid blocking implementation.

Critical (0)

None.

Important (2)

  1. Alternatives section (line 997-1000) states 'None' and defers rationale to 'PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion.' The design template requires at least one real alternative with rationale for rejection within the design document itself. The three alternatives mentioned (URI/ARN format, generic reference message, do nothing) should each have a brief description and rejection rationale inline.
  2. Open Question 2 (forward reference triggers, line 1017-1020) directly affects the correctness of the race condition mitigation described in Failure Handling. The choice between keeping triggers with updated JSON paths vs. moving FOR SHARE into the interceptor's DAO lookups should be resolved before implementation begins, as it changes the database migration strategy and interceptor implementation.

Suggestions (4)

  1. Goals mix developer-visible outcomes ('compile-time type safety') with user-visible outcomes ('standardize error reporting'). Consider reframing developer-focused goals in terms of user-observable benefits (e.g., 'Prevent cross-type reference assignment errors that currently surface as runtime 500s').
  2. Services in scope are never explicitly declared per osac-dimensions.md conventions. The delivery chunks implicitly cover BMaaS (Chunk 4), CaaS (Chunk 4), VMaaS (Chunk 2), and Core (Chunks 1, 3, 5), but a one-line declaration would help reviewers confirm coverage.
  3. Graduation criteria (line 1115-1118) partially defers: 'Graduation criteria will be defined when targeting a release.' While the per-chunk criteria are concrete, consider adding at least a target milestone (e.g., 'Chunk 1 targets 0.3') to anchor the delivery timeline.
  4. The design references 34 spec-level reference fields across 15 resources (line 61) — consider including this inventory as an appendix or linking to the architectural context document, so reviewers can verify completeness of the field mapping table (lines 456-492).

Review cost

Model: claude-opus-4-6
Cost: $0.7534
Tokens: 7 in / 4.9k out
Cache: 261.0k read
Active time: 2m 8s
API calls: 0

…ontext

Local references now resolve tenant/project from the owning resource's
metadata rather than the caller's auth context. This ensures correct
scoping when a Cloud Provider Admin creates resources in a different
tenant/project via the private API.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Haim Tayrie <htayrie@htayrie-thinkpadt14gen5.raanaii.csb>
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-121

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep implementation detail throughout. Full proto schemas for all three reference message patterns (public full, private full, local) with concrete before/after examples. Go code for the interceptor with specific types (ReferenceValidator, ReferenceLookupFunc, ResolvedRef). All lifecycle operations covered (Create/Update via interceptor, Delete via reverse triggers Z0003, Get/List unaffected). Error codes specified (InvalidArgument for validation failures, Internal for programming errors, NotFou
Testability 2/2 One of the strongest sections. Unit tests (Ginkgo) specify six distinct areas: message construction, reference detection, validation logic, resolution modes, request mutation, and per-server validation removal. Integration tests (kind cluster) enumerate 14 specific scenarios spanning all five delivery chunks, including concurrent create/delete serialization, cross-tenant resolution, oneof variant handling, project-scoped lookups, and id/name mismatch detection. E2E tests (pytest) cover full prov
Scope 1/2 Summary is concise (3 sentences), goals are user-visible outcomes not implementation tasks, non-goals are specific with locked design decision references (D1 for no backward compat, D5 for deferred name migration). No scope creep signals. PRD referenced in frontmatter and summary. All four personas appear in workflow descriptions. Delivery chunks provide clear boundaries. However, the Alternatives section is essentially empty — it names three alternatives (URI/ARN format, generic reference messa
Architecture 2/2 All OSAC patterns followed. Tenant isolation correctly identified as preserved (no new resources introduced, existing metadata unchanged). Public/private API distinction follows API guidelines — private is a superset with explicit tenant field. Reference detection via protoreflect naming convention is sound and self-documenting. Interceptor chain placement is well-reasoned (after Auth for tenant context, after Transaction for DAO access within the request's DB transaction). Spec/status ownership

Verdict: A thorough, well-structured design with deep technical detail across architecture, feasibility, and testability, held back from a perfect score only by an empty Alternatives section that defers trade-off analysis to an external PR discussion.

Feedback: The Alternatives section must document at least one real alternative with rationale within the design itself — naming URI/ARN, generic reference, and do-nothing options but deferring all analysis to 'PRD PR #113 discussion' makes the design not self-contained. Even a brief paragraph per alternative explaining why it was rejected would satisfy the requirement. Additionally, Open Question 2 (whether to keep forward triggers for FOR SHARE locking or move locking into the interceptor) has direct data consistency implications — the design should state the author's recommended option with reasoning, even if the final decision is deferred to implementation.

Critical (0)

None.

Important (2)

  1. Alternatives section is effectively empty: the design names three alternatives (URI/ARN format, generic reference message, do nothing) but says 'See the PRD PR PRD: Type-Safe Resource References (OSAC-1330) #113 discussion for details on the URI/ARN trade-off' instead of documenting trade-offs in the design. A self-contained design should include at least a brief rationale for rejecting each alternative so reviewers can evaluate the decision without chasing external references.
  2. Open Question 2 (forward trigger retention vs. removal) affects data consistency guarantees under concurrent operations. The design correctly identifies the race window under READ COMMITTED but defers the decision entirely. The author should state a recommended option — keeping triggers with updated JSON paths (option a) is the safer default since it preserves the existing locking safety net — and note what would change if the other option is chosen.

Suggestions (3)

  1. Top-level graduation criteria says 'will be defined when targeting a release' — consider adding a brief statement of expected graduation stages even if details are deferred, since the per-chunk criteria are already concrete and measurable.
  2. The Installation dimension from osac-dimensions.md is not explicitly addressed. A brief 'No installer changes required — all changes are within fulfillment-service and osac-ux' would close the gap more clearly than the current 'Infrastructure Needed: None' at the end.
  3. The design could briefly note whether the interceptor's protoreflect-based message walking has any performance characteristics worth monitoring beyond the proposed histogram metric — e.g., whether deeply nested messages (ComputeInstance with NetworkAttachment arrays) could produce non-trivial reflection overhead at scale.

Review cost

Model: claude-opus-4-6
Cost: $0.5211
Tokens: 6 in / 4.8k out
Cache: 212.6k read
Active time: 1m 58s
API calls: 0

@CrystalChun CrystalChun left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/lgtm

@openshift-ci

openshift-ci Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: CrystalChun, htayrie-rh

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-merge-bot
openshift-merge-bot Bot merged commit cfcc904 into osac-project:main Jul 22, 2026
5 checks passed
slintes pushed a commit to slintes/enhancement-proposals that referenced this pull request Jul 27, 2026
PRE_COMMIT_PR_BASE_SHA is a snapshot from the pull_request webhook
payload, captured at the PR's last open/synchronize event. It doesn't
advance as main gains new commits, and re-running an old CI job replays
that same stale payload rather than refreshing it.

This caused a false-positive class of failure: a long-lived PR that
hasn't been pushed to since some other, unrelated PR merged a
still-non-compliant enhancements/ directory into main would fail
check-ep-naming on that unrelated directory, even though the PR never
touches it and it's already correctly grandfathered on main itself.
Observed concretely on PR osac-project#121, which failed on
enhancements/storage-control-plane-osac-2872 (merged by PR osac-project#134) despite
never touching that path.

Fix: grandfathering now also checks the live tip of the base branch
(PRE_COMMIT_LIVE_BASE_REF, e.g. origin/main, fetched fresh at the start
of every CI run) in addition to the stale base SHA — a path is
grandfathered if it exists at either reference. This keeps enforcement
scoped to genuinely new paths, so contributors actively fixing their own
directory's naming are never blocked by an unrelated pre-existing
violation elsewhere in the repo.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Tommy Hughes <tohughes@redhat.com>
tchughesiv added a commit to tchughesiv/enhancement-proposals that referenced this pull request Jul 29, 2026
Full-repo sweep across all 41 retired directory names from the entire
OSAC-2870 effort (not just this PR's 5) turned up two more stale
references that earlier passes missed:

- OSAC-1330-type-safe-resource-references/design.md still linked to
  '/enhancements/networking' (renamed to OSAC-356-networking in osac-project#144).
  This directory was self-renamed by osac-project#121's own branch, so it never
  went through our cross-reference sweep.
- OSAC-985-metering-and-usage-tracking/design.md still linked to
  '/enhancements/vm-instance-types' (renamed to OSAC-46-vm-instance-types
  in osac-project#144). metering-and-usage-tracking was one of the directories
  deferred at that time due to an open PR, so it was excluded from
  that pass's cross-reference sweep and the reference went stale once
  the deferred rename landed in osac-project#149.

No open PRs conflict with either file (re-verified).

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Tommy Hughes <tohughes@redhat.com>
empovit pushed a commit to empovit/osac-enhancement-proposals that referenced this pull request Aug 2, 2026
…ure keys

Renames caas-cluster-storage, cluster-version-api, disk-image-OSAC-2540,
and secret-management to the OSAC-2836 naming convention
(enhancements/OSAC-NNNN-slug/), using each EP's confirmed Jira Feature
key. Also lowercases cluster-version-api's DESIGN.md to design.md, and
corrects caas-cluster-storage/design.md's stale tracking-link (was
pointing at the parent Epic OSAC-1123 instead of the Feature OSAC-1332
that prd.md already cited).

cluster-and-vm-provisioning-wizard, metering-and-usage-tracking, and
type-safe-resource-references are deferred to a fast-follow PR pending
enhancement-proposals#108, osac-project#131, and osac-project#121.

Signed-off-by: Tommy Hughes <tohughes@redhat.com>
empovit pushed a commit to empovit/osac-enhancement-proposals that referenced this pull request Aug 2, 2026
PRE_COMMIT_PR_BASE_SHA is a snapshot from the pull_request webhook
payload, captured at the PR's last open/synchronize event. It doesn't
advance as main gains new commits, and re-running an old CI job replays
that same stale payload rather than refreshing it.

This caused a false-positive class of failure: a long-lived PR that
hasn't been pushed to since some other, unrelated PR merged a
still-non-compliant enhancements/ directory into main would fail
check-ep-naming on that unrelated directory, even though the PR never
touches it and it's already correctly grandfathered on main itself.
Observed concretely on PR osac-project#121, which failed on
enhancements/storage-control-plane-osac-2872 (merged by PR osac-project#134) despite
never touching that path.

Fix: grandfathering now also checks the live tip of the base branch
(PRE_COMMIT_LIVE_BASE_REF, e.g. origin/main, fetched fresh at the start
of every CI run) in addition to the stale base SHA — a path is
grandfathered if it exists at either reference. This keeps enforcement
scoped to genuinely new paths, so contributors actively fixing their own
directory's naming are never blocked by an unrelated pre-existing
violation elsewhere in the repo.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Tommy Hughes <tohughes@redhat.com>
empovit pushed a commit to empovit/osac-enhancement-proposals that referenced this pull request Aug 2, 2026
Full-repo sweep across all 41 retired directory names from the entire
OSAC-2870 effort (not just this PR's 5) turned up two more stale
references that earlier passes missed:

- OSAC-1330-type-safe-resource-references/design.md still linked to
  '/enhancements/networking' (renamed to OSAC-356-networking in osac-project#144).
  This directory was self-renamed by osac-project#121's own branch, so it never
  went through our cross-reference sweep.
- OSAC-985-metering-and-usage-tracking/design.md still linked to
  '/enhancements/vm-instance-types' (renamed to OSAC-46-vm-instance-types
  in osac-project#144). metering-and-usage-tracking was one of the directories
  deferred at that time due to an open PR, so it was excluded from
  that pass's cross-reference sweep and the reference went stale once
  the deferred rename landed in osac-project#149.

No open PRs conflict with either file (re-verified).

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Tommy Hughes <tohughes@redhat.com>
@CrystalChun

CrystalChun commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

/retitle OSAC-2766: Design - Type-Safe Resource References

@openshift-ci openshift-ci Bot changed the title OSAC-1330: Design - Type-Safe Resource References OSAC-2766: Design - Type-Safe Resource References Aug 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants