PRD: Metering and Usage Tracking - #78
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughA new PRD defines metering and usage tracking for OSAC, covering VMaaS, CaaS, and MaaS. It adds terminology, scope, capabilities, operational requirements, acceptance criteria, assumptions, risks, open questions, and pricing formulas. ChangesMetering and Usage Tracking PRD
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~5 minutes Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 11✅ Passed checks (11 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Around line 92-93: Update the CAP-13 acceptance criteria in the
metering-and-usage-tracking PRD to explicitly cover hosted-control-plane
attribution, not just worker double-counting. In the CAP-13 section, add a
testable requirement that when clusters are nested via Hosted Control Plans, the
hosted control plane infrastructure is attributed to the hosted clusters and is
not billed independently alongside the parent cluster’s resources. Use the
existing CAP-13 wording as the anchor and expand the acceptance set so it
verifies both worker and hosted-control-plane resource accounting.
- Around line 46-49: Milestone 0.3 scope is inconsistent because it still
implies quota-related balance updates within 60 seconds even though quota
enforcement is deferred to 0.4. Update the PRD wording in the milestone 0.3 and
related sections to clearly exclude quota enforcement and any budget/quota
balance update requirement from 0.3, and move that requirement to the milestone
0.4 language tied to CAP-19 and the MaaS acceptance criteria.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: aa0f02cf-39a7-40f4-8353-2029d077dc5f
📒 Files selected for processing (1)
enhancements/metering-and-usage-tracking/prd.md
mhrivnak
left a comment
There was a problem hiding this comment.
In general, let's keep timelines out of PRDs. See the recent slack discussion about PRD process.
WDYT about getting more specific about each meter and why it's needed, down to the detail of the Use Case document I shared some time ago? What are we measuring, why, and in what units...
- Remove milestone references and roadmap section per reviewer guidance - Remove Enclave from services table - Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time) - Generalize Host type to Resource class - Remove project from Cloud Provider Admin view (CAP-1) - Add resource attribution to parent cluster (CAP-12) - Convert metering failure behavior to open question - Remove design/implementation open questions - Remove Cross-Cutting Dimensions and Academic sections - Tighten acceptance criteria Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Moti Asayag <masayag@redhat.com>
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Around line 64-68: The capability IDs in the metering PRD are out of sequence,
with CAP-17 appearing between CAP-3 and CAP-4. Update the affected entries in
the capability list so the numbering is sequential and easier to reference,
adjusting the nearby CAP labels consistently rather than leaving a gap in the
series.
- Line 87: CAP-12’s cluster attribution requirement is not reflected in the
acceptance criteria, so update the PRD acceptance section to add an explicit
criterion that validates a unified cluster-level cost query using the parent
cluster attribution. Use the existing CaaS acceptance criteria around control
plane and worker node metering, and add a check that these attributed resources
roll up into one cluster view without double-counting nested resources
(including hosted-control-plane cases). Keep the wording aligned with CAP-12 and
the existing CaaS criteria so the new criterion is easy to locate and verify.
- Line 45: Align the PRD so “quota enforcement” is treated consistently across
the non-goals and requirements: if it remains deferred, remove or rephrase
CAP-19 and the MaaS acceptance criterion that mention updating budget/quota
balances for subsequent request evaluation. Update the affected sections
(non-goal §2.2, CAP-19, and the MaaS acceptance criterion) to describe only
metering/usage emission and avoid implying the metering system performs quota
balance updates or enforcement.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: 5bad52d8-383d-4f55-a0ab-a39a6e0fc698
📒 Files selected for processing (2)
enhancements/caas-cluster-storage/prd.mdenhancements/metering-and-usage-tracking/prd.md
- Remove milestone references and roadmap section per reviewer guidance - Remove Enclave from services table - Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time) - Generalize Host type to Resource class - Remove project from Cloud Provider Admin view (CAP-1) - Add resource attribution to parent cluster (CAP-12) - Convert metering failure behavior to open question - Remove design/implementation open questions - Remove Cross-Cutting Dimensions and Academic sections - Tighten acceptance criteria Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Moti Asayag <masayag@redhat.com>
There was a problem hiding this comment.
♻️ Duplicate comments (1)
enhancements/metering-and-usage-tracking/prd.md (1)
118-118: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winAdd "when available" qualifier to align with deferred quota enforcement.
The acceptance criterion states the MaaS latency requirement is "so that downstream systems can evaluate against near-real-time balances," but omits the "(when available)" qualifier that CAP-19 uses. Since quota enforcement is explicitly deferred to a separate PRD (§2.2), the acceptance criterion should match CAP-19's wording to avoid implying that quota evaluation is currently required.
- and processed within 60 seconds so that downstream systems can evaluate against near-real-time balances + and processed within 60 seconds so that downstream systems (e.g., quota enforcement, when available) can evaluate against near-real-time balances🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@enhancements/metering-and-usage-tracking/prd.md` at line 118, Update the acceptance criterion in the metering-and-usage-tracking PRD to match CAP-19 wording by adding the “when available” qualifier to the near-real-time balances statement. Keep the criterion aligned with the deferred quota enforcement scope by editing the metering event latency requirement text in the affected bullet only, so it no longer implies quota evaluation is currently required.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Duplicate comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Line 118: Update the acceptance criterion in the metering-and-usage-tracking
PRD to match CAP-19 wording by adding the “when available” qualifier to the
near-real-time balances statement. Keep the criterion aligned with the deferred
quota enforcement scope by editing the metering event latency requirement text
in the affected bullet only, so it no longer implies quota evaluation is
currently required.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: 22c79e2e-256f-4fc5-b35f-119b9fc9f93b
📒 Files selected for processing (1)
enhancements/metering-and-usage-tracking/prd.md
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (4)
enhancements/metering-and-usage-tracking/prd.md (4)
116-116: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueConsider explicitly including cached tokens in the MaaS acceptance criteria.
CAP-17 explicitly lists cached tokens as a token type, and the charge model shows cached tokens with a discounted rate. The acceptance criterion at Line 116 mentions "input tokens, output tokens, and total tokens" — since total tokens encompasses all types, this is technically covered, but explicitly listing "cached tokens" would improve testability and align with CAP-17's granularity.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@enhancements/metering-and-usage-tracking/prd.md` at line 116, The acceptance criteria in the metering PRD should explicitly include cached tokens alongside input, output, and total tokens. Update the requirement tied to the inference request usage data in the metering-and-usage-tracking PRD to mention cached tokens by name, using the same terminology as CAP-17 and the charge model, so the criteria stay aligned and easier to test.
184-184: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueAlign VMaaS meter name with acceptance criteria terminology.
The acceptance criteria (Line 104) refer to "instance-type-seconds" for flat-rate pricing, but the charge model uses "vm uptime" as the meter name. Consider using "instance-type-seconds" or "vm uptime (instance-type-seconds)" to maintain terminological consistency between requirements and examples.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@enhancements/metering-and-usage-tracking/prd.md` at line 184, The flat per-instance-type example uses “vm uptime” as the meter name, which is inconsistent with the acceptance criteria terminology. Update the charge model entry in the PRD example to use “instance-type-seconds” or “vm uptime (instance-type-seconds)” so it matches the wording used in the acceptance criteria and keeps the terminology consistent across the document.
64-69: 📐 Maintainability & Code Quality | 🔵 TrivialRenumber CAP capabilities sequentially.
CAP-17 appears between CAP-3 and CAP-4, breaking numerical sequence. This reduces readability and can cause confusion when referencing capabilities. Renumber MaaS-related capabilities to maintain sequential order (e.g., CAP-4 through CAP-7 for the current set, with MaaS capabilities renumbered to fit the sequence).
Based on past review feedback, this ordering issue was previously flagged but remains unresolved.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@enhancements/metering-and-usage-tracking/prd.md` around lines 64 - 69, Renumber the capability list in the metering PRD so the CAP identifiers stay sequential and readable; the out-of-order CAP-17 in the capabilities section should be renumbered to follow CAP-3 and before CAP-4, and any related MaaS entries should be adjusted consistently across the same block. Update the numbered labels in the capability bullets only, keeping the meaning unchanged, and ensure the sequence in this section is contiguous when referenced from the metering-and-usage-tracking document.
36-36: 📐 Maintainability & Code Quality | 🔵 TrivialClarify the relationship between metering data and deferred quota enforcement.
Line 36 states that Cloud Provider Admins can query usage "for billing and quota enforcement," while Line 45 lists "quota enforcement" as a non-goal deferred to a separate PRD. The intent appears to be that metering provides data that enables quota enforcement, while the actual quota system is deferred. Consider rewording Line 36 to clarify this distinction — e.g., "query aggregated usage data per tenant as input for billing and quota enforcement systems" — to prevent reader confusion about whether quota enforcement is in scope.
Also applies to: 45-45
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@enhancements/metering-and-usage-tracking/prd.md` at line 36, Clarify the scope mismatch in the PRD by updating the metering requirement in the metering-and-usage-tracking document so Cloud Provider Admins can query aggregated usage data per tenant as input to billing and quota enforcement systems, while keeping the actual quota enforcement work deferred in the non-goals section; update the wording around the affected requirement and the quota-enforcement non-goal so the metering feature is clearly described as providing data, not implementing enforcement.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Line 116: The acceptance criteria in the metering PRD should explicitly
include cached tokens alongside input, output, and total tokens. Update the
requirement tied to the inference request usage data in the
metering-and-usage-tracking PRD to mention cached tokens by name, using the same
terminology as CAP-17 and the charge model, so the criteria stay aligned and
easier to test.
- Line 184: The flat per-instance-type example uses “vm uptime” as the meter
name, which is inconsistent with the acceptance criteria terminology. Update the
charge model entry in the PRD example to use “instance-type-seconds” or “vm
uptime (instance-type-seconds)” so it matches the wording used in the acceptance
criteria and keeps the terminology consistent across the document.
- Around line 64-69: Renumber the capability list in the metering PRD so the CAP
identifiers stay sequential and readable; the out-of-order CAP-17 in the
capabilities section should be renumbered to follow CAP-3 and before CAP-4, and
any related MaaS entries should be adjusted consistently across the same block.
Update the numbered labels in the capability bullets only, keeping the meaning
unchanged, and ensure the sequence in this section is contiguous when referenced
from the metering-and-usage-tracking document.
- Line 36: Clarify the scope mismatch in the PRD by updating the metering
requirement in the metering-and-usage-tracking document so Cloud Provider Admins
can query aggregated usage data per tenant as input to billing and quota
enforcement systems, while keeping the actual quota enforcement work deferred in
the non-goals section; update the wording around the affected requirement and
the quota-enforcement non-goal so the metering feature is clearly described as
providing data, not implementing enforcement.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml
Review profile: ASSERTIVE
Plan: Enterprise
Run ID: 27deda36-9408-4f81-8bf0-8718a92b0bbe
📒 Files selected for processing (1)
enhancements/metering-and-usage-tracking/prd.md
pgarciaq
left a comment
There was a problem hiding this comment.
A few observations. I think it's important that we align on terminology too.
I see a reference to FOCUS in the glossary but then it seems there's no reference to FOCUS anywhere ¿?
- Remove milestone references and roadmap section per reviewer guidance - Remove Enclave from services table - Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time) - Generalize Host type to Resource class - Remove project from Cloud Provider Admin view (CAP-1) - Add resource attribution to parent cluster (CAP-12) - Convert metering failure behavior to open question - Remove design/implementation open questions - Remove Cross-Cutting Dimensions and Academic sections - Tighten acceptance criteria Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Moti Asayag <masayag@redhat.com>
Updates — addressing mhrivnak's second CHANGES_REQUESTED (2026-07-06)Glossary — FOCUS alignment (r3530894536, r3530930037)Replaced non-standard terminology with FOCUS standard definitions:
Problem Statement — cost vs price (r3530930037)Reworked the second paragraph to be precise:
OQ 9.3 — metering unavailability (r3531110178)Added option (4) per mhrivnak's suggestion: emit events into a durable message bus that guarantees eventual delivery, decoupling OSAC from metering system availability entirely. Events buffer in the bus and are processed when the metering system recovers. CAP-5 — tooling prescription removed (r3494354429)Removed "using standard Kubernetes tooling" from CAP-5 — the phrase prescribes the deployment mechanism (a design decision) and unnecessarily constrains scope in a PRD. CAP numberingRenumbered all capabilities sequentially (CAP-1 through CAP-16) after removing the configurable meter enable/disable capability, and updated all cross-references in the open questions section. |
AI EP Review: EP-78Score: 9/10 | Verdict: PASS
Verdict: A strong, well-organized PRD with clear user-observable capabilities, concrete business justification, and specific testable acceptance criteria — held back slightly by bundling three separable service-specific metering implementations (VMaaS, CaaS, MaaS) that could ship independently. Feedback: Consider splitting MaaS metering into its own PRD — it has a fundamentally different model (per-token vs per-time), unique latency requirements, and a different data source, making it independently deliverable and prioritizable. Rewrite CAP-13's latency requirement in user-observable terms: instead of 'emitted within 30 seconds... processed within 60 seconds' (internal pipeline stages), state 'MaaS usage data is queryable within 90 seconds of an inference request completing.' Similarly, rephrase §7's 'Durable event pipeline' dependency in terms of what it provides to operators rather than naming internal architecture. Critical (0)None. Important (3)
Suggestions (3)
Review costModel: claude-opus-4-6 |
AI Design Review: EP-78Score: 6/8 | Verdict: PASS
Verdict: A well-structured PRD with strong requirements definition and disciplined scoping, but it reads as a pure requirements document missing the template's expected architectural proposal, implementation details, and test plan sections. Feedback: The PRD excels at defining what to build but needs a Proposal section covering how: add component architecture (event pipeline, aggregation layer, query API), define the event schema and transport mechanism, and sketch the API surface for usage queries. Add a Test Plan section with at least a high-level strategy covering integration testing of the event pipeline, load testing for concurrent event ingestion, and end-to-end validation of metering accuracy across VM lifecycle transitions. The open question in section 9.3 (metering system unavailability) is architecturally critical and should be resolved before implementation begins, as the choice between blocking provisioning, accepting gaps, reconciliation, or durable messaging fundamentally shapes the system's reliability guarantees and component topology. Critical (2)
Important (3)
Suggestions (3)
Review costModel: claude-opus-4-6 |
Add CAP-17 requiring that billing systems can determine the originating catalog offer and its bundled components for any metered resource, enabling charge decomposition into independently priceable layers. Extends the Service glossary definition to describe bundled components. Adds acceptance criteria for VMaaS (catalog item, instance type, OS image), CaaS (catalog item, cluster version, host type), and cross-cutting (full priceable component identification). Assisted-by: Claude Code <noreply@anthropic.com> Co-Authored-By: Moti Asayag <masayag@redhat.com> Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: Moti Asayag <masayag@redhat.com>
|
Added CAP-17 — catalog item traceability. Ensures billing systems can determine the originating catalog offer and its bundled components (compute, OS entitlement, boot storage for VMaaS; control plane, workers, cluster version for CaaS) for any metered resource, so charges can be decomposed into independently priceable layers using the provider's rate schedule. Also added corresponding acceptance criteria for VMaaS, CaaS, and cross-cutting. |
AI EP Review: EP-78Score: 10/10 | Verdict: PASS
Verdict: Exceptionally well-structured PRD with clear persona-driven capabilities, concrete business justification, measurable requirements, and well-bounded scope that properly defers billing and costing to separate work. Feedback: This PRD is ready for design. Minor refinements to consider: CAP-6 and CAP-14 use slightly implementation-leaning language ('emit metering events,' 'lifecycle events') — reframing as user outcomes ('track consumption of custom services,' 'integrate with external metering systems') would sharpen the user focus. The open questions (§9.1–9.6) are well-framed but §9.3 lists four implementation options — consider focusing that question on the user-observable behavior (should provisioning block or proceed?) rather than the mechanism. Critical (0)None. Important (2)
Suggestions (3)
Review costModel: claude-opus-4-6 |
Replace open questions (§9) with five decisions: - D-1: Tenant Users see project-scoped usage via RBAC - D-2: Metering failures must not affect provisioning - D-3: Catalog item traceability resolves tenant-defined Services - D-4: VMaaS/CaaS compute remains consumption-based only - D-5: Failed-state resources are not metered Remove design leakage from CAP-13 (pipeline emission/processing language), CAP-14 (lifecycle event emission), and D-2 (durable message bus). Make CAP-6 concrete (configuration-based meter registration). Fix stale §9.6 reference in CAP-11 to D-5. Number Charge Calculation Model as §10. Add MaaS catalog traceability acceptance criterion. Assisted-by: Claude Code <noreply@anthropic.com> Co-Authored-By: Moti Asayag <masayag@redhat.com> Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: Moti Asayag <masayag@redhat.com>
|
Resolved all open questions as decisions (§9):
Also: removed design leakage from CAP-13/CAP-14, made CAP-6 concrete (configuration-based meter registration), added MaaS catalog traceability acceptance criterion, numbered Charge Calculation Model as §10. |
AI EP Review: EP-78Score: 9/10 | Verdict: PASS
Verdict: A well-structured, user-focused PRD with strong persona coverage, concrete business justification, and testable acceptance criteria, held back slightly by bundling three independently shippable service meters into one document. Feedback: Consider whether VMaaS/CaaS (time-based) and MaaS (token-based with stricter latency requirements) should be separate PRDs, since they have different metering models and could be prioritized and delivered independently. The custom meter registration capability (CAP-6) could also stand alone. If keeping them bundled, add explicit phasing guidance in the PRD to clarify which services must ship together vs. which can be incrementally added. Critical (0)None. Important (2)
Suggestions (3)
Review costModel: claude-opus-4-6 |
|
/lgtm |
|
/approve |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: avishayt, masayag, ronniel1 The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
PRD: Metering and Usage Tracking
Jira: OSAC-985
Epic: OSAC-65
Milestone: 0.3
Summary
This PRD defines the metering requirements for OSAC — tracking consumption of VMaaS, CaaS, and MaaS resources with per-second granularity, per-tenant isolation, and per-project aggregation. It establishes the foundation for a sovereign cloud costing and billing stack in milestone 0.4.
Services in Scope
Requesting Review On
How to Review
Summary by CodeRabbit