Skip to content

PRD: Metering and Usage Tracking - #78

Merged
openshift-merge-bot[bot] merged 25 commits into
osac-project:mainfrom
masayag:prd/OSAC-985
Jul 14, 2026
Merged

openshift-merge-bot[bot] merged 25 commits into
osac-project:mainfrom
masayag:prd/OSAC-985

Conversation

@masayag

@masayag masayag commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

PRD: Metering and Usage Tracking

Jira: OSAC-985
Epic: OSAC-65
Milestone: 0.3

Summary

This PRD defines the metering requirements for OSAC — tracking consumption of VMaaS, CaaS, and MaaS resources with per-second granularity, per-tenant isolation, and per-project aggregation. It establishes the foundation for a sovereign cloud costing and billing stack in milestone 0.4.

Services in Scope

  • VMaaS — consumption-based (running VMs only)
  • CaaS — consumption-based (active clusters, per host type)
  • MaaS — consumption-based (per token, per request, 30s/60s latency for quota enforcement)

Requesting Review On

  • Capabilities completeness — do these cover all provider and tenant needs?
  • Consumption vs allocation model — VMaaS/CaaS are consumption-based; BMaaS is allocation-based (deferred). Is this the right default?
  • MaaS data source ownership — should OSAC or RHOAI emit MaaS metering events? (Open Question §9.5)
  • Nested cluster cost attribution (CAP-13) — is the deduplication requirement clear?
  • Future phase roadmap (§10) — does the costing/billing vision align with expectations?
  • Open questions — these need your input

How to Review

  • Comment inline on specific sections
  • Review the open questions (§9) — these need stakeholder input
  • Check that the charge calculation model (VMaaS, CaaS, MaaS) makes sense for your use case
  • Approve when the PRD accurately reflects the agreed requirements

Summary by CodeRabbit

  • Documentation
    • Added a “Metering and Usage Tracking” PRD covering billing-grade metering and consumption semantics for VMaaS, CaaS, and MaaS.
    • Introduced core terminology, glossary, and the Resource class/Template pricing grouping model, including the FOCUS interchange standard.
    • Defined requirements for per-second granularity, meter enable/disable behavior, duplicate-event deduplication, child-to-parent attribution, retention minimums, query/aggregation acceptance criteria, and upgrade safety.
    • Included example charge calculation models and scoped service roadmap (with BMaaS deferred).

@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

A new PRD defines metering and usage tracking for OSAC, covering VMaaS, CaaS, and MaaS. It adds terminology, scope, capabilities, operational requirements, acceptance criteria, assumptions, risks, open questions, and pricing formulas.

Changes

Metering and Usage Tracking PRD

Layer / File(s) Summary
Metadata, glossary, problem statement, and scope
enhancements/metering-and-usage-tracking/prd.md
Adds document metadata, metering terminology, the problem statement, and goals/non-goals scoping VMaaS/CaaS/MaaS with BMaaS deferred.
Capabilities, operational expectations, and acceptance criteria
enhancements/metering-and-usage-tracking/prd.md
Defines per-role metering and costing capabilities, cross-cutting meter behavior and latency requirements, operational retention/scalability targets, and acceptance criteria for VMaaS, CaaS, MaaS, and shared behavior.
Assumptions, risks, open questions, and charge model
enhancements/metering-and-usage-tracking/prd.md
Records deployment assumptions, dependencies, risks, open questions, and example pricing formulas for VMaaS, CaaS, and MaaS.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~5 minutes

Suggested reviewers

  • mhrivnak
  • rgolangh
  • avishayt

Poem

📘 A PRD takes shape in careful prose,
Metered time and usage in steady rows.
VM, CaaS, and MaaS align,
With pricing notes and questions in line.

🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely matches the PRD’s main subject: metering and usage tracking.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
No-Hardcoded-Secrets ✅ Passed The only changed file is a PRD markdown doc, and it contains no hardcoded secrets, credentialed URLs, or secret-like literals.
No-Weak-Crypto ✅ Passed Docs-only PRD; exact searches found no MD5/SHA1/DES/RC4/3DES/Blowfish/ECB or secret-comparison code, and the repo has no code files.
No-Injection-Vectors ✅ Passed Changed files are docs/config only; no shell=True, yaml.load, eval/exec, os.system, pickle.loads, or dangerouslySetInnerHTML in the touched files.
Container-Privileges ✅ Passed PR adds only a markdown PRD; repo-wide scan found no manifests using privileged/hostPID/hostNetwork/hostIPC/SYS_ADMIN/allowPrivilegeEscalation.
No-Sensitive-Data-In-Logs ✅ Passed PR only changes documentation; no logging code or log examples exposing secrets, PII, hostnames, or customer data were added.
Ai-Attribution ✅ Passed HEAD includes an Assisted-by trailer for Claude Code, and no Co-Authored-By trailers were found.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@masayag
masayag marked this pull request as ready for review June 25, 2026 14:23
@openshift-ci
openshift-ci Bot requested review from maorfr and omer-vishlitzky June 25, 2026 14:23

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Around line 92-93: Update the CAP-13 acceptance criteria in the
metering-and-usage-tracking PRD to explicitly cover hosted-control-plane
attribution, not just worker double-counting. In the CAP-13 section, add a
testable requirement that when clusters are nested via Hosted Control Plans, the
hosted control plane infrastructure is attributed to the hosted clusters and is
not billed independently alongside the parent cluster’s resources. Use the
existing CAP-13 wording as the anchor and expand the acceptance set so it
verifies both worker and hosted-control-plane resource accounting.
- Around line 46-49: Milestone 0.3 scope is inconsistent because it still
implies quota-related balance updates within 60 seconds even though quota
enforcement is deferred to 0.4. Update the PRD wording in the milestone 0.3 and
related sections to clearly exclude quota enforcement and any budget/quota
balance update requirement from 0.3, and move that requirement to the milestone
0.4 language tied to CAP-19 and the MaaS acceptance criteria.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: aa0f02cf-39a7-40f4-8353-2029d077dc5f

📥 Commits

Reviewing files that changed from the base of the PR and between c2c59cd and d68aa85.

📒 Files selected for processing (1)
  • enhancements/metering-and-usage-tracking/prd.md

Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md

@mhrivnak mhrivnak left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In general, let's keep timelines out of PRDs. See the recent slack discussion about PRD process.

WDYT about getting more specific about each meter and why it's needed, down to the detail of the Use Case document I shared some time ago? What are we measuring, why, and in what units...

Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
masayag added a commit to masayag/enhancement-proposals that referenced this pull request Jun 28, 2026
- Remove milestone references and roadmap section per reviewer guidance
- Remove Enclave from services table
- Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time)
- Generalize Host type to Resource class
- Remove project from Cloud Provider Admin view (CAP-1)
- Add resource attribution to parent cluster (CAP-12)
- Convert metering failure behavior to open question
- Remove design/implementation open questions
- Remove Cross-Cutting Dimensions and Academic sections
- Tighten acceptance criteria

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Moti Asayag <masayag@redhat.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Around line 64-68: The capability IDs in the metering PRD are out of sequence,
with CAP-17 appearing between CAP-3 and CAP-4. Update the affected entries in
the capability list so the numbering is sequential and easier to reference,
adjusting the nearby CAP labels consistently rather than leaving a gap in the
series.
- Line 87: CAP-12’s cluster attribution requirement is not reflected in the
acceptance criteria, so update the PRD acceptance section to add an explicit
criterion that validates a unified cluster-level cost query using the parent
cluster attribution. Use the existing CaaS acceptance criteria around control
plane and worker node metering, and add a check that these attributed resources
roll up into one cluster view without double-counting nested resources
(including hosted-control-plane cases). Keep the wording aligned with CAP-12 and
the existing CaaS criteria so the new criterion is easy to locate and verify.
- Line 45: Align the PRD so “quota enforcement” is treated consistently across
the non-goals and requirements: if it remains deferred, remove or rephrase
CAP-19 and the MaaS acceptance criterion that mention updating budget/quota
balances for subsequent request evaluation. Update the affected sections
(non-goal §2.2, CAP-19, and the MaaS acceptance criterion) to describe only
metering/usage emission and avoid implying the metering system performs quota
balance updates or enforcement.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 5bad52d8-383d-4f55-a0ab-a39a6e0fc698

📥 Commits

Reviewing files that changed from the base of the PR and between d68aa85 and f82302f.

📒 Files selected for processing (2)
  • enhancements/caas-cluster-storage/prd.md
  • enhancements/metering-and-usage-tracking/prd.md

Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
masayag added a commit to masayag/enhancement-proposals that referenced this pull request Jun 28, 2026
- Remove milestone references and roadmap section per reviewer guidance
- Remove Enclave from services table
- Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time)
- Generalize Host type to Resource class
- Remove project from Cloud Provider Admin view (CAP-1)
- Add resource attribution to parent cluster (CAP-12)
- Convert metering failure behavior to open question
- Remove design/implementation open questions
- Remove Cross-Cutting Dimensions and Academic sections
- Tighten acceptance criteria

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Moti Asayag <masayag@redhat.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
enhancements/metering-and-usage-tracking/prd.md (1)

118-118: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Add "when available" qualifier to align with deferred quota enforcement.

The acceptance criterion states the MaaS latency requirement is "so that downstream systems can evaluate against near-real-time balances," but omits the "(when available)" qualifier that CAP-19 uses. Since quota enforcement is explicitly deferred to a separate PRD (§2.2), the acceptance criterion should match CAP-19's wording to avoid implying that quota evaluation is currently required.

- and processed within 60 seconds so that downstream systems can evaluate against near-real-time balances
+ and processed within 60 seconds so that downstream systems (e.g., quota enforcement, when available) can evaluate against near-real-time balances
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/metering-and-usage-tracking/prd.md` at line 118, Update the
acceptance criterion in the metering-and-usage-tracking PRD to match CAP-19
wording by adding the “when available” qualifier to the near-real-time balances
statement. Keep the criterion aligned with the deferred quota enforcement scope
by editing the metering event latency requirement text in the affected bullet
only, so it no longer implies quota evaluation is currently required.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Duplicate comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Line 118: Update the acceptance criterion in the metering-and-usage-tracking
PRD to match CAP-19 wording by adding the “when available” qualifier to the
near-real-time balances statement. Keep the criterion aligned with the deferred
quota enforcement scope by editing the metering event latency requirement text
in the affected bullet only, so it no longer implies quota evaluation is
currently required.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 22c79e2e-256f-4fc5-b35f-119b9fc9f93b

📥 Commits

Reviewing files that changed from the base of the PR and between f82302f and 7c7f06b.

📒 Files selected for processing (1)
  • enhancements/metering-and-usage-tracking/prd.md

Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (4)
enhancements/metering-and-usage-tracking/prd.md (4)

116-116: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Consider explicitly including cached tokens in the MaaS acceptance criteria.

CAP-17 explicitly lists cached tokens as a token type, and the charge model shows cached tokens with a discounted rate. The acceptance criterion at Line 116 mentions "input tokens, output tokens, and total tokens" — since total tokens encompasses all types, this is technically covered, but explicitly listing "cached tokens" would improve testability and align with CAP-17's granularity.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/metering-and-usage-tracking/prd.md` at line 116, The acceptance
criteria in the metering PRD should explicitly include cached tokens alongside
input, output, and total tokens. Update the requirement tied to the inference
request usage data in the metering-and-usage-tracking PRD to mention cached
tokens by name, using the same terminology as CAP-17 and the charge model, so
the criteria stay aligned and easier to test.

184-184: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Align VMaaS meter name with acceptance criteria terminology.

The acceptance criteria (Line 104) refer to "instance-type-seconds" for flat-rate pricing, but the charge model uses "vm uptime" as the meter name. Consider using "instance-type-seconds" or "vm uptime (instance-type-seconds)" to maintain terminological consistency between requirements and examples.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/metering-and-usage-tracking/prd.md` at line 184, The flat
per-instance-type example uses “vm uptime” as the meter name, which is
inconsistent with the acceptance criteria terminology. Update the charge model
entry in the PRD example to use “instance-type-seconds” or “vm uptime
(instance-type-seconds)” so it matches the wording used in the acceptance
criteria and keeps the terminology consistent across the document.

64-69: 📐 Maintainability & Code Quality | 🔵 Trivial

Renumber CAP capabilities sequentially.

CAP-17 appears between CAP-3 and CAP-4, breaking numerical sequence. This reduces readability and can cause confusion when referencing capabilities. Renumber MaaS-related capabilities to maintain sequential order (e.g., CAP-4 through CAP-7 for the current set, with MaaS capabilities renumbered to fit the sequence).

Based on past review feedback, this ordering issue was previously flagged but remains unresolved.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/metering-and-usage-tracking/prd.md` around lines 64 - 69,
Renumber the capability list in the metering PRD so the CAP identifiers stay
sequential and readable; the out-of-order CAP-17 in the capabilities section
should be renumbered to follow CAP-3 and before CAP-4, and any related MaaS
entries should be adjusted consistently across the same block. Update the
numbered labels in the capability bullets only, keeping the meaning unchanged,
and ensure the sequence in this section is contiguous when referenced from the
metering-and-usage-tracking document.

36-36: 📐 Maintainability & Code Quality | 🔵 Trivial

Clarify the relationship between metering data and deferred quota enforcement.

Line 36 states that Cloud Provider Admins can query usage "for billing and quota enforcement," while Line 45 lists "quota enforcement" as a non-goal deferred to a separate PRD. The intent appears to be that metering provides data that enables quota enforcement, while the actual quota system is deferred. Consider rewording Line 36 to clarify this distinction — e.g., "query aggregated usage data per tenant as input for billing and quota enforcement systems" — to prevent reader confusion about whether quota enforcement is in scope.

Also applies to: 45-45

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/metering-and-usage-tracking/prd.md` at line 36, Clarify the
scope mismatch in the PRD by updating the metering requirement in the
metering-and-usage-tracking document so Cloud Provider Admins can query
aggregated usage data per tenant as input to billing and quota enforcement
systems, while keeping the actual quota enforcement work deferred in the
non-goals section; update the wording around the affected requirement and the
quota-enforcement non-goal so the metering feature is clearly described as
providing data, not implementing enforcement.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@enhancements/metering-and-usage-tracking/prd.md`:
- Line 116: The acceptance criteria in the metering PRD should explicitly
include cached tokens alongside input, output, and total tokens. Update the
requirement tied to the inference request usage data in the
metering-and-usage-tracking PRD to mention cached tokens by name, using the same
terminology as CAP-17 and the charge model, so the criteria stay aligned and
easier to test.
- Line 184: The flat per-instance-type example uses “vm uptime” as the meter
name, which is inconsistent with the acceptance criteria terminology. Update the
charge model entry in the PRD example to use “instance-type-seconds” or “vm
uptime (instance-type-seconds)” so it matches the wording used in the acceptance
criteria and keeps the terminology consistent across the document.
- Around line 64-69: Renumber the capability list in the metering PRD so the CAP
identifiers stay sequential and readable; the out-of-order CAP-17 in the
capabilities section should be renumbered to follow CAP-3 and before CAP-4, and
any related MaaS entries should be adjusted consistently across the same block.
Update the numbered labels in the capability bullets only, keeping the meaning
unchanged, and ensure the sequence in this section is contiguous when referenced
from the metering-and-usage-tracking document.
- Line 36: Clarify the scope mismatch in the PRD by updating the metering
requirement in the metering-and-usage-tracking document so Cloud Provider Admins
can query aggregated usage data per tenant as input to billing and quota
enforcement systems, while keeping the actual quota enforcement work deferred in
the non-goals section; update the wording around the affected requirement and
the quota-enforcement non-goal so the metering feature is clearly described as
providing data, not implementing enforcement.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 27deda36-9408-4f81-8bf0-8718a92b0bbe

📥 Commits

Reviewing files that changed from the base of the PR and between 7b43f70 and 65b4cac.

📒 Files selected for processing (1)
  • enhancements/metering-and-usage-tracking/prd.md

@pgarciaq pgarciaq left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few observations. I think it's important that we align on terminology too.

I see a reference to FOCUS in the glossary but then it seems there's no reference to FOCUS anywhere ¿?

Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md Outdated
Comment thread enhancements/metering-and-usage-tracking/prd.md
masayag added a commit to masayag/enhancement-proposals that referenced this pull request Jul 5, 2026
- Remove milestone references and roadmap section per reviewer guidance
- Remove Enclave from services table
- Remove CAP-13 (nested cluster dedup) and CAP-18 (GPU compute time)
- Generalize Host type to Resource class
- Remove project from Cloud Provider Admin view (CAP-1)
- Add resource attribution to parent cluster (CAP-12)
- Convert metering failure behavior to open question
- Remove design/implementation open questions
- Remove Cross-Cutting Dimensions and Academic sections
- Tighten acceptance criteria

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Moti Asayag <masayag@redhat.com>
@openshift-ci openshift-ci Bot removed the lgtm label Jul 5, 2026
@masayag

masayag commented Jul 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Updates — addressing mhrivnak's second CHANGES_REQUESTED (2026-07-06)

Glossary — FOCUS alignment (r3530894536, r3530930037)

Replaced non-standard terminology with FOCUS standard definitions:

  • Removed Cost Model — not a FOCUS term; conflates provider infrastructure cost with the pricing applied to tenants. Costing is deferred to a separate PRD.
  • Removed Metric — conflicts with the FOCUS definition (a numeric column type in a billing dataset); was unused in the PRD body.
  • Updated Service, Price List, FOCUS to use FOCUS definitions verbatim.
  • Added Billing Period using the FOCUS definition.

Problem Statement — cost vs price (r3530930037)

Reworked the second paragraph to be precise:

  • "costing layer" → "pricing layer" (OSAC enables providers to apply rates to usage data; it does not track infrastructure cost)
  • Added explicit statement: OSAC does not track the provider's actual infrastructure cost (hardware, power, cooling) — that remains internal to the provider
  • Clarified that providers and tenants have different views of the same usage data, not different cost models

OQ 9.3 — metering unavailability (r3531110178)

Added option (4) per mhrivnak's suggestion: emit events into a durable message bus that guarantees eventual delivery, decoupling OSAC from metering system availability entirely. Events buffer in the bus and are processed when the metering system recovers.

CAP-5 — tooling prescription removed (r3494354429)

Removed "using standard Kubernetes tooling" from CAP-5 — the phrase prescribes the deployment mechanism (a design decision) and unnecessarily constrains scope in a PRD.

CAP numbering

Renumbered all capabilities sequentially (CAP-1 through CAP-16) after removing the configurable meter enable/disable capability, and updated all cross-references in the open questions section.

@github-actions

github-actions Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

AI EP Review: EP-78

Score: 9/10 | Verdict: PASS

Criterion Score Notes
What 2/2 Excellent. 16 numbered capabilities organized by all four OSAC personas (Cloud Provider Admin, Cloud Infrastructure Admin, Tenant Admin, Tenant User). Services in scope explicitly listed (VMaaS, CaaS,
Why 2/2 Strong problem statement names concrete pain: no consumption tracking mechanism exists, providers build fragmented approaches, tenants have no cost visibility. Ties to strategic business model (sovere
How 2/2 Specific and measurable. 19 acceptance criteria are testable by a PM or QA engineer (e.g., 'A running VM generates usage data queryable as aggregated instance-type-seconds per tenant', 'A resource tha
Task 2/2 Clearly a proper enhancement adding a new capability (metering and usage tracking) that does not exist today. Not a bug fix or operational task. The PRD stays at the user-outcome level with only minor
Size 1/2 Bundles three service-specific metering implementations (VMaaS, CaaS, MaaS) that could each ship independently and provide value. MaaS in particular has a fundamentally different metering model (per-t

Verdict: A strong, well-organized PRD with clear user-observable capabilities, concrete business justification, and specific testable acceptance criteria — held back slightly by bundling three separable service-specific metering implementations (VMaaS, CaaS, MaaS) that could ship independently.

Feedback: Consider splitting MaaS metering into its own PRD — it has a fundamentally different model (per-token vs per-time), unique latency requirements, and a different data source, making it independently deliverable and prioritizable. Rewrite CAP-13's latency requirement in user-observable terms: instead of 'emitted within 30 seconds... processed within 60 seconds' (internal pipeline stages), state 'MaaS usage data is queryable within 90 seconds of an inference request completing.' Similarly, rephrase §7's 'Durable event pipeline' dependency in terms of what it provides to operators rather than naming internal architecture.

Critical (0)

None.

Important (3)

  1. Size: MaaS metering has a fundamentally different model (per-token vs per-time), unique latency requirements (30s/60s), and a different data source (inference service vs resource lifecycle). It could ship as its own PRD, enabling independent prioritization and delivery. Custom service metering (CAP-6) is also separable.
  2. Design leakage in CAP-13: 'Metering events must be emitted within 30 seconds of the inference request completing, and processed within 60 seconds of receipt' describes internal pipeline stages. Rewrite as a user-observable outcome: 'MaaS usage data is queryable within 90 seconds of an inference request completing.'
  3. Design leakage in §7 Dependencies: 'Durable event pipeline — a reliable message delivery layer between OSAC and the metering stack' names internal architecture. Rewrite in terms of what the operator needs: 'Reliable delivery of resource lifecycle data to the metering stack, ensuring no events are lost during transient failures.'

Suggestions (3)

  1. The Charge Calculation Model section appears after §9 but is unnumbered, breaking the document's section numbering convention.
  2. Consider adding a milestone or phasing section to indicate which service metering (VMaaS, CaaS, MaaS) is targeted for initial release vs subsequent iterations.
  3. §9.5 on allocation-based metering is thorough but could note whether any current provider has confirmed this requirement, to help prioritize design work.

Review cost

Model: claude-opus-4-6
Cost: $0.5634
Tokens: 811 in / 6.6k out
Cache: 83.8k read
Active time: 2m 16s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-78

Score: 6/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The described metering capabilities (instance-type-seconds, per-node-seconds by host class, per-token inference metering) are well-grounded in established commercial cloud patterns and clearly impleme
Testability 1/2 Section 5 provides detailed, specific acceptance criteria with checkboxes covering VMaaS, CaaS, MaaS, and cross-cutting concerns. Criteria like 'a resource that exists for 30 seconds appears in the us
Scope 2/2 Scope definition is a clear strength. The in-scope services table (VMaaS, CaaS, MaaS) with explicit deferrals (BMaaS, Storage-aaS, networking, bandwidth) is disciplined. Non-goals are specific and inc
Architecture 1/2 The document establishes sound architectural principles: event-driven with immutable events as source of truth, idempotent processing (CAP-16), independent deployment (CAP-14), and non-disruptive upgr

Verdict: A well-structured PRD with strong requirements definition and disciplined scoping, but it reads as a pure requirements document missing the template's expected architectural proposal, implementation details, and test plan sections.

Feedback: The PRD excels at defining what to build but needs a Proposal section covering how: add component architecture (event pipeline, aggregation layer, query API), define the event schema and transport mechanism, and sketch the API surface for usage queries. Add a Test Plan section with at least a high-level strategy covering integration testing of the event pipeline, load testing for concurrent event ingestion, and end-to-end validation of metering accuracy across VM lifecycle transitions. The open question in section 9.3 (metering system unavailability) is architecturally critical and should be resolved before implementation begins, as the choice between blocking provisioning, accepting gaps, reconciliation, or durable messaging fundamentally shapes the system's reliability guarantees and component topology.

Critical (2)

  1. Missing Proposal/Implementation Details section: The template requires a Proposal with Workflow Description, API Extensions, and Implementation Details. The PRD defines requirements but provides no architectural proposal, component design, API surface, event schema, or data model. Without these, reviewers cannot assess whether the design is sound or implementable within the OSAC ecosystem.
  2. Open Question 9.3 (metering unavailability) is unresolved but architecturally load-bearing: the choice between blocking provisioning, accepting gaps, reconciliation, or durable event bus fundamentally determines the system's reliability model, component dependencies, and operational complexity. This must be resolved before design proceeds.

Important (3)

  1. No Test Plan section: The template requires a Test Plan. The acceptance criteria in section 5 are a solid foundation but do not constitute a testing strategy. Missing: how to test metering accuracy under failure conditions, load/scale testing approach, integration test strategy for the event pipeline, and how to validate the 30s/60s MaaS latency SLAs.
  2. Missing Upgrade/Downgrade Strategy and Version Skew Strategy: CAP-15 states upgrades must not cause data loss or measurement gaps, but there is no discussion of how this is achieved — no mention of schema migration, backward compatibility of event formats, or how aggregated data survives a metering stack upgrade.
  3. MaaS data source ownership is unresolved (section 2.3 caveat): MaaS is listed as in-scope, but the PRD acknowledges 'data source ownership to be resolved during design.' If no OSAC component owns the inference event emission path, MaaS metering may not be implementable within the current OSAC architecture, making this a feasibility risk for 1/3 of the in-scope services.

Suggestions (3)

  1. Consider adding a data model section showing the event schema for each service (VMaaS lifecycle events, CaaS node events, MaaS inference events) — even at PRD level, this would help reviewers assess whether the metering units (instance-type-seconds, node-seconds, tokens) can be derived from available lifecycle data.
  2. The operational expectation '13 months retention for aggregated data' should specify whether this means 13 calendar months of rolling retention or 13 months from first data ingestion, and whether the retention clock is per-tenant or global.
  3. Section 9.5 (allocation-based metering) should note whether the current event schema design would need to be extended to support allocation events, or whether this can be added additively later — this affects whether the initial schema design needs to leave room for it.

Review cost

Model: claude-opus-4-6
Cost: $0.4518
Tokens: 789 in / 4.0k out
Cache: 79.3k read
Active time: 1m 28s
API calls: 0

@masayag
masayag requested a review from mhrivnak July 12, 2026 08:13
Add CAP-17 requiring that billing systems can determine the originating
catalog offer and its bundled components for any metered resource, enabling
charge decomposition into independently priceable layers.

Extends the Service glossary definition to describe bundled components.
Adds acceptance criteria for VMaaS (catalog item, instance type, OS image),
CaaS (catalog item, cluster version, host type), and cross-cutting
(full priceable component identification).

Assisted-by: Claude Code <noreply@anthropic.com>
Co-Authored-By: Moti Asayag <masayag@redhat.com>
Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: Moti Asayag <masayag@redhat.com>
@masayag

masayag commented Jul 13, 2026

Copy link
Copy Markdown
Contributor Author

Added CAP-17 — catalog item traceability. Ensures billing systems can determine the originating catalog offer and its bundled components (compute, OS entitlement, boot storage for VMaaS; control plane, workers, cluster version for CaaS) for any metered resource, so charges can be decomposed into independently priceable layers using the provider's rate schedule. Also added corresponding acceptance criteria for VMaaS, CaaS, and cross-cutting.

@github-actions

github-actions Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

AI EP Review: EP-78

Score: 10/10 | Verdict: PASS

Criterion Score Notes
What 2/2 Excellent persona coverage (all 4 OSAC personas) with 17 specific, user-observable capabilities organized by role. Services in scope (VMaaS, CaaS, MaaS) clearly identified with a scoping table. Non-go
Why 2/2 Concrete problem statement: OSAC has no mechanism to track resource consumption, providers build fragmented solutions, Cloud Provider Admins cannot generate bills or enforce quotas. Clear causal chain
How 2/2 Approach is specific and measurable: metering units defined (instance-type-seconds, per-token), charge calculation formulas with examples, MaaS latency requirements (30s emit, 60s processing), retenti
Task 2/2 Clear enhancement adding new metering and usage tracking capabilities that do not exist in OSAC today. Not a bug fix, operational task, or refactor.
Size 2/2 Coherent scope: metering infrastructure plus per-service meters are tightly coupled — the platform is useless without services to meter, and meters can't function without the platform. Three services

Verdict: Exceptionally well-structured PRD with clear persona-driven capabilities, concrete business justification, measurable requirements, and well-bounded scope that properly defers billing and costing to separate work.

Feedback: This PRD is ready for design. Minor refinements to consider: CAP-6 and CAP-14 use slightly implementation-leaning language ('emit metering events,' 'lifecycle events') — reframing as user outcomes ('track consumption of custom services,' 'integrate with external metering systems') would sharpen the user focus. The open questions (§9.1–9.6) are well-framed but §9.3 lists four implementation options — consider focusing that question on the user-observable behavior (should provisioning block or proceed?) rather than the mechanism.

Critical (0)

None.

Important (2)

  1. CAP-6 and CAP-14 use somewhat implementation-oriented language ('emit metering events for custom services,' 'OSAC emits lifecycle events'). These are not user-observable actions — consider reframing as 'track consumption of custom services' and 'support integration with external metering systems' to stay at the PRD level.
  2. Open Question 9.3 enumerates four implementation options (block provisioning, accept gaps, reconciliation service, durable message bus). The PRD-level question is about user-observable behavior: should provisioning proceed when metering is unavailable? The mechanism belongs in the design document.

Suggestions (3)

  1. The Charge Calculation Model section with pricing formulas and worked examples is useful domain context but straddles the PRD/design boundary. Consider keeping the metering units and grouping dimensions but moving the formula tables to the design document.
  2. Section 9.5 (allocation-based metering) and 9.6 (infrastructure-failure exemption) are rich open questions that could each become their own PRD capability if resolved affirmatively — flagging for the author to track resolution.
  3. Consider adding an explicit capability for audit trail / data integrity verification — the PRD mentions 'billing-grade metering' in the glossary but no capability directly addresses how a provider would verify metering accuracy.

Review cost

Model: claude-opus-4-6
Cost: $0.5447
Tokens: 811 in / 5.9k out
Cache: 84.7k read
Active time: 2m 2s
API calls: 0

Replace open questions (§9) with five decisions:
- D-1: Tenant Users see project-scoped usage via RBAC
- D-2: Metering failures must not affect provisioning
- D-3: Catalog item traceability resolves tenant-defined Services
- D-4: VMaaS/CaaS compute remains consumption-based only
- D-5: Failed-state resources are not metered

Remove design leakage from CAP-13 (pipeline emission/processing language),
CAP-14 (lifecycle event emission), and D-2 (durable message bus). Make
CAP-6 concrete (configuration-based meter registration). Fix stale §9.6
reference in CAP-11 to D-5. Number Charge Calculation Model as §10. Add
MaaS catalog traceability acceptance criterion.

Assisted-by: Claude Code <noreply@anthropic.com>
Co-Authored-By: Moti Asayag <masayag@redhat.com>
Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: Moti Asayag <masayag@redhat.com>
@masayag

masayag commented Jul 13, 2026 •

Copy link
Copy Markdown
Contributor Author

Resolved all open questions as decisions (§9):

  • D-1: Tenant Users see usage scoped to projects they have access to (existing RBAC model)
  • D-2: Metering failures never affect provisioning; all data is eventually captured with no billing gaps
  • D-3: Tenant-defined catalog items handled transparently via CAP-17 catalog traceability
  • D-4: VMaaS/CaaS compute stays consumption-based; allocation metering addressed in BMaaS Part 2 PRD
  • D-5: Failed resources are not metered regardless of failure cause; SLA credits are billing policy

Also: removed design leakage from CAP-13/CAP-14, made CAP-6 concrete (configuration-based meter registration), added MaaS catalog traceability acceptance criterion, numbered Charge Calculation Model as §10.

@github-actions

github-actions Bot commented Jul 13, 2026 •

Copy link
Copy Markdown

AI EP Review: EP-78

Score: 9/10 | Verdict: PASS

Criterion Score Notes
What 2/2 Excellent persona and service coverage. Four personas (Cloud Provider Admin, Cloud Infrastructure Admin, Tenant Admin, Tenant User) each have specific, numbered capabilities (CAP-1 through CAP-17). In
Why 2/2 Strong business justification with concrete pain points: 'no mechanism to track consumption over time' leading to 'fragmented approaches and inconsistent data models.' Ties directly to sovereign cloud
How 2/2 Acceptance criteria are specific, measurable, and verifiable by using the product. Examples: 'A stopped or paused VM does not generate compute usage data,' 'A metering event is emitted within 30 secon
Task 2/2 This is a proper enhancement PRD, not a task or design document. The document is remarkably clean of implementation details — no controllers, reconcilers, playbooks, or internal conditions are named.
Size 1/2 Bundles metering for three services (VMaaS, CaaS, MaaS) with meaningfully different characteristics — time-based vs token-based, different latency requirements (60s for MaaS vs polling interval for VM

Verdict: A well-structured, user-focused PRD with strong persona coverage, concrete business justification, and testable acceptance criteria, held back slightly by bundling three independently shippable service meters into one document.

Feedback: Consider whether VMaaS/CaaS (time-based) and MaaS (token-based with stricter latency requirements) should be separate PRDs, since they have different metering models and could be prioritized and delivered independently. The custom meter registration capability (CAP-6) could also stand alone. If keeping them bundled, add explicit phasing guidance in the PRD to clarify which services must ship together vs. which can be incrementally added.

Critical (0)

None.

Important (2)

  1. Scope bundles three independently shippable service meters (VMaaS, CaaS, MaaS) with different metering models and latency requirements, plus custom meter registration — consider splitting or adding explicit phasing.
  2. MaaS data source ownership is listed as 'to be resolved during design' (Section 2.3) — this open dependency could block an entire service's metering if not resolved early.

Suggestions (3)

  1. The Risks section (Section 8) is light — only two risks with generic mitigations. Consider adding risks around data volume at scale (high-cardinality metering dimensions), clock skew affecting per-second granularity, and the retroactive adjustment mechanism for failed-state resources (D-5).
  2. CAP-13 specifies 60-second end-to-end latency for MaaS but the acceptance criteria says 30 seconds for emission + 60 seconds for processing — clarify whether these are additive (90s total) or the 60s is the end-to-end target.
  3. Dependencies section could benefit from naming specific OSAC components that must emit lifecycle events and confirming they already do so or will need modification.

Review cost

Model: claude-opus-4-6
Cost: $0.5387
Tokens: 812 in / 5.0k out
Cache: 137.3k read
Active time: 1m 51s
API calls: 0

@ronniel1

Copy link
Copy Markdown

/lgtm

@ronniel1

Copy link
Copy Markdown

/approve

@openshift-ci

openshift-ci Bot commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: avishayt, masayag, ronniel1

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:
  • OWNERS [avishayt,masayag,ronniel1]

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants