Skip to content

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP) - #402

Merged
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit
May 26, 2026
Merged

feat(embeddings): emit UsageEvent on /v1/embeddings 200 (#226 MVP)#402
moonming merged 2 commits into
mainfrom
feat/issue-226-embeddings-usage-emit

Conversation

@moonming

@moonming moonming commented May 26, 2026

Copy link
Copy Markdown
Member

Summary

Refs #226 — embeddings only; remaining 5 endpoints tracked under #226 with checkboxes ticked as each follow-up PR lands.

Pre-fix, `/v1/embeddings` dropped the `UsageEvent` entirely. Every embeddings request was invisible to cp-api's budget ledger and the customer-facing `/logs` analytics. `chat.rs` has emitted `UsageEvent` since #302 M17 — the gap was specifically on the non-chat OpenAI-shape handlers.

This PR ships the embeddings half: a successful `/v1/embeddings` call now emits a `UsageEvent` on the configured sink with the upstream-reported `prompt_tokens`, the resolved `model_id`, the authenticated `api_key_id`, the status code, and `inbound_protocol = "openai"`.

Scope

MVP: /v1/embeddings only. Follow-ups for /v1/completions (#403), /v1/responses (#404), /v1/rerank (#405), /v1/audio/* (#406), /v1/images/* (#407) — same shape, same emit-on-success-only convention.

Per-PK telemetry attribution (`provider_kind` / `featured` / `branded_provider` / `pk_label` / `byo_label`) is wired in `chat.rs` only today. Adding it to the non-chat handlers is also a follow-up; backward-compat is preserved here because the `UsageEvent` defaults map empty/false → cp-api stores NULL.

Emission convention

Mirrors `chat.rs::emit_usage_event` — emit when an upstream call was made:

Path Emit? Reason
200 OK yes upstream returned a response — attribute spend (even when `prompt_tokens=0`)
501 NotImplemented no no upstream call happened (`upstream_called: false`) — skip
upstream 4xx/5xx error no no usage data to attribute

Gating is on the explicit `EmbedDispatchSuccess.upstream_called` flag (audit M1 fix — see commit 164c364) rather than a `prompt_tokens > 0` sentinel, so a future provider that legitimately reports zero tokens on a 200 still gets attributed.

Audit response

Independent agent review found 2 MEDIUM + helpful LOWs. Resolutions:

Test plan

  • Unit test: `emits_usage_event_on_200_with_prompt_tokens_issue_226` — drives a successful 200 with `prompt_tokens=42`, pins the wire shape (`prompt_tokens`, `completion_tokens=0`, `status_code=200`, `model_id`, `api_key_id`, `inbound_protocol="openai"`, non-empty `request_id` + `occurred_at`).
  • Audit M1 regression: `emits_usage_event_on_200_with_zero_prompt_tokens_audit_m1` — drives a successful 200 with `prompt_tokens=0` and asserts the event still arrives, pinning the `upstream_called`-based gate.
  • All 12 pre-existing `embeddings::tests` still pass — no regression in dispatch, input-shape preservation, or upstream error mapping.
  • `cargo clippy -p aisix-proxy -- -D warnings` clean.

References

Pre-#226, /v1/embeddings dropped the UsageEvent entirely — every
embeddings call was invisible to cp-api's budget ledger and the
customer-facing /logs analytics. Chat completions has emitted
UsageEvents since #302 M17, so the gap was specifically on the
non-chat OpenAI-shape handlers.

This PR ships the MVP for #226: /v1/embeddings now emits a
UsageEvent on 200 with the upstream-reported prompt_tokens, the
resolved model_id, the authenticated api_key_id, the status code,
and inbound_protocol = "openai" (matching chat.rs convention).

Emit-on-success-only mirrors chat.rs:
- 200 path → emit with real prompt_tokens from upstream usage block
- 501 Not Implemented (provider lacks embed support) → no upstream
  call happened, so prompt_tokens=0 signals "skip emit" to the
  handler (avoids attributing zero-token spend to the api_key and
  bloating /logs with noise)
- Upstream error path → no UsageEvent (no usage data to attribute)

Per-PK telemetry attribution (provider_kind / featured /
branded_provider / pk_label / byo_label) is wired for chat only;
filed as a follow-up so the non-chat handlers gain the same
dashboard-slicing surface.

Follow-ups (separate PRs) for /v1/completions, /v1/responses,
/v1/rerank, /v1/audio/*, /v1/images/* — same shape, same
emit-on-success-only convention.

Test: unit test asserts a successful /v1/embeddings call enqueues
exactly one UsageEvent on the sink with the expected prompt_tokens,
model_id, api_key_id, status_code, inbound_protocol, and non-empty
request_id + occurred_at. Mirrors chat.rs's existing UsageSink-
based unit-test pattern.
@coderabbitai

coderabbitai Bot commented May 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@moonming, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 minutes and 16 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Free

Run ID: bf0d9617-d2ab-4f08-a962-9c9651507583

📥 Commits

Reviewing files that changed from the base of the PR and between 869f617 and 164c364.

📒 Files selected for processing (1)
  • crates/aisix-proxy/src/embeddings.rs
📝 Walkthrough

Walkthrough

This PR enhances the embeddings endpoint to emit UsageEvent for successful requests. The handler now conditionally emits usage events when prompt_tokens > 0, backed by an internal dispatch refactor that captures and returns token counts. A new helper function builds and emits events to the usage sink and OTLP exporters, with test coverage validating the feature.

Changes

Embeddings usage event emission

Layer / File(s) Summary
Handler and imports update
crates/aisix-proxy/src/embeddings.rs
Updated embeddings handler to conditionally emit UsageEvent on successful requests when prompt_tokens > 0, using model, provider, api-key, status, latency, and resolved request identifiers. Added RequestOutcome to imports.
Dispatch structure and implementation
crates/aisix-proxy/src/embeddings.rs
Introduced EmbedDispatchSuccess struct containing Response, provider label, model_id, and prompt_tokens. Updated dispatch function to return this struct instead of a tuple. Refactored "not supported" error path to return prompt_tokens: 0 with 501 envelope, suppressing usage emission.
Usage event emission helper
crates/aisix-proxy/src/embeddings.rs
Added emit_usage_event helper that constructs a UsageEvent with embeddings-specific defaults and pushes it to the usage_sink and OTLP/HTTP exporter fan-out from the live snapshot, mirroring chat.rs conventions.
Regression test for usage event emission
crates/aisix-proxy/src/embeddings.rs
Added test validating that a successful 200 embeddings request emits exactly one UsageEvent with expected prompt_tokens, zero completion_tokens, HTTP status code, api_key_id, model_id, inbound_protocol = "openai", and valid timestamps/request IDs.

🎯 3 (Moderate) | ⏱️ ~25 minutes


Note

🎁 Summarized by CodeRabbit Free

Your organization has reached its limit of developer seats under the Pro Plan. For new users, CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please add seats to your subscription by visiting https://app.coderabbit.ai/login.If you believe this is a mistake and have available seats, please assign one to the pull request author through the subscription management page using the link above.

Comment @coderabbitai help to get the list of available commands and usage tips.

…mpt_tokens (#226 audit M1)

PR #402 audit found that gating emission on `prompt_tokens > 0`
conflates two distinct cases:

- 501 NotImplemented (no upstream call) — emit MUST be skipped
- 200 with upstream-reported `prompt_tokens=0` (rare but possible
  for empty input or provider-specific billing) — emit MUST happen

The audit-suggested fix: replace the numeric sentinel with an
explicit `upstream_called: bool` flag on EmbedDispatchSuccess.
Dispatch sets `true` in the 200 arm and `false` in the 501 arm; the
handler gates on the flag directly.

Regression test pins the post-fix contract: a 200 with
`usage.prompt_tokens=0` still produces exactly one UsageEvent on the
sink, attributed to the authenticated api_key for compliance /
audit even when the billable count is zero.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Non-chat endpoints don't emit UsageEvents — cp-api's spend ledger misses spend from /v1/responses, /v1/embeddings, /v1/audio*, /v1/images*, /v1/rerank

1 participant