Skip to content

test(e2e): typed per-test metadata for the e2e suite - #42044

Merged
ryan-crabbe-berri merged 7 commits into
mainfrom
feat/typed-e2e-test-metadata
Oct 1, 2026
Merged

ryan-crabbe-berri merged 7 commits into
mainfrom
feat/typed-e2e-test-metadata

Conversation

@ryan-crabbe-berri

@ryan-crabbe-berri ryan-crabbe-berri commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Builds on the recorded @step log from #42393, which has merged. This PR is the declared half.

This PR adds typed, declarative metadata to e2e tests, so coverage can be sliced by provider, model, endpoint and capability instead of by folder name. It is additive only. The coverage-registry YAML is untouched, and every existing @pytest.mark.covers("cell.id") keeps working.

The shape

The marker takes one frozen dataclass of enum members as its single argument, not kwargs. That way a typo is a type error rather than a silent miss:

@meta(Subject(
    domain=Domain.SPEND_BUDGETS,
    providers=(Provider.OPENAI,),
    models=(BACKEND,),
    capabilities=(Capability.REASONING,),
))
def test_priority_tier_bills_priority_rates(...): ...

dataclasses.asdict() turns that into <property> pairs with no per-field plumbing.

providers, models and capabilities are all tuples, because one test node often drives several of each. The claude_code matrix runs haiku, sonnet and opus in a single body, and a spend-attribution test calls a Gemini and an Anthropic model on one key.

  • Canonical form: all three share one canonicaliser, which dedupes them and sorts by serialised value.
  • JUnit encoding: each is written as a repeated property under the singular name, so there is no delimiter to corrupt.
  • No pairing: the lists are independent sets. A model is not tied to a provider by position.
  • Tuples only: anything but a tuple is refused at collection time. models=("gpt-5.5") is an error naming the file rather than one model per character.
  • Constants, not copies: a declared model names the constant the test calls, so the report can't keep naming a model the test stopped driving. A guard in test_e2e_metadata.py fails on any model typed out as a string literal in @meta

The declared fields go after the fixed package / covers / source prefix, which stays byte-identical.

Decisions worth a look

  • Every set-valued axis is plural. For capabilities, llm_translation/test_together_ai_e2e.py already hand-rolls a _Needs dataclass that ANDs two capabilities. coverage_registry/schema.py had to smuggle conjunctions in as ad-hoc values like thinking_with_tool_use. A tuple collapses both back.
  • class X(str, Enum), not StrEnum, because pyproject.toml floors at Python 3.10.
  • The suite stays free of a litellm import. tests/e2e is a black-box HTTP suite that imports litellm in zero files today, and making litellm a collection-time dependency would be a real regression. So Provider is its own enum rather than a re-export of LlmProviders, and TestProviderMirrorsLitellm fails on drift wherever litellm is importable.
  • New meta marker, not an overload of covers. A dataclass passed positionally to covers would be silently dropped by dedupe_covers's isinstance(str) filter, so the registry collector would read the test as uncovered.
  • route is the endpoint the test checks. A /team/update test gets team_management and a /spend/logs test gets spend_reporting. A budget or rate-limit test whose chat call only triggers the block leaves it unset, since its steps already name the call
  • One Subject per test item. subject_properties reads get_closest_marker("meta"), so a parametrized matrix builds its full Subject in one place per param. test(e2e): add conversational matrix across chat, messages and responses #42359's cells_covering() already has that shape.

Pilot usage

tests/e2e/quota_management/ is annotated for real: 29 files, all with domain=spend-budgets. Tests that drive more than one provider or model declare all of them, including the shared-key, budget-fallback, per-model-budget and key-attribution tests. The rest of the suite is a separate follow-up PR, and metadata is optional until that lands.

Verification

Check Result
test_e2e_metadata.py + test_e2e_junit_report.py (as the Code Quality test_e2e_metadata step runs them) + tests/e2e/test_junit_properties.py 73 passed
Offline harness run (-m "not e2e") 610 passed
Full-tree --collect-only 1345 collected, 2 errors
basedpyright tests/e2e 64 errors, none added
ruff clean

This is part of the e2e test-metadata rollout; see the rollout plan. It is consumed by BerriAI/project-releaser#252.


Note

Low Risk
Changes are confined to e2e harness metadata, JUnit reporting, and test annotations; production proxy behavior is untouched and existing @pytest.mark.covers behavior is preserved.

Overview
Adds a declared half to e2e test metadata alongside the existing recorded @step log: tests can attach @meta(Subject(...)) with closed enums for domain, route, providers, models, capabilities, and mode. The frozen Subject dataclass validates plural fields (tuple-only, deduped/sorted), serializes via subject_properties() into JUnit <property> entries (repeated singular names for provider/model/capability), and appends after the unchanged package/covers/source prefix in result_properties.

Registers the meta marker in conftest/pytest.ini, documents usage in AGENTS.md, and expands harness tests (test_e2e_metadata.py, test_e2e_junit_report.py) including end-to-end JUnit round-trip and collection failures for bare-string models=(x). Pilot backfill: annotates all tests/e2e/quota_management/** tests with @meta, refactors model strings to named constants where needed, and adds a guard that @meta must not hand-type model string literals.

Reviewed by Cursor Bugbot for commit dc70ca0. Bugbot is set up for automated code reviews on this repo. Configure here.

@greptile-apps

greptile-apps Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Low risk] Adds typed metadata decorators to test files.

The PR appears safe to merge based on this review

Summary

The PR adds typed, optional e2e test metadata to JUnit properties and annotates quota-management tests while retaining existing coverage markers. Since the previous review, changes trim explanatory docstrings

Reviews (11) · Last reviewed commit: "docs(e2e): trim excessive docstrings in ..."

Comment thread tests/e2e/conftest.py Outdated
Comment thread tests/e2e/conftest.py Outdated
Comment thread tests/e2e/quota_management/spend_tracking/test_spend_tracking_e2e.py Outdated
Comment thread tests/e2e/e2e_metadata.py Outdated
@codecov

codecov Bot commented Sep 19, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@ryan-crabbe-berri ryan-crabbe-berri changed the title Typed per-test metadata for the e2e suite test(e2e): typed per-test metadata for the e2e suite Sep 21, 2026
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the feat/typed-e2e-test-metadata branch from f272484 to f484cc2 Compare September 21, 2026 17:57
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptileai

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Comment thread tests/e2e/test_junit_report.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread tests/e2e/conftest.py Outdated
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the feat/typed-e2e-test-metadata branch from f484cc2 to 9d5a3ef Compare September 22, 2026 01:52
@ryan-crabbe-berri
ryan-crabbe-berri changed the base branch from main to feat/e2e-step-log September 22, 2026 01:54
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the feat/typed-e2e-test-metadata branch from 9d5a3ef to 11dbe70 Compare September 22, 2026 02:13
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the feat/typed-e2e-test-metadata branch from fad15f7 to 01d77ac Compare September 28, 2026 19:25
@ryan-crabbe-berri
ryan-crabbe-berri force-pushed the feat/typed-e2e-test-metadata branch from 01d77ac to f9988bb Compare September 30, 2026 20:53
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

Base automatically changed from feat/e2e-step-log to main October 1, 2026 02:33
Comment thread tests/e2e/quota_management/spend_tracking/test_spend_routes.py Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@meta(Subject(domain, route, providers, models, capabilities, mode)) declares
what a test is about with closed enums, and each field lands in the JUnit report
as a property. The quota_management suites are the first to declare it.
It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import
43 @meta declarations in quota_management typed the model name out again, so changing the call would leave the coverage report naming the old model. Each file now has one constant used by both, and a guard fails on any model written as a string literal in @meta
A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting
subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping
The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@ryan-crabbe-berri
ryan-crabbe-berri enabled auto-merge (squash) October 1, 2026 02:41
Comment thread tests/code_coverage_tests/test_e2e_metadata.py Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit dc70ca0. Configure here.

@ryan-crabbe-berri

Copy link
Copy Markdown
Contributor Author

@greptile re review

@ryan-crabbe-berri
ryan-crabbe-berri merged commit ae05f7d into main Oct 1, 2026
91 of 96 checks passed
@ryan-crabbe-berri
ryan-crabbe-berri deleted the feat/typed-e2e-test-metadata branch October 1, 2026 04:03
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...

This branch is waiting to be deployed

1 waiting deployment
e2e-changed — 8561cbf6 Waiting Oct 1, 2026 by ryan-crabbe-berri via Run changed e2e tests against the stage-mirror stack #15546
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants