Repository navigation
test(e2e): typed per-test metadata for the e2e suite - #42044
Conversation
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
f272484 to
f484cc2
Compare
|
bugbot run |
f484cc2 to
9d5a3ef
Compare
9d5a3ef to
11dbe70
Compare
5b6f0b4 to
8514ab8
Compare
fad15f7 to
01d77ac
Compare
01d77ac to
f9988bb
Compare
|
@greptile re review |
|
bugbot run |
|
@greptile re review |
@meta(Subject(domain, route, providers, models, capabilities, mode)) declares what a test is about with closed enums, and each field lands in the JUnit report as a property. The quota_management suites are the first to declare it.
It already imports pydantic and pytest, both of which the suite needs to collect. The rule that matters is no litellm import
A budget or rate-limit test whose chat call only triggers the block now leaves route unset, since its steps already name the call. Tests of an endpoint keep it: budget CRUD, key creation, spend reporting reads, and the per-endpoint spend tests for chat, messages, embeddings, batches and health. The two /spend/logs tests tagged chat_completions are now spend_reporting
subject_properties seeded a list and grew it with append and extend. It now flattens one tuple per field, and the plural-name table is a read-only mapping
The breadth test gave all 33 probes spend_reporting, so /key/list, /user/list, /team/list, /organization/list and /customer/list counted as spend reporting. Each case now carries its own route, with organization and customer management added to Route
fb7281c to
dc70ca0
Compare
|
@greptile re review |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit dc70ca0. Configure here.
|
@greptile re review |
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
Builds on the recorded
@steplog from #42393, which has merged. This PR is the declared half.This PR adds typed, declarative metadata to e2e tests, so coverage can be sliced by provider, model, endpoint and capability instead of by folder name. It is additive only. The coverage-registry YAML is untouched, and every existing
@pytest.mark.covers("cell.id")keeps working.The shape
The marker takes one frozen dataclass of enum members as its single argument, not kwargs. That way a typo is a type error rather than a silent miss:
dataclasses.asdict()turns that into<property>pairs with no per-field plumbing.providers,modelsandcapabilitiesare all tuples, because one test node often drives several of each. Theclaude_codematrix runs haiku, sonnet and opus in a single body, and a spend-attribution test calls a Gemini and an Anthropic model on one key.models=("gpt-5.5")is an error naming the file rather than one model per character.test_e2e_metadata.pyfails on any model typed out as a string literal in@metaThe declared fields go after the fixed
package/covers/sourceprefix, which stays byte-identical.Decisions worth a look
capabilities,llm_translation/test_together_ai_e2e.pyalready hand-rolls a_Needsdataclass that ANDs two capabilities.coverage_registry/schema.pyhad to smuggle conjunctions in as ad-hoc values likethinking_with_tool_use. A tuple collapses both back.class X(str, Enum), notStrEnum, becausepyproject.tomlfloors at Python 3.10.litellmimport.tests/e2eis a black-box HTTP suite that importslitellmin zero files today, and makinglitellma collection-time dependency would be a real regression. SoProvideris its own enum rather than a re-export ofLlmProviders, andTestProviderMirrorsLitellmfails on drift wherever litellm is importable.metamarker, not an overload ofcovers. A dataclass passed positionally tocoverswould be silently dropped bydedupe_covers'sisinstance(str)filter, so the registry collector would read the test as uncovered.routeis the endpoint the test checks. A/team/updatetest getsteam_managementand a/spend/logstest getsspend_reporting. A budget or rate-limit test whose chat call only triggers the block leaves it unset, since its steps already name the callSubjectper test item.subject_propertiesreadsget_closest_marker("meta"), so a parametrized matrix builds its fullSubjectin one place per param. test(e2e): add conversational matrix across chat, messages and responses #42359'scells_covering()already has that shape.Pilot usage
tests/e2e/quota_management/is annotated for real: 29 files, all withdomain=spend-budgets. Tests that drive more than one provider or model declare all of them, including the shared-key, budget-fallback, per-model-budget and key-attribution tests. The rest of the suite is a separate follow-up PR, and metadata is optional until that lands.Verification
test_e2e_metadata.py+test_e2e_junit_report.py(as the Code Qualitytest_e2e_metadatastep runs them) +tests/e2e/test_junit_properties.py-m "not e2e")--collect-onlybasedpyright tests/e2erufftests/code_coverage_tests/test_e2e_junit_report.pyruns real pytest with--junitxmlthrough the realtests/e2e/conftest.py, in-process and under-n 2. It asserts that the repeatedprovider/model/capabilityproperties land in the parsed XML, and that a bare-strmodels=fails collection.httpx2import errors as on feat(e2e): record each e2e test's steps, starting with ProxyClient #42393.mcp/oauth_chat_client.py, the same as on feat(e2e): record each e2e test's steps, starting with ProxyClient #42393.This is part of the e2e test-metadata rollout; see the rollout plan. It is consumed by BerriAI/project-releaser#252.
Note
Low Risk
Changes are confined to e2e harness metadata, JUnit reporting, and test annotations; production proxy behavior is untouched and existing
@pytest.mark.coversbehavior is preserved.Overview
Adds a declared half to e2e test metadata alongside the existing recorded
@steplog: tests can attach@meta(Subject(...))with closed enums for domain, route, providers, models, capabilities, and mode. The frozenSubjectdataclass validates plural fields (tuple-only, deduped/sorted), serializes viasubject_properties()into JUnit<property>entries (repeated singular names forprovider/model/capability), and appends after the unchangedpackage/covers/sourceprefix inresult_properties.Registers the
metamarker in conftest/pytest.ini, documents usage inAGENTS.md, and expands harness tests (test_e2e_metadata.py,test_e2e_junit_report.py) including end-to-end JUnit round-trip and collection failures for bare-stringmodels=(x). Pilot backfill: annotates alltests/e2e/quota_management/**tests with@meta, refactors model strings to named constants where needed, and adds a guard that@metamust not hand-type model string literals.Reviewed by Cursor Bugbot for commit dc70ca0. Bugbot is set up for automated code reviews on this repo. Configure here.