Repository navigation
feat(e2e): record each e2e test's steps, starting with ProxyClient - #42393
Conversation
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
|
@greptileai re review |
5b6f0b4 to
8514ab8
Compare
@step on a harness method records a plain-English line for every call, in order, as repeated JUnit step properties. Labels are templates filled from the call's parameters, like "Generate a virtual key with models: claude-haiku-4-5 and rpm limit: 3", and secret request fields are marked Field(repr=False) so they never print. ProxyClient and the rate-limit QuotaClient carry steps first; the other harnesses follow one area at a time. The recorder and JUnit tests run in the Code Quality workflow's test_e2e_metadata step.
5e8cbaf to
79ff294
Compare
|
@greptileai re review |
|
bugbot run |
The failed setup or call report of an mcp_oauth_live test copied user_properties before the steps were attached, so it carried no steps. Every setup and call report now takes its properties after the steps attach
|
bugbot run |
The create_credential label read credential_info, which defaults to {} and is never set by the live callers, so the step printed nothing after 'for'. It now reads the required credential_name, and a guard fails on any label that reads a field with a default
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 098c188. Configure here.
|
@greptileai re review |
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
Record what every e2e test actually did, step by step, from the harness it calls. The recorded steps become the test's user story in the JUnit report, and a failing test's last step is where it died. Nothing is hand-written, so the story can't drift from the code.
This is the steps half of #42044, split out so it can be reviewed on its own. #42044 is re-stacked on top of this branch and now holds only the declared
@meta(Subject(...))half.The shape
@stepgoes on harness helpers, never on tests:test_rpm_limit_blocks_over_limit's report then carries these as repeated<property name="step">entries:The label is recorded before the wrapped call, so a helper that raises still leaves its own label last.
This PR covers all 45
ProxyClientmethods, includingresponses_stream, which main added after this branch started, and the rate-limit suite'sQuotaClient.chat. The other harnesses (the domain clients,lifecycle,idp, the logging readers, migrations, the Claude CLI driver) get steps one area at a time in follow-ups, so each area's wording can be reviewed on its own. Until then their tests show only theProxyClientcalls they makeLabel templates
A label can name the helper's own parameters, so the story says what the test asked for.
@step("Generate a virtual key with {body}")records "Generate a virtual key with models: claude-haiku-4-5 and rpm limit: 3". Only what a label names reaches the report, aField(repr=False)field is never shown, and a placeholder the helper doesn't take fails at import. A dotted placeholder reads one field of a request model, as in{body.litellm_params.model}, and a guard test fails if one names a field the model doesn't have. Secret request fields (api_key, AWS and Vertex credentials,credential_values, and the Langfuse and W&B keys a logging callback carries in key metadata) are markedField(repr=False), so no template can print them. As a backstop, the recorder masks the value of every secret-named environment variable wherever it lands in a label, so a credential in a dict, a prompt or an unmarked field still comes out as***.Decisions worth a look
ProxyClient.create_modelgoes throughregister_model, and a domain client wraps the sharedProxyClient. So every layer carries a label, and the story still reads at the level the test called in at, one beat per action. A helper that fans work out to threads still records its workers' steps.@stepgoes above@contextmanager, and the setup and cleanup aroundyieldrun inside the step while thewithbody records normally. Without this, cleanup such asrestricted_user'sDROP ROLEwould append a step after the one a test died on.TypeErrorat import, because its body interleaves with the caller's.ProxyClient.delete_*helpers now usestacklevel=2 + STEP_FRAMES, so they still report at their caller rather than ate2e_metadata.py. A test pins the frame count.pytest_runtest_setup. An autouse fixture would run after wider-scoped fixtures and inherit a module finalizer's steps.--reruns 1doesn't double it.(N earlier steps not recorded)line, so the last step is still where the test died.package/covers/sourceprefix is byte-identical; steps only ever append after it.tests/e2e, which holds only tests that drive a live proxy. They aretests/code_coverage_tests/test_e2e_metadata.pyandtest_e2e_junit_report.py, run by thetest_e2e_metadatastep of the Code Quality GitHub Actions workflow, so they gate every PR. The unit file covers only the recorder's edge cases (dedupe, the cap, nesting, context managers). Call order, the last step, the per-test reset and the attach are asserted once, in the real-pytest report testVerification
test_e2e_metadata.py+test_e2e_junit_report.py, run as the Code Quality step runs them-m "not e2e", minus the liveclaude_codeCLI tests)--collect-onlybasedpyrighton every changed harness fileruff check --config ruff-tests.tomlassert_ci_coverage.pytest_e2e_junit_report.pyruns real pytest with--junitxmlthroughtests/e2e/conftest.py, in-process and under-n 2. It asserts on the parsed XML for the passing, failing, setup-error, rerun and wide-scope-fixture cases.ProxyClientand fails if one names a field its request model lacks. A deliberate{body.modle}typo fails it{body.credential_info}label oncreate_credentialfails itProxyClientlabel can print, nested ones included, and fails on any secret-named field that isn'trepr=False. Unhidingwandb_api_keyfails itmcp_oauth_livetest and of a plain one, and both must carry the steps. Attaching the steps after the oauth snapshot again fails ituser_propertiesinto the teardown report by then, so it never reaches the XMLThis is safe to merge on its own. The
stepproperties are ignored by the current emitter, and BerriAI/project-releaser#252 regroups them into the results JSON'sstepsarray. It is part of the e2e test-metadata rollout; see the rollout plan.Note
Medium Risk
Changes pytest reporting hooks and published JUnit metadata; incorrect attach timing or masking could mislead triage or leak secrets, though the PR adds broad guards and tests for both.
Overview
Adds runtime step recording for e2e harness calls via a new
@stepdecorator andStepRecorderine2e_metadata.py. Labels can use argument placeholders, dedupe consecutive repeats, cap at 50 steps, mask env secrets, and only record the outermost call in nested helpers.ProxyClient(andQuotaClient.chat) are annotated so each test’s story becomes repeated<property name="step">entries in JUnit XML.conftest.pyclears the log at setup and attaches steps after setup/call (not teardown), viaattach_step_propertiesinjunit_properties.py. Request models mark sensitive fieldsrepr=Falseso templates cannot leak credentials.CI runs new
test_e2e_metadata.py(recorder edge cases) andtest_e2e_junit_report.py(real pytest + junitxml, including xdist/reruns).tests/e2e/AGENTS.mddocuments how to write labels and how steps flow to reports.Reviewed by Cursor Bugbot for commit 098c188. Bugbot is set up for automated code reviews on this repo. Configure here.