Skip to content

fix(test): stabilize openai compat oversized-body regression - #839

Merged
zmanian merged 2 commits into
nearai:stagingfrom
CPU-216:fix/openai-compat-body-limit-flake
Mar 12, 2026
Merged

zmanian merged 2 commits into
nearai:stagingfrom
CPU-216:fix/openai-compat-body-limit-flake

Conversation

@CPU-216

@CPU-216 CPU-216 commented Mar 10, 2026

Copy link
Copy Markdown
Contributor

Summary

Stabilize the oversized-body regression coverage for the OpenAI-compatible /v1/chat/completions endpoint.

The previous integration test intermittently observed 503 Service Unavailable instead of the expected 413 Payload Too Large, even though the route-level body limit remained correctly configured. This change keeps the fix scoped to regression stability and coverage.

Root Cause

test_chat_completions_body_too_large relied on a spawned TCP gateway server inside the openai_compat_integration test binary. That made the assertion sensitive to non-deterministic gateway lifecycle behavior in the surrounding integration process.

The flaky failure path returned 503, which matches the OpenAI-compatible handler path where llm_provider is unavailable, instead of the deterministic body-limit rejection path the test intended to verify.

Changes

  • Reworked test_chat_completions_body_too_large to validate the real OpenAI-compatible route, auth middleware, and DefaultBodyLimit in-process
  • Kept the fix scoped to the regression itself instead of broadening it into unrelated test lifecycle cleanup
  • Preserved the intended behavioral assertion: oversized chat completion bodies return 413

Test Plan

  • cargo fmt --check
  • cargo clippy --all --benches --tests --examples --all-features
  • cargo test --test openai_compat_integration -- --test-threads=1 --nocapture
  • Repeated cargo test --test openai_compat_integration -- --test-threads=1 multiple times to confirm the flake no longer reproduces
  • Full cargo test passes except for a pre-existing SIGSEGV in tests/e2e_advanced_traces, which is unrelated to this change and reproducible on the upstream staging branch

Feature Parity

FEATURE_PARITY.md not updated; no tracked capability behavior changed.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added size: S 10-49 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: new First-time contributor labels Mar 10, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The stabilization approach (switching from a spawned TCP server to in-process oneshot) is sound and will eliminate the flaky 503 caused by server lifecycle non-determinism. The use of TestGatewayBuilder and direct Router construction is clean.

However, there is a correctness problem with the body limit value:

The test uses a 10 MB limit, but production uses 1 MB. The production server at src/channels/web/server.rs:354 applies DefaultBodyLimit::max(1024 * 1024) (1 MB), and the spec in src/channels/web/CLAUDE.md:200 confirms this. The test constructs its own router with DefaultBodyLimit::max(10 * 1024 * 1024), which means it is not testing the actual production configuration. It appears the old test had the same issue (the comment said "10 MB" and sent 11 MB), so this is a pre-existing bug being carried forward.

To actually regress against the production body limit, the test should use DefaultBodyLimit::max(1024 * 1024) (1 MB) and send a payload just over 1 MB (e.g., "x".repeat(1025 * 1024)). Alternatively, if 10 MB is the intended limit for the OpenAI-compat endpoint specifically, that should be reflected in the production router with a route-level override, and documented.

Everything else looks correct: auth middleware is wired, the handler and state types match production, and oneshot avoids all the TCP flakiness.

CLAUDE.md:200 still documented the pre-nearai#725 body limit of 1 MB, but
server.rs:354 was changed to 10 MB in nearai#725 (image upload support).
Update the documentation to match the actual production value.
@github-actions github-actions Bot added scope: channel/web Web gateway channel scope: docs Documentation risk: medium Business logic, config, or moderate-risk modules and removed risk: low Changes to docs, tests, or low-risk modules labels Mar 11, 2026
@CPU-216

CPU-216 commented Mar 11, 2026

Copy link
Copy Markdown
Contributor Author

Thank you for the prompt and thorough review! I dug into the source code and found that the discrepancy is actually in CLAUDE.md, not in the test:

  • server.rs:354 (production): DefaultBodyLimit::max(10 * 1024 * 1024) — 10 MB, with the comment // 10 MB max request body (image uploads)
  • CLAUDE.md:200 (documentation): still documented as DefaultBodyLimit::max(1024 * 1024) — 1 MB

Looking at the git history, #725 (feat: full image support across all channels) bumped the limit from 1 MB to 10 MB to support image uploads, but CLAUDE.md was not updated in that commit.

So the test's 10 MB limit + 11 MB payload is consistent with the actual production configuration. I've pushed a follow-up commit (af95b46) that fixes the stale documentation in CLAUDE.md to match the real value.

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The previous review feedback has been fully addressed. The core concern was a perceived mismatch between the test's 10 MB body limit and production. The second commit (af95b46) correctly identifies that production (server.rs:354) was already updated to 10 MB in #725 for image upload support -- it was the CLAUDE.md documentation that was stale, not the test.

Review of changes:

  1. Test stabilization (first commit): Switching from a spawned TCP server to in-process oneshot is the right fix. The flaky 503 was caused by server lifecycle non-determinism where the LLM provider appeared unavailable. The new test constructs the router directly with TestGatewayBuilder, wires auth middleware, applies DefaultBodyLimit::max(10 * 1024 * 1024) matching production, and uses oneshot -- no TCP, no sleeps, no flake vectors.

  2. Documentation fix (second commit): CLAUDE.md line 200 updated from "1 MB" to "10 MB" with a reference to #725. Matches server.rs:354 exactly.

No new issues found. Test correctly sends 11 MB to exceed the 10 MB limit and asserts 413.

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The previous review feedback has been fully addressed. The core concern was a perceived mismatch between the test's 10 MB body limit and production. The second commit (af95b46) correctly identifies that production (server.rs:354) was already updated to 10 MB in #725 for image upload support -- it was the CLAUDE.md documentation that was stale, not the test.

Review of changes:

  1. Test stabilization (first commit): Switching from a spawned TCP server to in-process oneshot is the right fix. The flaky 503 was caused by server lifecycle non-determinism where the LLM provider appeared unavailable. The new test constructs the router directly with TestGatewayBuilder, wires auth middleware, applies DefaultBodyLimit::max(10 * 1024 * 1024) matching production, and uses oneshot -- no TCP, no sleeps, no flake vectors.

  2. Documentation fix (second commit): CLAUDE.md line 200 updated from 1 MB to 10 MB with a reference to #725. Matches server.rs:354 exactly.

No new issues found. Test correctly sends 11 MB to exceed the 10 MB limit and asserts 413.

@zmanian
zmanian merged commit c372c99 into nearai:staging Mar 12, 2026
2 checks passed
@ironclaw-ci ironclaw-ci Bot mentioned this pull request Mar 12, 2026
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
)

* fix(test): stabilize openai compat oversized-body regression

* docs(web): fix stale body limit in CLAUDE.md (1 MB → 10 MB)

CLAUDE.md:200 still documented the pre-nearai#725 body limit of 1 MB, but
server.rs:354 was changed to 10 MB in nearai#725 (image upload support).
Update the documentation to match the actual production value.
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
)

* fix(test): stabilize openai compat oversized-body regression

* docs(web): fix stale body limit in CLAUDE.md (1 MB → 10 MB)

CLAUDE.md:200 still documented the pre-nearai#725 body limit of 1 MB, but
server.rs:354 was changed to 10 MB in nearai#725 (image upload support).
Update the documentation to match the actual production value.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: medium Business logic, config, or moderate-risk modules scope: channel/web Web gateway channel scope: docs Documentation size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants