Skip to content

test(api): lock chat developer role and multimodal content honesty - #405

Closed
seonghobae wants to merge 4 commits into
mainfrom
feat/chat-message-content-honesty-http-20260813223006
Closed

test(api): lock chat developer role and multimodal content honesty#405
seonghobae wants to merge 4 commits into
mainfrom
feat/chat-message-content-honesty-http-20260813223006

Conversation

@seonghobae

Copy link
Copy Markdown
Contributor

Summary

  • Real HTTP tests for chat Completions message honesty:
    • developer role → 400 invalid_message_role (use system)
    • multipart content (image_url, input_audio) → 400 invalid_message_content
    • non-string content → 400
  • Full tip substrate from bare main.

Test plan

  • python tests/test_chat_message_content_honesty.py (HTTP 5/5)
  • CI Full unit + Semgrep

Real HTTP coverage for fail-closed message validation: developer role
rejected with system migration; multipart image/audio content parts
rejected as unsupported on the text-only gateway. Ships tip substrate
from bare main.
@seonghobae
seonghobae enabled auto-merge (squash) August 13, 2026 13:30
@coderabbitai

coderabbitai Bot commented Aug 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@seonghobae, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 15 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 359ecbef-cb02-4ce4-8a30-1eafad219131

📥 Commits

Reviewing files that changed from the base of the PR and between 6841b71 and d508009.

📒 Files selected for processing (14)
  • contextual_orchestrator/cost_ledger.py
  • contextual_orchestrator/orchestrator.py
  • contextual_orchestrator/server.py
  • tests/test_analytics_runtime.py
  • tests/test_budget_enforcement.py
  • tests/test_chat_message_content_honesty.py
  • tests/test_chat_model_required_honesty.py
  • tests/test_commercial_readiness.py
  • tests/test_cost_review_server.py
  • tests/test_openai_passthrough.py
  • tests/test_sales_readiness.py
  • tests/test_security_hardening.py
  • tests/test_streaming.py
  • tests/test_true_streaming.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Align test_http_over_budget_returns_429 with fail-closed model-required
chat Completions so the request reaches budget enforcement (429) instead
of invalid_model (400).

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head d508009ec43e79b258130d1b2596374e779b5f13.

  • Head SHA: d508009ec43e79b258130d1b2596374e779b5f13

  • Workflow run: 31870693033

  • Workflow attempt: 1

Coverage evidence

Coverage evidence job did not run or did not publish coverage evidence.

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Changed file (3 files)"]
  S1 --> I1["repository behavior"]
  I1 --> R1["Review risk: Changed file (3 files)"]
  R1 --> V1["required checks"]
  Evidence --> S2["Test (11 files)"]
  S2 --> I2["regression suite"]
  I2 --> R2["Review risk: Test (11 files)"]
  R2 --> V2["targeted test run"]
Loading

@opencode-agent

Copy link
Copy Markdown
Contributor

OpenCode Review Overview

  • Head SHA: d508009ec43e79b258130d1b2596374e779b5f13
  • Workflow run: 31870693033
  • Workflow attempt: 1
  • Gate result: REQUEST_CHANGES (approval step)

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head d508009ec43e79b258130d1b2596374e779b5f13.

  • Head SHA: d508009ec43e79b258130d1b2596374e779b5f13

  • Workflow run: 31870693033

  • Workflow attempt: 1

Coverage evidence

Coverage evidence job did not run or did not publish coverage evidence.

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Changed file (3 files)"]
  S1 --> I1["repository behavior"]
  I1 --> R1["Review risk: Changed file (3 files)"]
  R1 --> V1["required checks"]
  Evidence --> S2["Test (11 files)"]
  S2 --> I2["regression suite"]
  I2 --> R2["Review risk: Test (11 files)"]
  R2 --> V2["targeted test run"]
Loading

@opencode-agent
opencode-agent Bot disabled auto-merge August 15, 2026 07:54

Copy link
Copy Markdown
Contributor Author

Closing as obsolete and superseded. PR #563 owns supported OpenAI-style multimodal message content, while cumulative PR #565 owns the remaining message validation contracts. This older reject-all multimodal slice would conflict with the canonical vision path. Reviews and checks remain head-bound.

@seonghobae seonghobae closed this Aug 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant