Skip to content

feat(canary): report Reborn inference cost - #5931

Closed
serrrfirat wants to merge 6 commits into
mainfrom
codex/reborn-canary-inference-cost
Closed

serrrfirat wants to merge 6 commits into
mainfrom
codex/reborn-canary-inference-cost

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • log structured Reborn model usage events from the runner without prompts, responses, or secrets
  • aggregate product and semantic-judge token usage into live QA case/lane artifacts
  • show run-wide and per-lane estimated inference cost in the Slack canary report, with unpriced calls called out explicitly

Change Type

  • Test/canary infrastructure
  • Observability/reporting
  • Product behavior change
  • Database/schema change
  • Security boundary change

Linked Issue

  • None.

Security Impact

  • No secrets, prompts, or full model responses are intentionally written to cost telemetry.
  • Semantic-judge failure artifacts store only sanitized verdict metadata and usage fields; response excerpts are excluded.
  • Existing artifact scrubbing remains in place.

Blast Radius

  • Reborn WebUI v2 live canary runner and Slack reporter.
  • Reborn model gateway diagnostic logging for model usage events.
  • No canary prompts, pass/fail criteria, model routing, or user-facing product workflows are changed.

Rollback Plan

  • Revert this PR to remove inference-cost telemetry from live canary artifacts and Slack reports.
  • Canary pass/fail behavior should return to the prior reporting-only state because prompts and assertions are unchanged.

Reborn Trust-Boundary Checklist

  • No new trusted ingress path.
  • No auth, OAuth, secret storage, or outbound delivery permission change.
  • No prompt/response content is required for cost aggregation.
  • Live-secret usage remains limited to the existing canary workflow gates.

Notes

  • Cost is an estimate from provider-reported token usage and configured provider rates; calls without usage remain visible as unpriced.
  • Reporter-side Slack enrichment cost is not included in the Reborn inference estimate.

Verification

  • python3 -m unittest scripts.reborn_webui_v2_live_qa.test_run_live_qa scripts.live-canary.test_notify_slack
  • cargo fmt --check -- crates/ironclaw_runner/src/model_gateway.rs crates/ironclaw_llm/src/nearai_chat.rs
  • cargo test -p ironclaw_runner model_gateway --lib
  • cargo build --profile dist --package ironclaw_reborn_cli --features webui-v2-beta,slack-v2-host-beta,libsql,postgres,inmemory-turn-state --bin ironclaw-reborn
  • git diff --check

Known unrelated check

  • cargo test -p ironclaw_runner model_gateway --lib emits existing dead-code warnings under crates/ironclaw_runner/src/subagent; this branch does not touch those files.
  • python3 -m ruff check was not runnable locally because ruff is not installed in the system Python environment.

@ironloopai

ironloopai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 90d0cbc27b090fb08b0ca81ca69d78ca7680c846
Result: No reviewer jobs are scheduled yet.
Next: Run @ironloopai review to start reviewers.
Updated: 2026-07-10T16:54:35.462Z

Current reviewers:

Reviewer State Verdict Findings Last update
none Queued N/A No reviewer jobs scheduled yet. N/A
Reviewer summaries
Reviewer Detail
none No reviewer jobs scheduled yet.
Recent activity
Time Reviewer State Detail
N/A N/A Waiting No progress events recorded yet.
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
  • @ironloopai status
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 11:02 Destroyed
@github-actions github-actions Bot added scope: docs Documentation size: XL 500+ changed lines labels Jul 10, 2026
@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds model inference usage pricing and tracing, aggregates priced and unpriced calls into live QA results, and exposes per-case and cross-lane token and cost summaries in Slack notifications.

Changes

Inference usage telemetry

Layer / File(s) Summary
Provider pricing fallback
crates/ironclaw_llm/src/costs.rs, crates/ironclaw_llm/src/nearai_chat.rs, crates/ironclaw_runner/Cargo.toml
Adds structured model rates, DeepSeek cache pricing, remote fallback pricing, explicit free-model handling, and runtime decimal support.
Gateway usage tracing
crates/ironclaw_runner/src/model_gateway.rs
Completion paths trace token counts, cache usage, estimated cost, or unavailable usage for tool, repair, and text requests.
Live QA usage aggregation
scripts/reborn_webui_v2_live_qa/run_live_qa.py, scripts/reborn_webui_v2_live_qa/semantic_judge.py, scripts/reborn_webui_v2_live_qa/test_run_live_qa.py, .github/workflows/live-canary.yml
Semantic-judge responses and server logs are sanitized and aggregated into per-case and run-level inference_usage data written to results.json, with workflow preparation-job gating.
Slack usage reporting
scripts/live-canary/notify_slack.py, scripts/live-canary/test_notify_slack.py, scripts/live-canary/README.md
Lane and case reports parse inference metrics, and Slack renders group and cross-lane token, cost, and unpriced-call summaries.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant ModelGateway
  participant SemanticJudge
  participant LiveQaRunner
  participant ResultsJson
  participant SlackNotifier
  ModelGateway->>LiveQaRunner: emit model usage events
  SemanticJudge->>LiveQaRunner: return inference_usage
  LiveQaRunner->>ResultsJson: persist case and aggregate usage
  ResultsJson->>SlackNotifier: provide inference_usage
  SlackNotifier->>SlackNotifier: render lane and case summaries
Loading

Possibly related issues

Possibly related PRs

Suggested labels: scope: dependencies

Poem

I’m a rabbit counting tokens bright,
Pricing each hop by lantern light.
Free models glow, cache rates hum,
Unpriced calls say where they’re from.
Slack shares the totals—hare-brained fun!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately summarizes the main change: reporting Reborn inference cost in canary workflows.
Description check ✅ Passed The description covers the required sections well, including summary, change type, security impact, rollback, trust-boundary, notes, and verification.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 10, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements telemetry for tracking LLM provider usage, token counts, and estimated USD costs across Reborn and semantic-judge inferences, updating the Slack notification and test reporting systems to aggregate and display these metrics. A critical issue was identified in run_live_qa.py where summing Decimal values without an explicit Decimal(0) start value will raise a TypeError at runtime.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

"input_tokens": sum(_non_negative_int(item.get("input_tokens")) for item in case_summaries),
"output_tokens": sum(_non_negative_int(item.get("output_tokens")) for item in case_summaries),
"estimated_usd": str(
sum((_decimal(item.get("estimated_usd")) or Decimal(0)) for item in case_summaries)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

In Python, summing Decimal objects using the built-in sum() function without an explicit Decimal start value will raise a TypeError: unsupported operand type(s) for +: 'int' and 'decimal.Decimal'. This is because sum() defaults to a start value of 0 (an integer), and Python does not allow implicit addition of int and Decimal in this context.

To fix this, provide Decimal(0) as the second argument to sum() to initialize the accumulator correctly.

Suggested change
sum((_decimal(item.get("estimated_usd")) or Decimal(0)) for item in case_summaries)
sum((_decimal(item.get("estimated_usd")) or Decimal(0) for item in case_summaries), Decimal(0))

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 2 2 7384b476bb69

Head: 7384b476bb69046637e432587206e8631336513b
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

No blocking runtime or security issues found. I found two non-blocking telemetry accounting gaps in the new inference-cost reporting path.

Findings

Blocking: 0 / Notes: 2

Non-blocking notes (2)
1. 💬 [LOW] Usage telemetry is enabled on an env var Reborn does not read

Location: scripts/reborn_webui_v2_live_qa/run_live_qa.py:739-745
ironclaw-reborn serve initializes tracing from IRONCLAW_REBORN_LOG / IRONCLAW_REBORN_OPERATOR_LOG, not RUST_LOG. Appending ironclaw_runner::model_usage=info to RUST_LOG here does not force the new REBORN_INFERENCE_USAGE events when the Reborn log filter is tightened, so an environment with IRONCLAW_REBORN_LOG=warn will under-report product inference usage. Set or append the target on the Reborn log env var for the child process, and update the test to cover that variable.

2. 💬 [LOW] Failed semantic-judge verdicts are omitted from usage totals

Location: scripts/reborn_webui_v2_live_qa/run_live_qa.py:4546-4549
_case_inference_usage() only searches result.details for semantic-judge usage. When _wait_for_assistant_reply() calls the judge and the judge returns a non-passing verdict, that payload is only embedded in the assertion text; _live_chat_case() then builds failure details from observed, which never receives the judge payload. Those judge calls are therefore missing from per-case and top-level inference_usage on failed canary cases. Persist the judge payload or its inference_usage into the failure details before raising, or carry it via a typed exception.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.
  4. Use @ironloopai status to check queued/running/completed/failed/superseded state while reviewers run.

"RUST_LOG",
"ironclaw=warn,ironclaw_reborn=warn,ironclaw_reborn_webui_ingress=info",
)
if "ironclaw_runner::model_usage" not in rust_log:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This appends the usage target to RUST_LOG, but ironclaw-reborn serve reads IRONCLAW_REBORN_LOG for its stderr tracing filter. If that Reborn log filter is set to warn, the new usage events will still be dropped and product inference usage will report as zero.


def _case_inference_usage(output_dir: Path, result: ProbeResult) -> dict[str, object]:
events = _product_inference_usage(output_dir)
events.extend(_semantic_judge_usage(result.details))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This only finds semantic-judge usage that made it into result.details. If the judge runs and returns a failing verdict, _wait_for_assistant_reply() raises with the judge payload only in the assertion string, so the judge inference is not counted for failed cases.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/live-canary/notify_slack.py`:
- Around line 254-259: Update _non_negative_decimal to reject non-finite Decimal
values before comparing or returning: after parsing, check parsed.is_finite()
and return Decimal(0) for NaN or Infinity; retain the existing handling for
parse errors and negative values.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 41e41b9f-d3f5-4da5-96f9-494e19ef49c9

📥 Commits

Reviewing files that changed from the base of the PR and between 5ae88dd and 7384b47.

📒 Files selected for processing (7)
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/README.md
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread scripts/live-canary/notify_slack.py Outdated
@github-actions

github-actions Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.2% (290574 / 341066 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 341066 lines now vs 320188 at floor capture (+20878 lines, +6.52%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.2% — 290574 / 341066 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 52.36% 222 / 424
ironclaw_authorization 53.66% 462 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_reborn_cli 62.5% 3766 / 6026
ironclaw_webui_v2 62.62% 2632 / 4203
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_migration 66.93% 1168 / 1745
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_filesystem 67.25% 3833 / 5700
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_first_party_extensions 77.66% 5400 / 6953
ironclaw_llm 78.39% 20370 / 25984
ironclaw_product_context 78.57% 11 / 14
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_reborn_openai_compat 81.23% 978 / 1204
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_wasm 82.72% 996 / 1204
ironclaw_events 82.86% 1765 / 2130
ironclaw_auth 83.87% 3078 / 3670
ironclaw_turns 84.31% 13392 / 15884
ironclaw_reborn_config 84.33% 1814 / 2151
ironclaw_processes 84.44% 993 / 1176
ironclaw_host_api 84.9% 2608 / 3072
ironclaw_product_workflow 85.19% 10762 / 12633
ironclaw_threads 85.88% 4226 / 4921
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_common 86.46% 1514 / 1751
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_product_adapters 86.98% 3207 / 3687
ironclaw_reborn_identity 87.03% 557 / 640
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.77% 9914 / 11296
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_extensions 88.35% 2638 / 2986
ironclaw_host_runtime 88.45% 17056 / 19283
ironclaw_reborn_composition 88.96% 74433 / 83673
ironclaw_approvals 89.24% 1584 / 1775
ironclaw_runner 89.29% 16769 / 18781
ironclaw_conversations 90.33% 3120 / 3454
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_support 92.49% 14749 / 15946
ironclaw_resources 92.81% 4722 / 5088
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.63% 8811 / 9311
ironclaw_safety 94.8% 3668 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@serrrfirat
serrrfirat force-pushed the codex/reborn-canary-inference-cost branch from 7384b47 to 59bd4b3 Compare July 10, 2026 11:13
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 11:13 Destroyed
@railway-app

railway-app Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5931 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 10, 2026 at 4:54 pm

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Line 51: Change the model usage telemetry helpers trace_model_usage and
trace_unpriced_model_call from info! to debug! while preserving the
REBORN_INFERENCE_USAGE target and all call sites. Update run_live_qa.py’s
_log_filter_with_model_usage filter directive from =info to =debug so the
harness matches the new logging level.
- Around line 1382-1389: Update the `trace_model_usage` call for the
`provider_complete` event to pass `response.cache_read_input_tokens` and
`response.cache_creation_input_tokens` instead of hardcoded zeros, matching the
tool-call usage logging path.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c644d507-ccf3-410b-8ae5-8e0e14ffe189

📥 Commits

Reviewing files that changed from the base of the PR and between 7384b47 and 59bd4b3.

📒 Files selected for processing (7)
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/README.md
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread crates/ironclaw_runner/src/model_gateway.rs Outdated
Comment thread crates/ironclaw_runner/src/model_gateway.rs
@serrrfirat
serrrfirat force-pushed the codex/reborn-canary-inference-cost branch from 59bd4b3 to 118b9fe Compare July 10, 2026 11:28
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 11:28 Destroyed
@github-actions github-actions Bot added scope: ci CI/CD workflows risk: medium Business logic, config, or moderate-risk modules and removed risk: low Changes to docs, tests, or low-risk modules labels Jul 10, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/reborn_webui_v2_live_qa/semantic_judge.py`:
- Around line 194-216: Update the token extraction logic in the response-usage
parsing function: require both prompt_tokens and completion_tokens to be present
integer values greater than or equal to zero, and return None when either is
missing or invalid. Avoid defaulting absent provider counts to zero so
downstream aggregation treats incomplete usage as unpriced; retain valid
cached_tokens handling.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 491bbd37-a8e3-4a60-885d-08bc3851bab6

📥 Commits

Reviewing files that changed from the base of the PR and between 59bd4b3 and 118b9fe.

📒 Files selected for processing (8)
  • .github/workflows/live-canary.yml
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/README.md
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread scripts/reborn_webui_v2_live_qa/semantic_judge.py
@serrrfirat
serrrfirat force-pushed the codex/reborn-canary-inference-cost branch from 118b9fe to 3bb49c6 Compare July 10, 2026 11:38
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 11:38 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Around line 73-98: Update trace_model_usage and its callers to accept the
effective request model name and input/output token rates captured at request
start, rather than reading provider.active_model_name() or
provider.cost_per_token() after completion. In NearAiChatProvider::complete,
resolve req.take_model_override() and the corresponding pricing before
performing the request, then pass these captured values through the
usage-tracing path so concurrent switches cannot alter attribution.

In `@scripts/live-canary/test_notify_slack.py`:
- Around line 390-400: Extend the test around the existing QA and cross-lane
Slack assertions to cover a payload with unpriced_call_count > 0. Assert that
the QA-group text includes the required unpriced-call marker and that the
cross-lane context text includes its corresponding unpriced marker, preserving
the existing priced-call assertions.

In `@scripts/reborn_webui_v2_live_qa/run_live_qa.py`:
- Around line 4500-4511: Update _record_assistant_reply_wait_result() and
routine_confirmation_follow_up handling so every semantic-judge usage event is
preserved instead of overwriting observed["semantic_judge"]; store payloads as a
list or aggregate calls immediately, then update _semantic_judge_usage() to
traverse and count all preserved payloads, including follow-up calls, tokens,
and cost.
- Around line 1317-1332: The failure handling around `_result` and the
corresponding assertion path must not persist the raw `semantic_judge` payload.
Replace the full `semantic_judge` dict with only safe usage fields and redacted
verdict metadata, excluding `response_excerpt` and any prompt/response content,
and ensure assertion text uses the sanitized representation rather than
embedding the original payload.
- Around line 739-750: Use the merged environment mapping when deriving log
filters: in the code that assigns rust_log and reborn_log, replace both
os.environ.get calls with env.get calls so extra_env overrides for RUST_LOG and
IRONCLAW_REBORN_LOG are honored by the child process.
- Around line 4434-4439: Update _decimal to reject non-finite Decimal values
before the nonnegative comparison: after parsing inside the try block, check
parsed.is_finite() and return None when false, then perform the existing >= 0
validation.

In `@scripts/reborn_webui_v2_live_qa/test_run_live_qa.py`:
- Around line 1257-1298: The test’s helper functions, nested exception context,
and dynamic attribute access trigger Ruff warnings. In
test_wait_for_assistant_reply_attaches_failed_semantic_judge_to_error, annotate
fake_judge and fake_sleep parameters and return types, flatten the nested
assertRaisesRegex context using the repository’s preferred unittest style or a
targeted exemption, and replace getattr on raised.exception with direct
semantic_judge access (or an approved static alternative).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: da7736ac-50fb-40c7-9537-3f4e20c54d31

📥 Commits

Reviewing files that changed from the base of the PR and between 118b9fe and 3bb49c6.

📒 Files selected for processing (9)
  • .github/workflows/live-canary.yml
  • crates/ironclaw_llm/src/nearai_chat.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/README.md
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread crates/ironclaw_runner/src/model_gateway.rs Outdated
Comment thread scripts/live-canary/test_notify_slack.py
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
Comment thread scripts/reborn_webui_v2_live_qa/test_run_live_qa.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Around line 74-100: Remove the duplicate is_explicit_free_model predicate from
model_gateway.rs and reuse a shared helper from the owning ironclaw_llm crate,
exposing it if necessary. Update fallback_usage_rates_for_model and related call
sites to use that helper, preserving the existing :free, openrouter/free, and
free handling in one centralized implementation.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 57358bc8-64f7-4488-a93d-29ec10c3f4b8

📥 Commits

Reviewing files that changed from the base of the PR and between 3bb49c6 and 5f0452c.

📒 Files selected for processing (6)
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread crates/ironclaw_runner/src/model_gateway.rs Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 13:38 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 14:48 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
scripts/reborn_webui_v2_live_qa/test_run_live_qa.py (1)

1420-1435: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert that extra_env overrides inherited log filters.

These assertIn checks pass even if server_env accidentally preserves the process-level RUST_LOG or IRONCLAW_REBORN_LOG values instead of applying extra_env precedence. Seed sentinel inherited filters and assert they are absent, or compare the complete expected values while retaining the telemetry directive.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/reborn_webui_v2_live_qa/test_run_live_qa.py` around lines 1420 -
1435, Strengthen test_server_env_honors_extra_env_log_filters by seeding
sentinel inherited RUST_LOG and IRONCLAW_REBORN_LOG values before calling
server_env, then assert those sentinels are absent and the resulting values
exactly reflect extra_env plus the required ironclaw_runner::model_usage=debug
directive. This verifies extra_env precedence rather than only checking
substring presence.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_llm/src/costs.rs`:
- Around line 77-79: Normalize model IDs to lowercase in model_cost() before the
pricing match, or make the DeepSeek-V4-Flash match case-insensitive, so
mixed-case identifiers resolve to the configured DeepSeek rates rather than the
zero-cost fallback used by cache_read_input_rate_for_model().

In `@crates/ironclaw_runner/src/model_gateway.rs`:
- Around line 103-114: Move cache-read and base pricing ownership into a shared
ironclaw_llm::costs API that returns input, output, and optional cache-read
rates for a model, then update cache_read_input_rate_for_model and
locked_usage_rates_for_model to use it instead of hardcoded model checks and
direct model_cost calls. Reuse the existing nearai_chat.rs
remote_model_fallback_cost behavior so free-model handling remains consistent,
and remove the duplicated deepseek-v4-flash pricing logic from model_gateway.rs.

In `@scripts/reborn_webui_v2_live_qa/run_live_qa.py`:
- Around line 4600-4610: Use an explicit None check when assigning
cache_read_input_rate in the cost calculation: select rates[2] whenever it is
not None, including Decimal("0"), and fall back to rates[0] only when it is
unset. Keep the behavior consistent with the existing rates[2] is not None
check.

---

Outside diff comments:
In `@scripts/reborn_webui_v2_live_qa/test_run_live_qa.py`:
- Around line 1420-1435: Strengthen test_server_env_honors_extra_env_log_filters
by seeding sentinel inherited RUST_LOG and IRONCLAW_REBORN_LOG values before
calling server_env, then assert those sentinels are absent and the resulting
values exactly reflect extra_env plus the required
ironclaw_runner::model_usage=debug directive. This verifies extra_env precedence
rather than only checking substring presence.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 3bb56e33-78c9-4a72-b157-1e62d010de1c

📥 Commits

Reviewing files that changed from the base of the PR and between bc7a695 and 0a1b328.

📒 Files selected for processing (5)
  • crates/ironclaw_llm/src/costs.rs
  • crates/ironclaw_llm/src/nearai_chat.rs
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread crates/ironclaw_llm/src/costs.rs
Comment thread crates/ironclaw_runner/src/model_gateway.rs Outdated
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
@serrrfirat
serrrfirat force-pushed the codex/reborn-canary-inference-cost branch from 0a1b328 to f8280ad Compare July 10, 2026 15:30
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 15:30 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_llm/src/costs.rs`:
- Around line 170-181: The test
test_deepseek_v4_flash_uses_remote_pricing_before_local_heuristic only checks
two casing variants; add a mixed-case assertion using
model_cost("deepseek-ai/Deepseek-v4-Flash") and verify it does not return
Some((ZERO, ZERO)), preserving the regression coverage.

In `@crates/ironclaw_llm/src/nearai_chat.rs`:
- Around line 1207-1218: Centralize the remote model fallback pricing logic in a
single public function in ironclaw_llm::costs, including the shared
explicit-free-model check and zero-cost guard. Replace nearai_chat.rs functions
remote_model_fallback_cost and is_explicit_free_model, the inline logic in
costs.rs, and model_gateway.rs functions fallback_usage_rates_for_model and
is_explicit_free_model with calls to this shared function, removing duplicate
implementations.

In `@scripts/reborn_webui_v2_live_qa/run_live_qa.py`:
- Around line 6981-7007: Malformed priced-usage records are being discarded
instead of counted as unpriced. In the event-building logic, replace the
continue triggered when input_rate, output_rate, or estimated_usd is missing or
invalid with creation of the same unpriced event shape used by the
usage_available=="false" branch, preserving operation, model, token counts, and
unpriced accounting before appending it.
- Around line 1214-1250: The _safe_semantic_judge_payload function still
persists free-text judge reasoning that may contain response content. Remove the
reason field from the sanitized payload, or replace it with a bounded,
predefined classification value; do not persist the model-generated text while
retaining the existing structured fields and usage sanitization.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: fb3939ab-f779-40fe-9beb-69b622b7b06d

📥 Commits

Reviewing files that changed from the base of the PR and between 0a1b328 and f8280ad.

📒 Files selected for processing (11)
  • .github/workflows/live-canary.yml
  • crates/ironclaw_llm/src/costs.rs
  • crates/ironclaw_llm/src/nearai_chat.rs
  • crates/ironclaw_runner/Cargo.toml
  • crates/ironclaw_runner/src/model_gateway.rs
  • scripts/live-canary/README.md
  • scripts/live-canary/notify_slack.py
  • scripts/live-canary/test_notify_slack.py
  • scripts/reborn_webui_v2_live_qa/run_live_qa.py
  • scripts/reborn_webui_v2_live_qa/semantic_judge.py
  • scripts/reborn_webui_v2_live_qa/test_run_live_qa.py

Comment thread crates/ironclaw_llm/src/costs.rs
Comment thread crates/ironclaw_llm/src/nearai_chat.rs Outdated
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
Comment thread scripts/reborn_webui_v2_live_qa/run_live_qa.py
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 16:42 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5931 July 10, 2026 16:54 Destroyed
@serrrfirat serrrfirat closed this Jul 13, 2026

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5931 — 90d0cbc2 Deployed Jul 10, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant