Skip to content

feat(proxy): add native ROI calculator for gateway spend vs merged PRs - #43669

Merged
moe-berri merged 29 commits into
mainfrom
litellm_roi_calculator_v2
Sep 30, 2026
Merged

moe-berri merged 29 commits into
mainfrom
litellm_roi_calculator_v2

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Compare gateway spend with estimated merged-PR effort inside LiteLLM
  • Preserve the standalone calculator's simple setup and scheduled updates

How it solves it:

  • Three steps: GitHub, repositories, estimator and schedule
  • Daily reruns, persistent estimate reuse, progress and cancellation
  • Overview, People, email matching, CSV export and sample preview
  • Admin permissions, encrypted credentials and normal gateway inference accounting

Intentional product change: the calculator runs inside the gateway, using a GitHub token instead of the standalone GitHub App. Full gateway admins configure it; read-only admins view results. Team admins and regular users cannot access the calculator

Scope: ROI calculator only. Dependency manifests, lockfiles, Google API tests, and shared encryption behavior match main. Includes main through 50f5cc9bbb

User Flow

Before: the initial native implementation exposes all settings together and requires manual refreshes

  1. Open /roi-calculator/ as a gateway admin
  2. Enter GitHub connection, repositories, model and prompt together
  3. Save settings and start analysis
  4. Return later and manually refresh the report

Original native setup

After: guided setup leads to the first report and scheduled updates

  1. Open /roi-calculator/ as a gateway admin
  2. Connect GitHub, select repositories, then choose an estimator
  3. Keep daily updates or select manual updates, then start backfill
  4. Review Overview, inspect PR reasoning, and match emails in People
  5. Return to automatically refreshed results; change preferences in Settings

Guided setup

Parity with the standalone calculator

Capability Gateway implementation
Calculation Same matched-spend / estimated-hours calculation and cohort exclusions
Estimator Same default prompt, temperature 0, metadata-only evidence and three concurrent PRs
Initial window and schedule Seven-day backfill, daily refresh, manual option, minimum five-minute schedule
Reuse Successful unchanged PRs reused before details or inference; failed estimates retried
Progress Elapsed time, estimated remaining time, reused count and cancellation
Reporting Overview, People, calculation details, PR reasoning and CSV export
Email matching GitHub profile/commit emails plus manual matches to observed gateway identities
Setup tools Connection checks, sample preview, advanced prompt/key options and restart setup

Intentional differences:

  • The gateway connection is already available, so no separate URL/key setup or standalone service is needed
  • GitHub token authentication replaces the GitHub App creation/install flow; Enterprise API URLs remain supported
  • Gateway PostgreSQL replaces local SQLite/files, and secrets are encrypted using gateway secret storage
  • Shared database leases coordinate multiple gateway workers, including cancellation and recovery after interruption
  • Gateway admin roles protect reports and settings; the standalone requires an external access-control layer
  • Superseded cache versions are removed after successful re-estimation; estimates outside the selected date/repository scope are retained
  • One unreadable PR is marked for review while other results publish; a complete metadata outage preserves the last report, and failed metadata is retried
  • Unlike standalone, a repository outage permits an explicitly incomplete report from healthy repositories, with ratios withheld; complete outages preserve the previous report
  • Profile emails refresh for reused PRs without another estimate; transient lookup failures retain known addresses, while confirmed email removal clears them
  • Total estimator failure and partial repository outages with no usable PRs preserve the previous report

Screenshot walkthrough

Screenshots use an isolated local gateway, two real public PRs from BerriAI/litellm-admin-agent, and paid estimator calls. engineer@example.com is a synthetic QA identity used only to demonstrate matching

Setup: GitHub, repositories, estimator, first backfill

Connect GitHub
Choose repositories
Estimator and schedule
Backfill progress

Overview, pull requests and estimate reasoning

Overview
Pull requests
Estimate reasoning

People: before matching, matching dialog, matched result

Unmatched people
Match email
Matched people

Settings, advanced options, restart confirmation and sample preview

Settings
Advanced settings
Restart setup confirmation
Sample preview

Matching calculator icons in the sidebar and page heading

Matching calculator icons

Screenshots / Proof of Fix

Shared setup: local gateway on port 4036, dashboard on port 3036, isolated PostgreSQL database, real public GitHub PRs, roi-estimator routed to paid gpt-6-luna. Authentication below uses the local admin key through an environment variable, never a committed secret

Before (1afcc91, initial PR implementation)

  1. Open /roi-calculator/ with no saved configuration: all setup fields appear together, shown above
  2. Configure and analyze the public repository: a manual refresh is required for another run; the initial implementation has no scheduled refresh setting

After (5d642e6)

  1. Complete the guided setup shown above; the live report contains saved estimates of 36 and 32 hours from paid gateway calls
  2. Send curl -X POST http://localhost:4036/roi-calculator/connections/test -H "Authorization: Bearer $ROI_ADMIN_KEY": HTTP 200
  3. Send curl -X POST http://localhost:4036/roi-calculator/sync -H "Authorization: Bearer $ROI_ADMIN_KEY", then GET the same URL: {"phase":"complete","total":2,"estimated":2,"reused":2,"needs_attention":0}
  4. Set a five-minute schedule and age the isolated test database's last-run timestamp: the real gateway timer starts a run without a button press, reusing both estimates; restore the daily schedule afterward
  5. Compare the saved real report with both standalone and native calculation functions: every metric matches exactly with automatic matching and with a manual QA mapping
  6. Test a restricted inference key: connection validation returns HTTP 409 and unauthorized estimates fail instead of using admin privileges; restore the valid key and complete both estimates
  7. Add BerriAI/roi-qa-missing-repository alongside the real public repository in the isolated QA configuration and start analysis: two existing PR estimates remain, the unavailable repository is named, and cost-per-hour is absent
  8. Remove that unavailable repository and rerun: both estimates are reused, the warning clears, and the ratio returns
  9. Select the empty BerriAI/litellm-roi-calculator repository alongside the unavailable QA repository and analyze: an error explains that no report was published, and both previous estimates remain; restore the original repositories
  10. In isolated QA, use an invalid estimator key with a changed prompt and run analysis: the authentication failure preserves the previous report; restore the original prompt and key, then rerun successfully using both cached estimates
  11. Send PUT /roi-calculator/identity-map with {"github_login":"invalid.name","email":"engineer@example.com"}: HTTP 422, with settings unchanged

Incomplete report during a real GitHub 404

Empty partial outage preserves report
Incomplete calculation explanation
Estimator outage preserves report

Additional regression checks: 77 ROI backend/database tests and 15 ROI UI tests passed using the unchanged dependencies from main. Real PostgreSQL coverage verifies scope-preserving cache cleanup, exclusive sync ownership, writer routing with a separate reader, the scheduled interval gate, expired-lease fencing, and UTC timestamp recovery. A repeated-startup regression also verifies the scheduler retains exactly one ROI job; it fails with the original duplicate-job error before the fix. Full make check passes

Validation on 5d642e6211: all required CI checks pass, Codecov patch coverage passes at 84.13% against 82.51%, Greptile is 5/5 with no outstanding findings, and Veria reports no security issues. All reported review findings have tested fixes

Outage regressions cover unavailable repositories, empty partial results, estimator failures, profile lookup failures versus confirmed email removal, and recovery. Refreshed identities persist across restarts; unchanged identities skip database writes. Live PostgreSQL row versions confirm no cache writes during an unchanged rerun

Bugbot found no issues on deb9d2d5e5; its final rerun is unavailable because the team spending limit was reached. The only subsequent change skips unchanged cache writes and is covered by regression tests and live database checks

Validation limitation: Python CodeQL fails because Security/CWE-117/LogInjection.ql exceeds its 2 GiB result-set limit on both this PR and main. This is an incomplete scan, not a reported security finding; scan configuration and protections are unchanged. Required checks pass on 5d642e6211

OSV baseline: a fresh scan of main at 264b09ac8d reproduces the same five findings as this PR, affecting the existing PyJWT, urllib3 and Next.js versions. All scanned lockfiles and scan configuration are identical to main; dependency upgrades remain outside this ROI-only change

Unrelated CI baseline: the unchanged Google API compliance tests fail against upstream schema changes. Those tests and dependencies remain identical to main

Pre-Submission checklist

  • Added meaningful behavioral and real-database regression tests
  • Focused backend and UI tests pass locally
  • Latest required CI checks pass
  • Changes are scoped to the ROI calculator, gateway integration, tests and screenshots
  • Latest Greptile confidence is at least 4/5

Caveats

Medium

  • GitHub App onboarding is replaced by the accepted token flow
  • Scheduled updates require a running gateway and available database
  • Partial repository outages hide ratios until access recovers
  • Authorized admins share the connection and repository report visibility

Low

  • Estimated effort is not measured savings or financial return
  • Source patches and personal emails are excluded from estimator evidence
  • PR descriptions and commit messages may themselves contain code
  • Exact model outputs can vary despite temperature 0
  • CSV content is tested; native browser download completion remains unverified

Type

New feature, bug fixes and regression tests

Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 29, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ moe-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds ROI calculator feature with GitHub integration and spend analysis.

The PR appears safe to merge based on this review; the latest cache-write change leaves no outstanding finding.

Summary

The PR adds a gateway-native ROI calculator with guided setup, scheduled analysis, persistent estimate reuse, reporting, and admin-only access.

  • The latest change avoids rewriting saved PR records when a reused PR’s identity has not changed.
  • Changed identities continue to be persisted, including confirmed email removal.

Reviews (18) · Last reviewed commit: "perf(roi): skip writes for unchanged cac..."

Comment thread litellm/proxy/roi_calculator/github.py Outdated
Comment thread litellm/proxy/roi_calculator/estimator.py Outdated
Comment thread litellm/proxy/management_endpoints/roi_calculator_endpoints.py Outdated
Comment thread litellm/proxy/roi_calculator/sync.py Outdated
Comment thread litellm/proxy/roi_calculator/pull_cache.py
Comment thread ui/litellm-dashboard/src/app/(dashboard)/roi-calculator/page.tsx Outdated
Comment thread litellm/proxy/roi_calculator/sync.py Outdated
Comment thread litellm/proxy/roi_calculator/analytics.py Fixed
Comment thread litellm/proxy/roi_calculator/sync.py Fixed
@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_roi_calculator_v2 (5d642e6) with main (50f5cc9)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (657bb77) during the generation of this report, so 50f5cc9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

devin-ai-integration Bot and others added 8 commits September 29, 2026 06:29
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
@moe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@moe-berri

Copy link
Copy Markdown
Contributor

@greptileai

@moe-berri

Copy link
Copy Markdown
Contributor

@veria-ai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread litellm/proxy/roi_calculator/github.py Outdated
Co-Authored-By: Ishaan Jaffer <155045088+ishaan-berri@users.noreply.github.com>
@moe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@moe-berri

Copy link
Copy Markdown
Contributor

@greptileai

@moe-berri

Copy link
Copy Markdown
Contributor

@veria-ai

Comment thread litellm/proxy/roi_calculator/sync.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@moe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@moe-berri

Copy link
Copy Markdown
Contributor

@greptileai

@moe-berri

Copy link
Copy Markdown
Contributor

@veria-ai

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit deb9d2d. Configure here.

Comment thread litellm/proxy/roi_calculator/sync.py Outdated
@moe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@moe-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@moe-berri

Copy link
Copy Markdown
Contributor

@veria-ai

@moe-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@moe-berri
moe-berri merged commit 6b9766f into main Sep 30, 2026
102 of 104 checks passed
@moe-berri
moe-berri deleted the litellm_roi_calculator_v2 branch September 30, 2026 21:16
yuneng-berri added a commit that referenced this pull request Oct 1, 2026
…route sweep

GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
yuneng-berri added a commit that referenced this pull request Oct 1, 2026
…fixtures (#43958)

* test(ci): repair stale request fakes, spend-log golden, auto-router labels, and Interactions spec lookups

Request fakes now carry the scope a real Starlette request has, the GCS pub/sub
spend-log golden gains the agent identity keys from #43722, the auto-router
session tests follow the baseline_models contract from #43348, and the
Interactions spec checks resolve the create body and resource paths from the
live spec instead of hardcoded names

* test(ci): move retired OpenAI text-completion fixtures to live vehicles

OpenAI still serves native /v1/completions on the gpt-5.4 family, so the
single-prompt cases move to text-completion-openai/gpt-5.4-nano. Multi-prompt
batches and echo with logprobs now 500 on every OpenAI model, so those cases
keep the same text-completion-openai transport pointed at Fireworks, which
documents both. The optional-params test asserts the request body actually
sent instead of a success callback whose assertions were swallowed

* test(ci): use a serverless Fireworks model for the text-completion batch and echo cases

gpt-oss-20b is on-demand only on Fireworks, so the CI key got 404 model not
deployed; glm-5p3-flash is listed as serverless

* test(ci): skip the ROI calculator repository listing in the security route sweep

GET /roi-calculator/repositories (#43669) lists repositories from the configured
GitHub API, api.github.com by default, so the S2 sweep's GET of every route made
the owned proxy reach an external host and failed the egress check in 31
integration-security tests. It joins /get/latest_release_info in the deny list
@moe-berri moe-berri mentioned this pull request Oct 10, 2026
3 of 4 tasks

This branch was successfully deployed

1 active deployment
e2e-changed — 5d642e62 Deployed Sep 30, 2026 by moe-berri via oauth #2071
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants