Skip to content

fix:optimize quota usage for Copilot and metered models - #1662

Closed
Gustave-241021 wants to merge 3 commits into
charmbracelet:mainfrom
Gustave-241021:fix/reduce-copilot-quota
Closed

Gustave-241021 wants to merge 3 commits into
charmbracelet:mainfrom
Gustave-241021:fix/reduce-copilot-quota

Conversation

@Gustave-241021

Copy link
Copy Markdown
Contributor

Root Cause Analysis

GitHub Copilot quotas are counted by requests (turns). Since crush uses an agentic architecture, it iterates through several turns to solve a problem:

  1. Agent Loop: Every turn of "Think -> Call Tool -> Get Result -> Think Again" consumes one request.
  2. Helper Tasks: Auto-generating session titles (generateTitle()) and auto-summarizing (Summarize()) consume additional requests.
  3. Tool Call Granularity: If the model tends to call tools sequentially rather than in batch, the number of interaction turns increases significantly.

Proposed Design

1. Metered Provider Identification

Add an isMetered() helper function in agent.go to identify GitHub Copilot based on its Provider ID or Base URL.

2. Disable Non-Essential Requests

  • Skip generateTitle() for metered providers when starting a new session to save an initial request.

3. Efficiency-First System Prompt

Inject efficiency instructions into promptPrefix() for these providers:

  • Encourage the model to call multiple independent tools in parallel within a single turn.
  • Minimize conversational filler or redundant confirmation steps.
  • Guide the model to provide more comprehensive initial analysis to reduce future follow-up requests.

Assisted-by: gemini3 flash

@Gustave-241021
Gustave-241021 requested a review from a team as a code owner December 18, 2025 14:16
@Gustave-241021
Gustave-241021 requested review from andreynering and aymanbagabas and removed request for a team December 18, 2025 14:16
@domdeger

Copy link
Copy Markdown

Is that really the right way? In VSCode Copilot one Request can be like 100 requests to the model back and forth. It is counted by one request of the user and then the Agent keeps going until it decides that it has finished. So there must be some provider specific mechanism how long a requests "session" can go. I think the approach of this PR won't solve the root cause of the issue.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

Is that really the right way? In VSCode Copilot one Request can be like 100 requests to the model back and forth. It is counted by one request of the user and then the Agent keeps going until it decides that it has finished. So there must be some provider specific mechanism how long a requests "session" can go. I think the approach of this PR won't solve the root cause of the issue.

I'm just trying to correct this as best as I can based on the information I currently have.😂

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

Is that really the right way? In VSCode Copilot one Request can be like 100 requests to the model back and forth. It is counted by one request of the user and then the Agent keeps going until it decides that it has finished. So there must be some provider specific mechanism how long a requests "session" can go. I think the approach of this PR won't solve the root cause of the issue.

Thank you for your suggestion. I might start by looking into HTTP requests to understand how it works.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

You hit the nail on the head, @domdeger . Thanks for the insight.

I performed a reverse analysis (packet capture) on VS Code's Agent mode traffic, and it confirms your theory: we are currently treating every step as a standalone request, which triggers the rate limiter. VS Code, however, groups these steps into a single billing session.

The Root Cause:

  • VS Code maintains a "Session" context by generating a UUID and reusing it across the entire chain of thought (plan -> tool -> observe -> plan).

To fix this properly, we need to modify the request headers in Crush. Specifically, for a single Agent task loop, we must:

  • Generate a UUID at the start of the task.

  • Reuse this UUID in the x-interaction-id header for every subsequent HTTP request in that loop.

  • Explicitly set the following headers to signal the backend that this is an Agent session:

  • openai-intent: conversation-agent

  • x-interaction-type: conversation-agent

  • copilot-integration-id: vscode-chat

I will work on implementing this session-keeping logic to ensure Crush consumes quota per task, not per step, similar to the native VS Code experience.

@domdeger

Copy link
Copy Markdown

Thank you for your hard work! Really interesting to learn how they handle this. @Gustave-241021

@kujtimiihoxha

Copy link
Copy Markdown
Contributor
image

(@Gustave-241021 that is awesome let me know if you need any help implementing this)

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

@kujtimiihoxha Sorry to bother you. While pushing forward with this change, I noticed that my current modifications seem to bypass GitHub Copilot's usage statistics when the corresponding model can be used. Could you help me identify the problem? If necessary, please report this issue to GitHub Copilot in a timely manner.

https://github.com/Gustave-241021/crush/tree/hack/github-copilot-with-no-Usage

@Gustave-241021
Gustave-241021 force-pushed the fix/reduce-copilot-quota branch 2 times, most recently from cf8fee2 to d8f3070 Compare December 19, 2025 15:36
@Gustave-241021

Copy link
Copy Markdown
Contributor Author

This is the latest development:

Implementation Complete

Feature Status
VSCode headers (User-Agent, Editor-Version, etc.) ✅
/models/session API call for auto mode ✅
x-interaction-id for request grouping ✅
vscode-sessionid, vscode-machineid ✅
vscode-abexpcontext (A/B experiment flags) ✅
Token refresh before session init ✅
CopilotHeaderTransport for header injection ✅

Current Behavior

Headers confirmed matching VSCode exactly:

User-Agent: GitHubCopilotChat/0.35.2
Copilot-Integration-Id: vscode-chat
Editor-Version: vscode/1.107.1
x-interaction-id: (same UUID for all requests in interaction)

Remaining Issue

Request grouping not working

Possible Causes

  1. Server-side anti-spoofing - GitHub may use additional signals to verify authentic VSCode clients
  2. Telemetry verification - VSCode sends telemetry that may establish client identity
  3. HTTP/2 connection behavior - Connection pooling differences between clients
  4. Unknown server logic - Other backend mechanisms we haven't identified

Recommendation

The implementation correctly replicates all observable VSCode HTTP headers and API behavior. The remaining issue is likely server-side and may require:

  1. Further investigation of VSCode telemetry payloads
  2. Engagement with GitHub Copilot team
  3. Acceptance of current behavior

@slhad

slhad commented Dec 25, 2025

Copy link
Copy Markdown

@Gustave-241021 thank you for your great work investigating the issue,

Just to check if Crush has the same problem of some other VS Code extension using "VS Code LM API" or alike, I burned through 10% of my premium requests (soon new reset, it's ok :D ) instead of 0.2 to 1% with Haiku 4.5

It feels like curent implementation is doing a new request for each turn.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

@slhad Yes, that's why solving the same problem often consumes more quotas for crushes than for VSCode GitHub Copilot

@dcominottim

Copy link
Copy Markdown

I know there is some recent bad blood between the projects, but anyway, it might be worth checking out how sst/opencode is doing this; I used it quite extensively in the last week with Copilot and my quota usage seems to be as expected.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

I know there is some recent bad blood between the projects, but anyway, it might be worth checking out how sst/opencode is doing this; I used it quite extensively in the last week with Copilot and my quota usage seems to be as expected.

Thank you for your suggestion, I will refer to opencode.

@meowgorithm

Copy link
Copy Markdown
Member

Small note here that this is important to us and we do fully intend to sort this one out. Thanks, everyone, for all the work in this so far.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

If you have better ideas, please communicate with me in a timely manner. I will continue to push forward this feature until it is completed

@meowgorithm

Copy link
Copy Markdown
Member

@Gustave-241021 No, by all means please keep going. But if you need anything from our end, by all means please let us know.

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

@meowgorithm Apologies for the confusion. To clarify, I just wanted to say that I’m fully committed to this and will follow up until the matter is completely resolved

@PeroSar

PeroSar commented Dec 30, 2025 •

Copy link
Copy Markdown

@Gustave-241021 I think you need this, maybe?

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

@Gustave-241021 I think you need this, maybe?

thanks

@Gustave-241021
Gustave-241021 force-pushed the fix/reduce-copilot-quota branch from d8f3070 to ca06767 Compare December 30, 2025 14:17
@Gustave-241021

Copy link
Copy Markdown
Contributor Author

In my tests, this change has worked. You can pull my branch to help verify whether it is effective
If there are any problems, please point them out in the comments in a timely manner

Here is the summary👇

Status: ✅ FIXED


Phase 1: VSCode Header Replication (Partial Success)

Approach

Replicated all observable VSCode Copilot Chat HTTP headers:

  • User-Agent, Editor-Version, Editor-Plugin-Version, Copilot-Integration-Id
  • vscode-sessionid, vscode-machineid, vscode-abexpcontext
  • x-interaction-id, x-interaction-type, openai-intent
  • /models/session API call for auto mode initialization

Result

❌ Request grouping still not working

Conclusion

Header matching alone was insufficient. Server-side logic required something else.


Phase 2: X-Initiator Fix (Complete Success)

Discovery

Referenced opencode PR #595 and codecompanion.nvim fix.

Key insight: The X-Initiator header determines billing:

  • X-Initiator: user → Counts against quota
  • X-Initiator: agent → Does NOT count

Implementation

Dynamic X-Initiator based on message roles:

func determineInitiator(body []byte) string {
    // Parse messages, check for "tool" or "assistant" roles
    for _, msg := range reqBody.Messages {
        if msg.Role == "tool" || msg.Role == "assistant" {
            return "agent"  // Follow-up request, no charge
        }
    }
    return "user"  // Initial request, normal charge
}

Result

✅ Complete fix working

Thank you to everyone who has paid attention to this matter and provided help💖💖💖🥰🥰🥰

@PeroSar

PeroSar commented Dec 30, 2025

Copy link
Copy Markdown

@Gustave-241021 great job!

@domdeger

Copy link
Copy Markdown

Thank you for your hard work @Gustave-241021! This is makes crush way more useful!

@dcominottim

Copy link
Copy Markdown

You're awesome, @Gustave-241021, thank you so much!

Is there any official doc on how to set up the integration with GH Copilot? I haven't found any.

@kujtimiihoxha

Copy link
Copy Markdown
Contributor

@Gustave-241021 I simplified the implementation and added you as a co-author in #1738

@Gustave-241021

Copy link
Copy Markdown
Contributor Author

@dcominottim In fact, I couldn't find the corresponding official documentation. I did it by capturing packages and referencing open code

@dcominottim

Copy link
Copy Markdown

Oh, thanks for the clarification, @Gustave-241021.

@kujtimiihoxha @meowgorithm, is it just a case of missing documentation for something that already exists or is some kind of workaround/bridge needed for Crush to use GH Copilot?

@meowgorithm

Copy link
Copy Markdown
Member

@dcominottim Copilot integration in built into Crush. Use ctrl+l to open the model chooser, filter for "Copilot" and choose a model. You'll be prompted to authenticate if necessary.

@meowgorithm

Copy link
Copy Markdown
Member

Just a note that this was fixed in #1738 based on @Gustave-241021’s excellent work here and should be in a release later today.

Thanks, everyone, for your support with this one and let us know if you experience any further issues.

@meowgorithm meowgorithm closed this Jan 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants