fix:optimize quota usage for Copilot and metered models - #1662
Gustave-241021 wants to merge 3 commits into
Conversation
|
Is that really the right way? In VSCode Copilot one Request can be like 100 requests to the model back and forth. It is counted by one request of the user and then the Agent keeps going until it decides that it has finished. So there must be some provider specific mechanism how long a requests "session" can go. I think the approach of this PR won't solve the root cause of the issue. |
I'm just trying to correct this as best as I can based on the information I currently have.😂 |
Thank you for your suggestion. I might start by looking into HTTP requests to understand how it works. |
|
You hit the nail on the head, @domdeger . Thanks for the insight. I performed a reverse analysis (packet capture) on VS Code's Agent mode traffic, and it confirms your theory: we are currently treating every step as a standalone request, which triggers the rate limiter. VS Code, however, groups these steps into a single billing session. The Root Cause:
To fix this properly, we need to modify the request headers in Crush. Specifically, for a single Agent task loop, we must:
I will work on implementing this session-keeping logic to ensure Crush consumes quota per task, not per step, similar to the native VS Code experience. |
|
Thank you for your hard work! Really interesting to learn how they handle this. @Gustave-241021 |
(@Gustave-241021 that is awesome let me know if you need any help implementing this) |
|
@kujtimiihoxha Sorry to bother you. While pushing forward with this change, I noticed that my current modifications seem to bypass GitHub Copilot's usage statistics when the corresponding model can be used. Could you help me identify the problem? If necessary, please report this issue to GitHub Copilot in a timely manner. https://github.com/Gustave-241021/crush/tree/hack/github-copilot-with-no-Usage |
cf8fee2 to
d8f3070
Compare
|
This is the latest development: Implementation Complete
Current BehaviorHeaders confirmed matching VSCode exactly: Remaining IssueRequest grouping not working Possible Causes
RecommendationThe implementation correctly replicates all observable VSCode HTTP headers and API behavior. The remaining issue is likely server-side and may require:
|
|
@Gustave-241021 thank you for your great work investigating the issue, Just to check if Crush has the same problem of some other VS Code extension using "VS Code LM API" or alike, I burned through 10% of my premium requests (soon new reset, it's ok :D ) instead of 0.2 to 1% with Haiku 4.5 It feels like curent implementation is doing a new request for each turn. |
|
@slhad Yes, that's why solving the same problem often consumes more quotas for crushes than for VSCode GitHub Copilot |
|
I know there is some recent bad blood between the projects, but anyway, it might be worth checking out how sst/opencode is doing this; I used it quite extensively in the last week with Copilot and my quota usage seems to be as expected. |
Thank you for your suggestion, I will refer to opencode. |
|
Small note here that this is important to us and we do fully intend to sort this one out. Thanks, everyone, for all the work in this so far. |
|
If you have better ideas, please communicate with me in a timely manner. I will continue to push forward this feature until it is completed |
|
@Gustave-241021 No, by all means please keep going. But if you need anything from our end, by all means please let us know. |
|
@meowgorithm Apologies for the confusion. To clarify, I just wanted to say that I’m fully committed to this and will follow up until the matter is completely resolved |
|
@Gustave-241021 I think you need this, maybe? |
thanks |
d8f3070 to
ca06767
Compare
|
In my tests, this change has worked. You can pull my branch to help verify whether it is effective Here is the summary👇Status: ✅ FIXED Phase 1: VSCode Header Replication (Partial Success)ApproachReplicated all observable VSCode Copilot Chat HTTP headers:
Result❌ Request grouping still not working ConclusionHeader matching alone was insufficient. Server-side logic required something else. Phase 2: X-Initiator Fix (Complete Success)DiscoveryReferenced opencode PR #595 and codecompanion.nvim fix. Key insight: The
ImplementationDynamic func determineInitiator(body []byte) string {
// Parse messages, check for "tool" or "assistant" roles
for _, msg := range reqBody.Messages {
if msg.Role == "tool" || msg.Role == "assistant" {
return "agent" // Follow-up request, no charge
}
}
return "user" // Initial request, normal charge
}Result✅ Complete fix working Thank you to everyone who has paid attention to this matter and provided help💖💖💖🥰🥰🥰 |
|
@Gustave-241021 great job! |
|
Thank you for your hard work @Gustave-241021! This is makes crush way more useful! |
|
You're awesome, @Gustave-241021, thank you so much! Is there any official doc on how to set up the integration with GH Copilot? I haven't found any. |
|
@Gustave-241021 I simplified the implementation and added you as a co-author in #1738 |
|
@dcominottim In fact, I couldn't find the corresponding official documentation. I did it by capturing packages and referencing open code |
|
Oh, thanks for the clarification, @Gustave-241021. @kujtimiihoxha @meowgorithm, is it just a case of missing documentation for something that already exists or is some kind of workaround/bridge needed for Crush to use GH Copilot? |
|
@dcominottim Copilot integration in built into Crush. Use ctrl+l to open the model chooser, filter for "Copilot" and choose a model. You'll be prompted to authenticate if necessary. |
|
Just a note that this was fixed in #1738 based on @Gustave-241021’s excellent work here and should be in a release later today. Thanks, everyone, for your support with this one and let us know if you experience any further issues. |

CONTRIBUTING.md.fix:Copilot Models burn through request quota #1657
Root Cause Analysis
GitHub Copilot quotas are counted by requests (turns). Since
crushuses an agentic architecture, it iterates through several turns to solve a problem:generateTitle()) and auto-summarizing (Summarize()) consume additional requests.Proposed Design
1. Metered Provider Identification
Add an
isMetered()helper function inagent.goto identify GitHub Copilot based on its Provider ID or Base URL.2. Disable Non-Essential Requests
generateTitle()for metered providers when starting a new session to save an initial request.3. Efficiency-First System Prompt
Inject efficiency instructions into
promptPrefix()for these providers:Assisted-by: gemini3 flash