[Feature]: KVPress #10491

flozi00 · 2024-11-20T11:06:03Z

🚀 The feature, motivation and pitch

https://github.com/NVIDIA/kvpress

Deploying long-context LLMs is costly due to the linear growth of the key-value (KV) cache in transformer models. For example, handling 1M tokens with Llama 3.1-70B in float16 requires up to 330GB of memory. This repository implements multiple KV cache pruning methods and benchmarks using 🤗 transformers, aiming to simplify the development of new methods for researchers and developers in this field.

Alternatives

No response

Additional context

No response

Before submitting a new issue...

Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

flozi00 added the feature request label Nov 20, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Feature]: KVPress #10491

[Feature]: KVPress #10491

flozi00 commented Nov 20, 2024

[Feature]: KVPress #10491

[Feature]: KVPress #10491

Comments

flozi00 commented Nov 20, 2024

🚀 The feature, motivation and pitch

Alternatives

Additional context

Before submitting a new issue...