Simone Scardapane

@sscardapane.bsky.social

I fall in love with a new #machinelearning topic every month 🙄 Ass. Prof. Sapienza (Rome) | Author: Alice in a differentiable wonderland (https://www.sscardapane.it/alice-book/)

🚀 New Paper Alert! 🚀 We introduce Q-Filters, a training-free method for efficient KV Cache compression! It is compatible with FlashAttention and can compress along generation which is particularly useful for reasoning models ⚡ TLDR: we make Streaming-LLM smarter using the geometry of attention

Bild

Q-Filters is very efficient which allows streaming compression at virtually no latency cost, just like Streaming-LLM... ...but it is also much better at retaining relevant KV pairs compared to fast alternatives (and can even beat slower algorithms such as SnapKV)

Bild