@simjeg.bsky.social

Senior LLM Technologist @NVIDIA Views and opinions are my own

Fresh news from kvpress, our open source library for KV cache compression 🔥 1. We published a blog post with @huggingface 2. We published a Space for you to try it 3. Following feedback from the research community, we added a bunch of presses and benchmarks Links👇(1/2)

Bild

Hidden states in LLM ~ follow normal distributions. Consequently, both queries and keys also follow a normal distribution and if you replace all queries and keys by their average counterpart, this magically explains the slash pattern observed in attention matrices

Bild

Ever noticed that the attention mechanism in transformers is essentially a two-layer MLP? 🤔 A(q, K, V) = V @ softmax(K / √d @ q) Weights: K / √d and V nonlinearity: softmax 💡This offers fresh insights into KV cache compression research 🧵(1/3)