Sparse attention is one of the most promising strategies to unlock long-context processing and long-generation reasoning in LLMs. We performed the most comprehensive study on training-free sparse attention to date. Here is what we found:
@pnawrot.bsky.social
Several incredible NeurIPS tutorials this year. Worth navigating through the Swifties.
Another nano gem from my amazing student Piotr Nawrot! A repo & notebook on sparse attention for efficient LLM inference: github.com/PiotrNawrot/... This will also feature in my #NeurIPS 2024 tutorial "Dynamic Sparsity in ML" with André Martins: dynamic-sparsity.github.io Stay tuned!
Another nano gem from my amazing student Piotr Nawrot! A repo & notebook on sparse attention for efficient LLM inference: github.com/PiotrNawrot/... This will also feature in my #NeurIPS 2024 tutorial "Dynamic Sparsity in ML" with André Martins: dynamic-sparsity.github.io Stay tuned!