Pete Cheslock

@petecheslock.com

๐Ÿฅฉ He/Him ๐Ÿ– "Anything worth doing is worth overdoing."

Part 2 of our ๐——๐—ถ๐˜€๐˜๐—ฟ๐—ถ๐—ฏ๐˜‚๐˜๐—ฒ๐—ฑ ๐—”๐—œ ๐—œ๐—ป๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ series is now live on Red Hat Developer: ๐˜–๐˜ฑ๐˜ต๐˜ช๐˜ฎ๐˜ช๐˜ป๐˜ช๐˜ฏ๐˜จ ๐˜‹๐˜ช๐˜ด๐˜ต๐˜ณ๐˜ช๐˜ฃ๐˜ถ๐˜ต๐˜ฆ๐˜ฅ ๐˜ˆ๐˜ ๐˜๐˜ฏ๐˜ง๐˜ฆ๐˜ณ๐˜ฆ๐˜ฏ๐˜ค๐˜ฆ: ๐˜ˆ๐˜ฅ๐˜ท๐˜ข๐˜ฏ๐˜ค๐˜ฆ๐˜ฅ ๐˜‹๐˜ฆ๐˜ฑ๐˜ญ๐˜ฐ๐˜บ๐˜ฎ๐˜ฆ๐˜ฏ๐˜ต ๐˜—๐˜ข๐˜ต๐˜ต๐˜ฆ๐˜ณ๐˜ฏ๐˜ด. In Part 1, we covered prefill/decode phases and the 5D parallelism framework.

Excited to share Part 1 of our blog series on Red Hat Developer: ๐˜‹๐˜ฆ๐˜ด๐˜ช๐˜จ๐˜ฏ๐˜ช๐˜ฏ๐˜จ ๐˜‹๐˜ช๐˜ด๐˜ต๐˜ณ๐˜ช๐˜ฃ๐˜ถ๐˜ต๐˜ฆ๐˜ฅ ๐˜ˆ๐˜ ๐˜๐˜ฏ๐˜ง๐˜ฆ๐˜ณ๐˜ฆ๐˜ฏ๐˜ค๐˜ฆ: ๐˜Š๐˜ฐ๐˜ณ๐˜ฆ ๐˜Š๐˜ฐ๐˜ฏ๐˜ค๐˜ฆ๐˜ฑ๐˜ต๐˜ด ๐˜ข๐˜ฏ๐˜ฅ ๐˜š๐˜ค๐˜ข๐˜ญ๐˜ช๐˜ฏ๐˜จ ๐˜‹๐˜ช๐˜ฎ๐˜ฆ๐˜ฏ๐˜ด๐˜ช๐˜ฐ๐˜ฏ๐˜ด. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache.

Hey Boston friends, we're cooking up another great event in the area. Workshop + evening sessions covering: - vLLM project update - Model compression and speculative decoding - Agentic AI with vLLM - Distributed inference at scale with llm-d and k8s luma.com/4rmkrrb7

vLLM Inference Meetup ยท Boston ยท Luma

Deep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,โ€ฆ

luma.com

๐Ÿ“ข ๐—ง๐—ต๐—ฒ ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ ๐—ผ๐—ณ ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฆ๐—ฒ๐—ฟ๐˜ƒ๐—ถ๐—ป๐—ด ๐—–๐—ผ๐—บ๐—บ๐˜‚๐—ป๐—ถ๐˜๐—ถ๐—ฒ๐˜€: ๐— ๐—ฎ๐—ฟ๐—ฐ๐—ต ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐˜€ ๐—ผ๐˜‚๐˜! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. Weโ€™ve gained over ๐Ÿญ๐Ÿฏ๐Ÿฌ๐Ÿฌ ๐˜€๐˜‚๐—ฏ๐˜€๐—ฐ๐—ฟ๐—ถ๐—ฏ๐—ฒ๐—ฟ๐˜€!

The agenda is still evolving, and weโ€™ve got even more awesomeness in the works! ๐Ÿ“ˆ Whether you're running GenAI in production or building the platforms to support it, this is the room to be in. ๐Ÿ“… March 11 | 4:30 PM ๐Ÿ“ 1 Madison Ave, NYC ๐ŸŽŸ๏ธ RSVP: luma.com/0crwqwg4

Distributed Inference Meetup NYC ยท Luma

llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What toโ€ฆ

luma.com

๐Ÿ“ข ๐—ง๐—ต๐—ฒ ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ ๐—ผ๐—ณ ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฆ๐—ฒ๐—ฟ๐˜ƒ๐—ถ๐—ป๐—ด ๐—–๐—ผ๐—บ๐—บ๐˜‚๐—ป๐—ถ๐˜๐—ถ๐—ฒ๐˜€: ๐—™๐—ฒ๐—ฏ๐—ฟ๐˜‚๐—ฎ๐—ฟ๐˜† ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐˜€ ๐—ผ๐˜‚๐˜! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. Weโ€™ve gained over ๐Ÿญ๐Ÿฎ๐Ÿฌ๐Ÿฌ ๐˜€๐˜‚๐—ฏ๐˜€๐—ฐ๐—ฟ๐—ถ๐—ฏ๐—ฒ๐—ฟ๐˜€!

Standardizing high-performance inference requires deep ecosystem collaboration. ๐Ÿš€ Huge shoutout to @vllm_project and @IBMResearch on the new KV Offloading Connector. Weโ€™re seeing up to 9x throughput gains on H100s and massive TTFT reductions. ๐Ÿงต blog.vllm.ai/2026/01/08/k...

Inside vLLMโ€™s New KV Offloading Connector: Smarter Memory Transfer for Maximizing Inference Throughput

In this post, we will describe the new KV cache offloading feature that was introduced in vLLM 0.11.0. We will focus on offloading to CPU memory (DRAM) and its benefits to improving overall inferenceโ€ฆ

blog.vllm.ai

So a long time ago when buying new headphones and reading reviews, I noticed how the reviews often sounded similar to reviews for a bottle of wine. Like: "Rich and full-bodied with excellent depth. The bass notes are particularly impressive, with a smooth finish that lingers pleasantly."