vLLM at #PyTorchCon NA 2026 covers attention, KV cache management, disaggregated serving, expert parallelism, and hardware portability. Register by Sept 4: https://hubs.la/Q04tDf0j0 Guide: https://pytorch.org/blog/vllm-sessions-at-pytorch-conference-north-america-2026/
PyTorch
@pytorch.org
Tensors and neural networks in Python with strong hardware acceleration. PyTorch is an open source project at the Linux Foundation. #PyTorchFoundation pytorch.org
At #PyTorchCon NA 2025, Christian Jacobi (IBM) spoke about the innovation ahead in deploying AI across core IT and business processes at scale with enterprise qualities of service. #PyTorchCon North America returns Oct 20–21. Register by Sept 4: https://hubs.la/Q04tDf0j0
At PyTorch Conference North America 2026, hear from the people working across PyTorch Foundation projects. Simon Mo (vLLM, Inferact) says vLLM’s goal is to become “the easiest to use and most efficient inference engine.” #PyTorchCon Register by Sept 4: https://hubs.la/Q04tDf0j0
Core PyTorch at #PyTorchCon NA 2026 spans torch.compile, dynamic shapes, distributed communication, release/CI, stable ABI, and accelerator backends. Register by Sept 4: https://hubs.la/Q04tDf0j0 Guide: https://pytorch.org/blog/core-pytorch-sessions-at-pytorch-conference-north-america-2026/
NVIDIA’s blog shows how QAD improves Nemotron 3.5 Lightning using NVIDIA Model Optimizer, outperforming PTQ on agentic benchmarks. https://developer.nvidia.com/blog/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer/
Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer | NVIDIA Technical Blog
Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models, developers can find the right-sized model for their needs.
developer.nvidia.com
Network with developers, researchers, maintainers, and practitioners across the PyTorch community at #PyTorchCon NA, October 20–21 in San Jose. Watch PyTorch Foundation Executive Director Mark Collier preview the conference. Register by September 4 to save: https://hubs.la/Q04tDf0j0
We’re proud to announce the inaugural PyTorch Day Japan, hosted by PyTorch Foundation, Hugging Face, IBM, and Mitsubishi Electric, on Dec. 10 in Tokyo. Technical talks and discussions will span training, inference, responsible AI, and more. https://hubs.la/Q04vtQxX0
What began as experimentation with a Python framework evolved into PyTorch, the underlying framework for developers in the fast-moving AI space. Jana van Greunen (Meta) reflects on that evolution at #PyTorchCon NA 2025. Register by Sept 4: https://hubs.la/Q04tDf0j0
TRANSIT is a runtime that makes unified virtual memory practical for large-scale LLM training, enabling models to train on up to 50% fewer GPUs without requiring modification to existing PyTorch training code. Learn how to train more with less at PyTorch Conference North America: hubs.la/Q04v4SL60
Leverage PyTorch-native building blocks for constructing GPU-accelerated materials simulation workflows with coding agents. NVIDIA follows an end-to-end NVIDIA ALCHEMI Toolkit workflow. Read more: developer.nvidia.com/blog/how-ai-...
How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit | NVIDIA Technical Blog
Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the simulation stack. The first remains the…
developer.nvidia.com
Congrats to Mark Gatere for creating the winning design in our 2026 PyTorch Foundation Flare Pin Community Design Contest! You'll be able to collect this flare pin at the Flare Party during PyTorch Conference North America. Learn more & get your ticket for #PyTorchCon hubs.la/Q04qXq570 #PyTorchPin
The keynote sessions at PyTorch Conference North America will focus on: strengthening the core, expanding multi-hardware support, accelerating efficient training and inference, and building next-generation intelligent systems. Learn more about our keynote speakers here: https://bit.ly/4xeMiwu
"The PyTorch Conference is really unique in bringing together users, practitioners and developers across the AI spectrum in a way that really no other event does." Niles Burbank, AMD #PyTorchCon Join us October 20–21 for PyTorch Conference North America 2026: https://bit.ly/4wJLBub
Tristan Rice (Meta) on hearing directly from users at #PyTorchCon NA 2025. Save $200 by Sept. 4: https://events.linuxfoundation.org/pytorch-conference-north-america/register/?utm_campaign=49768803-26Q3_PyTorch%20Conference%20North%20America&utm_source=bluesky&utm_medium=social#register-now
Leverage a PyTorch-native fine-tuning library with Day-0 Hugging Face checkpoint support for the recently released Alibaba open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model. Read the full post: https://bit.ly/4xdN25l
Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72 | NVIDIA Technical Blog
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem. It has 2.4T total parameters with 95B…
developer.nvidia.com
To bridge the gap between rapid model evolution and hardware compiler updates, the IBM Spyre team used AI coding agents to write lightweight PyTorch runtime adapters, enabling over 6,800 Hugging Face models on new silicon. Read the full technical deep dive here 👉 https://bit.ly/4wFGMlB
#PyTorchCon North America brings the PyTorch community together for technical exchange, collaboration, and connection. Sharon Zhou, PhD (AMD) shared that perspective at the 2025 conference. Register by September 4 to save $200 and secure your pass: https://bit.ly/3U6lMa8
Ant Group has joined PyTorch Foundation as a Gold Member. Ant Group is one of the organizations behind AReaL, a PyTorch Ecosystem Landscape project, and will continue investing in its development while deepening technical integration and collaboration across the PyTorch ecosystem landscape.
AMD has been upstreaming optimizations for improved FP8 training support in PyTorch/TorchTitan and PyTorch/TorchAO, making FP8 training work out of the box on AMD Instinct GPUs! Read our latest blog: pytorch.org/blog/fp8-tra...
The PyTorch Certified Associate (PTCA) program offers early-stage practitioners a direct route to build credibility and demonstrate hands-on expertise in AI engineering. Learn more about the certification and how to enroll here 👉 https://bit.ly/4z5oZGN
3 days left to enter the 2026 PyTorch Foundation flare pin contest. Win a ticket to PyTorch Conference North America in San Jose, and see your design become the conference pin. Use the optional AI design prompt: https://pytorch.org/blog/pytorch-foundation-flare-pin-community-design-contest/
Win a free ticket to PyTorch Conference North America in San Jose & see your design come to life --> Enter the 2026 Flare Pin Community Design Contest. Submit your original 1" enamel pin design featuring the PyTorch logo & "2026" by Aug 14th. Post your entries here using #PyTorchPin and #PyTorchCon
Today Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark for on-device agentic workflows & ExecuTorch is adding end-to-end support for running Muse Glimmer on NVIDIA GPUs and Macs with Apple silicon. Learn more: pytorch.org/blog/fast-on...
TorchTitan uses NVIDIA contributions to deliver approximately 6x higher performance on the same GB300 NVL72 infrastructure, reaching a record 1,648 TFLOPs per GPU for DeepSeek-V3 671B pre-training. https://developer.nvidia.com/blog/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72/
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | NVIDIA Technical Blog
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls…
developer.nvidia.com
PyTorch Conference NA 2026 features keynotes from PyTorch Foundation, Agentic AI Foundation, Meta, Google Cloud, Red Hat, NVIDIA, Cohere, Inferact, Core Automation and Adaption in San Jose, Oct. 20–21. Full lineup: https://pytorch.org/blog/pytorch-conference-north-america-announces-2026-keynotes/
UCSC OSPO and Red Hat hosted their inaugaral Santa Cruz PyTorch Meetup! Check out their origin story, the content of their first event, and why you should start your own community PyTorch Meetup 👉 https://bit.ly/4fLDjLK
Design the 2026 PyTorch Foundation flare pin. Your work could become the conference pin and win you a complimentary ticket. Use our blog’s AI prompt. Submit by Aug. 14 with #PyTorchPin and #PyTorchCon: https://pytorch.org/blog/pytorch-foundation-flare-pin-community-design-contest/
In this post, you’ll learn how to use the PyTorch Torch Inductor compiler and kernel fusion to improve memory bandwidth and reduce kernel launch overhead. Read the full post: https://bit.ly/45Emn5g
Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead | NVIDIA Technical Blog
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead, along with multiple ways to apply it in…
developer.nvidia.com
Design the 2026 PyTorch Foundation flare pin, and you could win a complimentary PyTorch Conference North America ticket. Try our AI design prompt. Submit by Aug. 14 at 11:59 p.m. PT with #PyTorchPin and #PyTorchCon. https://pytorch.org/blog/pytorch-foundation-flare-pin-community-design-contest/
FBTriton is the Triton repo where Meta develops its experimental GPU optimization solutions (including TLX/torchTLX and autoWS). Read the full blog to learn more here: https://bit.ly/4fQDDcu..*