Qwen just announced Qwen3.8-27B along with Qwen3.8-Max! 🔥 Qwen3.8-27B will run locally on 17GB RAM/VRAM setups and is expected to be the best performing model for its size.
Unsloth AI
@unsloth.ai
Making open-source AI more accessible! 🦥 Github: http://github.com/unslothai/unsloth
DeepSeek V4 Flash 0731 can now be run locally! 🐳 Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM. V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Deep...
You can now run Inkling-Small, a new 276B model by Thinking Machines. Inkling-Small is the strongest open model for its size and runs local on 128GB RAM. Apache-2.0 Licensed, it has image, audio + 1M context support. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Inkl...
We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts. 1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s. GitHub repo: github.com/unslothai/un...
Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
We signed the Open Weights letter because we believe the future of AI should be shaped by everyone, not controlled by a select few. That belief has always been at the heart of Unsloth: everyone should be able to train and run models on their own local device.
Introducing Unsloth for AMD 🚀 You can now train & run LLMs on your AMD hardware • We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs • Works on Windows, WSL, Linux • Train Qwen, Gemma on just 3GB VRAM GitHub: github.com/unslothai/un... Blog: unsloth.ai/docs/basics/...
Gemma 4 is now faster and much more accurate! 🚀 Google made huge improvements to tool-calling and chat accuracy, reliability + speed. To get fixes, re-download our updated GGUF, MLX, NVFP4 quants! Unsloth quants: huggingface.co/collections/... Gemma 4 Guide: unsloth.ai/docs/models/...
Inkling, a new 975B parameter open model is here! From Thinking Machines, Inkling supports image, audio, text & 1M context. We quantized Inkling to Dynamic 1-bit (-86% size) and retained 74.2% of top-1% accuracy. Run on 280GB. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/inkl...
We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: unsloth.ai/docs/basics/... Gemma NVFP4: huggingface.co/collections/...
We collaborated with AWS on a complete guide to LLM Quantization and Deployment. Learn about: • Model formats, dynamic quants & making your own • Choosing GGUF, NVFP4 or FP8 • Picking the right tools & deploy on AWS SageMaker • Benchmark quality, latency & cost Read: aws.amazon.com/blogs/machin...
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡ Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/... Qwen3.6 NVFP4: huggingface.co/collections/...
DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳 Run lossless DeepSeek-V4-Flash on 168GB RAM. 3-bit works on 110GB Mac, RAM, VRAM setups. Run via Unsloth Studio or llama.cpp. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Deep...
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5 We gave 3 models the same prompt and compared one-shot outputs. The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra 256GB RAM at ~21.6 tok/s. Which do you like best? GGUF: huggingface.co/unsloth/GLM-... Guide: unsloth.ai/docs/models/...
GLM-5.2 can now be run locally! 🔥 The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/GLM-...
You can now run Kimi K2.7 Code locally! 🌘 We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted. Run at >40 tok/s on 330GB RAM/VRAM setups. Run full precision on 610 GB. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
DiffusionGemma can now run at 2000+ tokens/sec! ⚡ We made local DiffusionGemma inference 1.8× faster. Run it on 18GB RAM via Unsloth Studio. GitHub: github.com/unslothai/un... Guide: unsloth.ai/docs/models/...
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/...
Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF: huggingface.co/unsloth/diff... Guide: unsloth.ai/docs/models/...
Google releases Gemma 4 QAT. ✨ You can now run Gemma 4 at 3x less memory with near original performance. Quantization-Aware Training (QAT) makes it possible to run Gemma 4 26B-A4B on 16GB RAM. GGUFs: huggingface.co/collections/... QAT Guide: unsloth.ai/docs/models/...
NVIDIA releases Nemotron 3 Ultra, a new 550B model. 💚 Nemotron-3-Ultra-550B-A55B is NVIDIA's largest LLM yet, with 1M context, frontier coding & chat. Run 2-bit on 200GB RAM, 3-bit on 256GB, 8-bit on 600GB. GGUF: huggingface.co/unsloth/NVID... Guide: unsloth.ai/docs/models/...
Google releases Gemma 4 12B, a new model that can run locally on 8GB RAM. Gemma 4 12B Unified model supports image, audio and 256K context. Run and train the model via Unsloth Studio. GGUF: huggingface.co/unsloth/gemm... Guide: unsloth.ai/docs/models/...
We made a guide on using MCP with local LLMs. Connect Qwen3.6 and Gemma 4 for controlled access to tools, files, APIs, enabling private automated workflows. Learn to use OAuth, Exa, Context7, Hugging Face & more. Guide: unsloth.ai/docs/basics/... GitHub: github.com/unslothai/un...
Qwen3.6 now runs 2x faster with MTP GGUFs! Run locally on just 18GB RAM. ⚡️ MTP enables Qwen3.6 to generate ~1.4–2.2× faster with no accuracy change. Qwen3.6-27B MTP runs at 160 tokens/s. 35B-A3B reaches 240 t/s. GGUFs: huggingface.co/unsloth/Qwen... Guide: unsloth.ai/docs/models/...
We’re excited to share that Unsloth has joined the PyTorch Ecosystem! Unsloth is an open-source project that makes training & running models faster, more accurate with less compute. We want AI to be accessible to everyone. Blog: unsloth.ai/blog/pytorch GitHub: github.com/unslothai/un...
We collaborated with NVIDIA to teach you how we made LLM training ~25% faster! 🚀 Learn how 3 optimizations help your home GPU train models faster: 1. Packed-sequence metadata caching 2. Double-buffered checkpoint reloads 3. Faster MoE routing Guide: unsloth.ai/blog/nvidia-...
We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw. Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp Guide: unsloth.ai/docs/basics/...
NVIDIA releases Nemotron-3-Nano-Omni, a new 30B open multimodal MoE model. Nemotron-3-Nano-Omni-30B-A3B is the strongest omni model for its size and supports audio, video, image and text. Run on ~25GB RAM. GGUF: huggingface.co/unsloth/NVID... Guide: unsloth.ai/docs/models/...
DeepSeek releases DeepSeek-V4. 🐋 - DeepSeek-V4-Pro: 1.6T params - DeepSeek-V4-Flash: 284B params DeepSeek-V4-Pro rivals Claude-Opus-4.6-Max, GPT-5.4-xHigh and Gemini-3.1-Pro-High. They support 1M context length, thinking and set new records for Codeforces.