Unsloth AI

@unsloth.ai

Making open-source AI more accessible! 🦥 Github: http://github.com/unslothai/unsloth

Qwen just announced Qwen3.8-27B along with Qwen3.8-Max! 🔥 Qwen3.8-27B will run locally on 17GB RAM/VRAM setups and is expected to be the best performing model for its size.

Bild

We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts. 1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s. GitHub repo: github.com/unslothai/un...

Unsloth AI@unsloth.ai · last wk.

Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...

We signed the Open Weights letter because we believe the future of AI should be shaped by everyone, not controlled by a select few. That belief has always been at the heart of Unsloth: everyone should be able to train and run models on their own local device.

Bild

We collaborated with AWS on a complete guide to LLM Quantization and Deployment. Learn about: • Model formats, dynamic quants & making your own • Choosing GGUF, NVFP4 or FP8 • Picking the right tools & deploy on AWS SageMaker • Benchmark quality, latency & cost Read: aws.amazon.com/blogs/machin...

Bild

Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide: unsloth.ai/docs/models/...

Bild

We collaborated with NVIDIA to teach you how we made LLM training ~25% faster! 🚀 Learn how 3 optimizations help your home GPU train models faster: 1. Packed-sequence metadata caching 2. Double-buffered checkpoint reloads 3. Faster MoE routing Guide: unsloth.ai/blog/nvidia-...

Bild

We made a guide on how to run open LLMs in Claude Code, Codex and OpenClaw. Use Gemma 4 and Qwen3.6 GGUFs for local agentic coding on 24GB RAM Run with self-healing tool calls, code execution, web search via the Unsloth API endpoint and llama.cpp Guide: unsloth.ai/docs/basics/...

Bild

DeepSeek releases DeepSeek-V4. 🐋 - DeepSeek-V4-Pro: 1.6T params - DeepSeek-V4-Flash: 284B params DeepSeek-V4-Pro rivals Claude-Opus-4.6-Max, GPT-5.4-xHigh and Gemini-3.1-Pro-High. They support 1M context length, thinking and set new records for Codeforces.

Bild