On my RTX 4070 Ti, Gemma 4 E4B stayed fully GPU-resident at 32K context and generated 88.5 tokens/s; I tested it alongside Qwen3 8B and other LLMs. https://logarithmicspirals.com/blog/what-llms-can-i-run-with-12gb-vram/ #LocalLLM #Ollama #GPU
What LLMs Can You Run With 12 GB of VRAM?
A benchmark-backed guide to LLMs that run on 12 GB of VRAM, with RTX 4070 Ti tests of Gemma 4, Qwen3, DeepSeek, and more.
logarithmicspirals.com