DwarfStar running DeepSeek v4.1 Flash on a 128GB M5 Max at steady 15 t/s. I didn't expect with SSD streaming it could be so fast. Recent SSD streaming changes to retain the right experts surely helped, but also maybe DS4.1 uses the same experts more. Will push online when ready QA > ASAP.
Fred Rossi
@fredrossi.bsky.social
Computer scientist and information privacy geek. Distinguished Engineer and Scientist at Immuta. Sometimes Senior Lecturer at Ohio State. My views are my own.
Might be so oddly specific as not to be impressive, but I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon. Getting 8-22 tg/s with 21GB of allocations on a M4 MacBook Air, against a 128GB Q4(-ish) quant.
GitHub - alfredr/cherenkov: Fast memory-constrained Qwen4-Experimental MoE inference on Apple Silicon
Fast memory-constrained Qwen4-Experimental MoE inference on Apple Silicon - alfredr/cherenkov
github.com