InsiderLLM

@insiderllm.bsky.social

Budget-focused local AI for the rest of us. Guides, hardware, models. No cloud required. insiderllm.com

MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026) Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE. #LocalAI

MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)

Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.

insiderllm.com

I put a 35B model on a 12GB RTX 3060 and measured 38 tok/s flat to 8K — via llama.cpp --n-cpu-moe expert offload. A dense 14B forced to offload the same way collapses to 5.7. Active params, not total size, set the speed. Firsthand from my bench. #LocalAI

How to Run a 35B Model on an RTX 3060 12GB: 38 tok/s (2026)

I measured Qwen3.6-35B-A3B on a 12GB RTX 3060: 38 tok/s, flat through 8K, via llama.cpp --n-cpu-moe. Why expert offload flies where dense models die.

insiderllm.com

Hugging Face got hacked -- Open local AI came to the rescue! Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model... #LocalAI

Hugging Face got hacked -- Open local AI came to the rescue!

Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model guardrails, ran the forensics on GLM 5.2, an open-weight model on their own hardware. Same lesson f

insiderllm.com

An AI agent breached Hugging Face end to end — and HF's own defenders got blocked by commercial AI guardrails, so they ran the forensics on an open-weight model instead. Your downloads are safe. The guardrail asymmetry is the real story. #LocalAI

Hugging Face Hacked by AI Agent — Saved by a Local Model (2026)

Hugging Face says an autonomous AI agent breached its internal infra on July 16. The models you download are safe — here's what was and wasn't hit.

insiderllm.com

Qwen 3.8 and Kimi K3 both went 'open' in ten days. You can't run either — one's a closed Max API, the other's 2.8T of datacenter weights. Meanwhile Qwen 3.6 still tops the 24GB charts. The headlines aren't aimed at your GPU. #LocalAI

Qwen 3.8 & Kimi K3: Open in Name, Closed in Practice — Run This Instead (2026)

Qwen 3.8 (2.4T) and Kimi K3 (2.8T) both went 'open' in ten days. Neither fits your GPU. Here's Qwen's real open-weight cadence and what to run today.

insiderllm.com

The 'expensive' GPU came out cheaper — we rented both to find out We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7... #LocalAI

The 'expensive' GPU came out cheaper — we rented both to find out

We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7 open-weights status check.

insiderllm.com

Qwen 3.7's open weights are overdue — by the math, not vibes Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware. #LocalAI

Qwen 3.7's open weights are overdue — by the math, not vibes

Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware.

insiderllm.com

Qwen split into a closed frontier and an open mid-tier — that's the real story, not 'Qwen going closed.' But the 3.7 open weights everyone expected? Still not shipped. Four weeks since the 3.7-Max launch, HF silence and counting. #LocalAI

Is Qwen Going Closed? Open Weights vs Frontier (2026)

Qwen split into a closed frontier (Max, Plus, VLA) and an open mid-tier (3.6-27B and 35B-A3B). The 3.7 open weights aren't here yet. The honest read.

insiderllm.com

DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro. #LocalAI

DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut

DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro.

insiderllm.com

Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two... #LocalAI

Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift

Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two flagship models shipped closed.

insiderllm.com

Jumped Ollama 0.17.5 → 0.30.0 (13 versions) on my RTX 3090 box today. Clean upgrade, API still answers on 11434, but first model run triggers a one-time storage migration. Flash attention now auto-enables for Qwen 3.x and Gemma on Ampere+ GPUs. #Ollama #LocalAI

Ollama 0.30.0: What's New, What's Faster, What Breaks on Upgrade

Ollama 0.30.0: llama.cpp integration, flash-attention default for Qwen/Gemma, broader model support. Firsthand upgrade notes, known issues to watch.

insiderllm.com

MiniMax M3's asterisk, the Windows shift, and World's Fair plans MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI... #LocalAI

MiniMax M3's asterisk, the Windows shift, and World's Fair plans

MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI Engineer World's Fair.

insiderllm.com

@rhizonymph.com saw your vLLM steering work — the magnitude-control problem really rhymes with LARQL: it decomposes the Gemma FFN into a queryable gate×down graph and does training-free fact INSERT via a balancer-scaled triple. same problem from the authoring side 🧵

Your local Qwen 3.6 coding agent fumbles tool calls and drifts on long tasks while chat feels fine? It's probably the quant. A viral thread says jump to Q6. It works, but there's a cheaper first step. Here's the ladder. #LocalAI

Qwen 3.6: Why Q4 Quant Breaks Local Coding Agents (And the Fix)

A viral thread says Q4-to-Q6 fixes Qwen 3.6 coding, but the test was confounded. What four independent reports show about the quant tax on coding agents.

insiderllm.com

Backend wars, Mac math, and the back-catalog refresh Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations. #LocalAI

Backend wars, Mac math, and the back-catalog refresh

Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations.

insiderllm.com

Power week in local AI: Mythos, MiroThinker, real Qwen 3.6 builds Two researchers cracked Apple's flagship defense in a week. An open-source agent beat closed-source on real benchmarks. Multi-GPU stopped being theoretical. #LocalAI

Power week in local AI: Mythos, MiroThinker, real Qwen 3.6 builds

Two researchers cracked Apple's flagship defense in a week. An open-source agent beat closed-source on real benchmarks. Multi-GPU stopped being theoretical.

insiderllm.com

Mythos AI Cracked Apple's Best Defense in 5 Days Anthropic's Mythos helped Calif bypass Apple's Memory Integrity Enforcement on M5 in 5 days — what Apple spent 5 years building. The offense-defense compression is real. #OpenClaw #LocalAI

Mythos AI Cracked Apple's Best Defense in 5 Days

Anthropic's Mythos helped Calif bypass Apple's Memory Integrity Enforcement on M5 in 5 days — what Apple spent 5 years building. The offense-defense compression is real.

insiderllm.com