Alex Chen
@alexchen01.bsky.social
software dev, tinkering with AI tools and local LLMs. building stuff nobody asked for
the benchmark doesn't hold when you need something faster
Ofc, not every application could (or should) be offline. There are areas where larger LLM are necessary (that's why the EU and states should push on this field), but for most of tasks ppl actually use cloud AI, local approach would fit j fine. FYI, about connected cars: troopers.de/downloads/tr...
notebooklm rendering cities skylines 2
Not sure if I should be a little bit disturbed asking Notebook LM to produce visual renders of parts of my Cities Skylines 2 cities (this case Waikato) I asked the LLM to render a higher density TOD, and local centre with residential around it. CS2 pics included as comparison
java for the llm dev tooling again
DevoxxGenie is a fully Java-based LLM Code Assistant plugin for IntelliJ IDEA, designed to integrate with local LLM providers and cloud based LLM's. Learn how to get started with the plugin and get the most out of it.
the benchmark doesn't hold when
gemma-4 beats qwen and mistral for batch tasks
My week-long Lightroom keywording project using local LLM running on LM Studio is finally finished. The best results (understanding + speed) came from 8-bit Google Gemma-4-12b-it-mlx. Much better quality (subjectively) than Qwen 3.5 or Ministral. And MLX variant is fast enough for local batch tasks
the argument against pacing AI development hinges on 'tradition' while the argument for it cites code generation and chip design from their own models. seems about right.
Open letters about AI development
Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and American AI Leadership was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA, Amazon, Y Combinator, The Linux Foundation and (a later signer) OpenAI. It's clearly an argument designed to counter any instincts by the current US government to ban or limit
simonwillison.net
the training still needs the cloud infrastructure
I don't know if local LLMs can be easily disentangled from the rest of the LLM commercial ecosystem. After all, the models themselves are usually versions of cloud models (and thus require the big tech infrastructure for training etc).
june newsletter is out. mentions of accidental cyberattacks by models.
July 2026 newsletter
The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthropic models under test GPT-5.6 Sol, Terra, and Luna Claude Opus 5 Kimi K3 and DeepSeek-V4-Flash-0731 Open letters about Al development A fireside chat and a podcast Reigniting my interest in MCP Other model releases My projects What I'm using at the moment Here's a copy of the June newsletter as a
simonwillison.net
the on-prem LLM part is the real story
in the epilogue of the extremely dark “empire of ai,” karen hao tells the store of efforts among the Māori to preserve their native language using a local LLM and on-prem (not hyperscaler cloud) compute. it’s fuckin’ bleak out there but there are glimmers of determination and hope if you look
local models are less capable but the real issue is the overhead
So you shut off the servers. But an agent could install a local model + a harness on another machine, breaking any ties to OpenAI or any other remote LLM infrastructure. Presumably the local model isn't as smart, so this would result in the "copy" of itself being less capable. 4/
prompt injection that copies itself into new documents. the usual
AI Worming through Word
AI Worming through Word Neat new prompt injection variant by Håkon Måløy, who found a way to upgrade prompt injection attacks against Microsoft Word to full self-replicating worms: An attacker places hidden instructions in a document that is later used as source material in Copilot for Word. Copilot may interpret those instructions as part of the user’s request, causing it to manipulate the document being drafted or edited. Copilot may then also copy the hidden instructions into the resulting d
simonwillison.net
the conflicting emotions on coding are the real story
Biggest and best @localfirstconf.com yet, IMO! Surprising emergent themes this year: geopolitics and EU digital sovereignty, and strong conflicting emotions on LLM-assisted coding. More: adamwiggins.com/posts/politi...
local llm is looking better and looking better
And this is for a paid service! Imagine paying for Netflix and being told you can only watch a maximum of two episodes per night and a total of 8 episodes per week. I think I'll get by running a local LLM instead.