Marco Abis

@marcoabis.it

Techie, recovering entrepreneur - Runner 🏃 Currently living (again) the life of the indie solo product developer at nearly 50 Trace your local AI workload on Apple Silicon 👉 https://ziraph.com

Building Ziraph, I kept scratching little itches around LLM-assisted dev and Claude Code. So I'm extracting a few: • a little comms bus for agents to coordinate • a meta-IDE with observability panels • an SDLC policy enforcer The last one might be useful to others, so I'll polish it up.

If only I'd waited 48h before posting this. My "on-device AI is a margin decision" post weighs 6 axes but assumes the cloud is a stable utility you rent. The US suspending Fable 5 + Mythos 5 adds a 7th: availability / continuity / sovereignty. https://ziraph.com/blog/on-device-ai-margin-decision

On-device AI is a margin decision - Ziraph blog

The aspects engineering leaders should weigh before betting on on-device AI: margin, privacy, battery, a fragmenting stack, the hardware your users actually carry - and the one phrase it all hinges on.

ziraph.com

Apple fm is the odd one out among local AI runners. It runs on the Neural Engine, not the GPU. No token count without a second call. It trades speed for power, by design. Apple tells you which silicon ran your model, not what it cost. https://ziraph.com/blog/apple-fm-odd-runner

Apple fm: the odd runner - Ziraph blog

The newest runner I added to ziraph, and the odd one out: it hides its token count, runs on the Neural Engine instead of the GPU, and quietly trades speed for power.

ziraph.com

Ran Apple's new fm CLI (Foundation Models) on macOS 27 and watched where the work actually went. It's a Neural Engine workload, full stop: ANE ~3W, GPU 0.17W (whose only clients are WindowServer + my terminal drawing this). Apple's on-device LLM runs on the ANE 🧵

Ziraph's terminal dashboard on an Apple M1 MacBook Air. The ANE panel reads ~2.97 W (peak 3.04 W); the GPU panel reads 0.17 W with only WindowServer and the ghostty terminal as GPU clients, and GPU clock OFF 84.7% of the time. The workload row shows "fm respond" with Ziraph detecting the worker as TGUIDeviceInferenceProviderService, which appears in the ANE client list - i.e. Apple's Foundation Models inference is running on the Neural Engine, not the GPU.

Apple's two on-device models, per Subramanya at WWDC: AFM Core - dense AFM Core Advanced - sparse, natively multimodal Both are "custom builds for Apple Silicon, trained using proprietary data, and refined using outwards from Gemini frontier models." 9to5mac.com/2026/06/08/c...

Craig Federighi details Apple’s collaboration with Google for Siri AI in iOS 27 - 9to5Mac

Apple’s Siri team, led by Craig Federighi, held a post-WWDC keynote tech talk with members of the press this afternoon...

9to5mac.com

Apple quietly rebuilt how AI runs on Apple silicon - and the docs are clearer than the keynote. On-device AI is now a platform. The hard problems are the ones we all fight. And the local/cloud line just blurred. What changed, in 8 parts 🧵 https://ziraph.com/blog/apple-on-device-ai-wwdc-2026

Apple rebuilt its on-device AI stack at WWDC 2026 - Ziraph blog

No new chips - but a structural rebuild of how AI runs on Apple silicon, and the interesting parts are in the details Apple didn

ziraph.com

Every #LocalLLM tool prints tokens/sec, none the bill. tok/s is vanity. Energy per token is sanity. Joules are reality - the metric I'm building around on #AppleSilicon. Matched-quant Gemma 4: decode a tie, 4.5x the CPU energy. https://ziraph.com/blog/energy-per-token-vanity-sanity-reality

Why I care so much about energy per token - Ziraph blog

Every local AI tool prints tokens per second; none of them prints what those tokens cost. Why energy - and energy per token specifically - is the metric I built Ziraph around.

ziraph.com

to trust the numbers Ziraph reads off Apple Silicon I check the resolver on real silicon. here are my gaps (✗ = no real hardware). got one of these Macs? reply or DM and I'll set you up 🙏 #AppleSilicon #LocalLLM #MLX

Coverage grid of Apple Silicon chip variants (M1 to M5 in base/pro/max/ultra, plus A18 Pro) marking which ones Ziraph has validated on real hardware (green check) versus gaps with no real-hardware run yet (red cross).

yesterday I said I'd share more on my little product soon and... while I still need a bit more time before opening the website and beta 🤞 I'm gonna share one of the draft posts I've been writing in preparation, here goes nothing 🦒

Hello, World! 👋 haven't posted here in years - and when I did, it was mostly reposting other people, not my own stuff. let's see if I can get back into the groove after a long stretch obsessively building a little product 🦒. More on what it is soon.

What do you do when your media org is captured? You start your own. Introducing....The Nerve!!! @thenerve-news.bsky.social We're all-female, journalist-owned & launching next week. Please help us build a truly independent, progressive new media!👊👊👊

Press Gazette@pressgazette.co.uk · 10mo ago

Former Observer big-hitters inc @carolecadwalla.bsky.social launch new title The Nerve with redundancy payouts After Observer sale: "We decided to launch a title that the journalists themselves would own, that would be truly independent" pressgazette.co.uk/news/former-...

As you may have seen in last night's video, MY NEW BOOK IS HERE! This practical handbook is packed with real-world advice and actionable insights, distilling decades of software engineering experience. Early readers have found it a valuable complement to my other books... 1/2

Yesterday I told the story of an apparent request from a Trump-run government agency that Amazon de-list my book. As a result, in the last 16 hours, the book has shot to number 6 across all of Amazon out of millions of books. I am absolutely stunned.

Bild