Daniel Moll

@danielmoll.com

Building the web

Cloudflare Wallets are now available. You can secure your account / name via cloudflare.pay if you like. If that leaves you with a bunch of question marks – take a look at the blog. What Cloudflare is "tinkering" with right now actually makes perfect sense.

Just hit 59 tok/s on 3× GTX 1080 Ti (Pascal, 2017) Qwen3.6-35B-A3B MoE + MTP, no CPU offload: • 59 tok/s TG128 @ 32K • 3× faster than with --cpu-moe • llama.cpp PR #22673 + community MTP GGUF Key learning: --cpu-moe killed MTP performance. Without it, the MoE's 3.8B active params fly.

I bought last Sunday (26/04/2026) HP Z840 Workstation from 2017 with this specs: 2× Intel Xeon E5‑2650 v3 (2,3 GHz), 128GB Samsung ECC DDR4‑2133 (8×16GB) 3× NVIDIA GeForce GTX 1080 Ti (33VRAM) Now i try the best to get local LLM working on this machine! For now Qwen 3.6 27B - runs with 20/ts

Bild

Don't be evil ... but ... Why do coding agents produce so much code? And why will commercial AI always generate to much code? They're selling tokens. The more tokens you use, the more profit they make!

Do you know OpenCode(.)ai? If not, you can hear me talk about this tool on November 25, 2025, at the 3rd edition of the Hamburger Agentic Coding Meetup. "OpenCode – The Freedom of Choice for Agentic Coding" --- See you ... there! Daniel Registration: www.moll.tv/some/ac/

Bild

OpenAI is not open. ↳ Its proprietary research. Artificial Intelligence is not intelligence. ↳ Its autosuggest. Social media is not social. ↳ Its algorithmic engagement engine.