Edward J. Schwartz
@ejschwar.bsky.social
Computer security researcher at CMU's Software Engineering Institute; {computer,car lease} hacker; rescue dog daddy; soccer player/referee; skier. https://edmcman.github.io/
www.latent.space/p/attention-... 1. Interesting view of AI timeline 2. I agree that models and harnesses are being intertwined, and I'm not sure it's a good thing.
The Evolution of the Agent Harness
Models keep absorbing the harness into their weights — soon, it will be a harness for human attention rather than for the model.
latent.space
Well I'll be. NDSS is moving from San Diego! www.internetsociety.org/blog/2026/06...
NDSS Symposium 2027 Heads to Seoul: Expanding Global Collaboration in Cybersecurity Research - Internet Society
We are pleased to announce that NDSS Symposium 2027 will take place in Seoul, Republic of Korea, from 22–26 March 2027.
internetsociety.org
Harbor's hub now lets you upload benchmark runs, which is pretty neat. Here's an example run from the auto-bench blogs I've been writing: hub.harborframework.com/jobs/4149711... You can view each run and its trajectory: hub.harborframework.com/jobs/4149711...
Harbor Hub
Browse Harbor Hub
hub.harborframework.com
github.com/QwenLM/Qwen3... Something is wrong when outsiders need to reverse engineer how you are getting your published benchmarks.
Reproducing Qwen3.6-27B SWE-bench Pro (53.5): a `str_replace` file-edit tool + reliable-test conditioning closes the gap (bash-only ~28% → ~51% pass@1 / ~71% pass@8) · Issue #179 · QwenLM/Qwen3.6
✅ Update (2026-07-08) — largely resolved With a str_replace file-edit tool + reliable-test conditioning, the official Qwen/Qwen3.6-27B reproduces the model-card number: bash-only ~28% → ~50.7% pass...
github.com
Congratulations to Dr Luke Dramko (my lucky 13th PhD graduate!) for his successful and stellar defense yesterday of “Neural Decompilation with Minimal Risk.” The last chapter isn’t published yet but STAY TUNED because it’s a symbolic approach for proving the fidelity of neurally-decompiled code. 1/
🚨 Blog Post: "Benchmarking Quantized LLMs for Local Coding Agents Part 4: Investigating the Impact of Thinking on Qwen3.5-35B-A3B" https://edmcman.github.io/blog/2026-07-22--benchmarking-quantized-ll-ms-for-local-coding-agents-part-4-investigating-the-impact-of-thinking-on-qwen3-5-35-b-a3-b/
I'm old enough that I remember legitimately using this! www.osnews.com/story/145534...
Microsoft releases its weird ’90s IRC client as open source – OSnews
osnews.com
🚨 Blog Post: "Benchmarking Quantized LLMs for Local Coding Agents Part 3: Investigating the Performance of Qwen3.5-35B-A3B (Take 2)" https://edmcman.github.io/blog/2026-07-04--benchmarking-quantized-ll-ms-for-local-coding-agents-part-3-investigating-the-performance-of-qwen3-5-35-b-a3-b-take-2/
🚨 Blog Post: "Benchmarking Quantized LLMs for Local Coding Agents Part 3: Investigating the Performance of Qwen3.5-35B-A3B" https://edmcman.github.io/blog/2026-06-12--auto-bench-benchmarking-quantized-ll-ms-for-local-coding-agents-part-3/
🚨 Blog Post: "auto-bench: Benchmarking Quantized LLMs for Local Coding Agents Part 2" https://edmcman.github.io/blog/2026-06-04--auto-bench-benchmarking-quantized-ll-ms-for-local-coding-agents-part-2/
🚨 Blog Post: "ASP: Towards the Next Generation OOAnalyzer" https://edmcman.github.io/blog/2026-05-28--towards-the-next-generation-oo-analyzer/
🚨 Blog Post: "Self-hosted MCP Servers for Daily Life" https://edmcman.github.io/blog/2026-05-20--self-hosted-mcp-servers-for-daily-life/
I went to the trouble of automating the creation of CAPEv2 sandbox so you don't have to: github.com/edmcman/cape...
GitHub - edmcman/cape-sandbox-vm: Packer project for building CAPEv2 malware analysis sandbox VMs
Packer project for building CAPEv2 malware analysis sandbox VMs - edmcman/cape-sandbox-vm
github.com
my google account is heaven@gmail (long story), so I get a steady trickle of emails from people writing to God. some of them are real winners
🚨 Blog Post: "auto-bench: Benchmarking Quantized LLMs for Local Coding Agents" https://edmcman.github.io/blog/2026-05-15--commonsense-benchmarks-for-local-coding-agents/
Ubuntu is being ddosed and their repositories are down. Bad news if you need apt because you want to build a devcontainer or run your CI!
I really wish there was a leaderboard for GGUFs on swe-bench or something similar...
www.reddit.com/r/ClaudeAI/c... AI providers are beginning to slow down their subsidies on LLM usage. For now, it is only external agents like Opencode. But LLM access as a subscription model is a loss leader. All these vibe coders are going to be very upset when they have to pay per token.
From the ClaudeAI community on Reddit: Claude subscriptions will no longer be usable in Opencode.
Explore this post and more from the ClaudeAI community
reddit.com
I designed my first real part! www.printables.com/model/162187... No, I don't really know what I'm doing. Yes, it is satisfying anyway!
Case for Waveshare ESP32-C6-Geek Development Board by Edward Schwartz | Download free STL model | Printables.com
printables.com
TIL about pypi-timemachine. You're welcome. (You know, for those 10-year old research projects without lock files)
snap: When you only care about 90% of your apps to work. It's been years. Why do major snaps (firefox!) still have usability issues?!