Gal Sapir

@sapir.bsky.social

https://sparsethought.com/

more and more I feel like the days resulting in my best work, are the days in which I hold back on "letting agents roam free" and being much more careful and well, full of care

1/ a nature medicine paper claiming general-purpose llms beat specialized clinical tools (openevidence, uptodate) is going around. i pushed on it and landed somewhere different. what does it actually measure?

General-purpose large language models outperform specialized clinical AI tools on medical benchmarks

small skill: for long-horizon agent runs, have it keep an implementation-notes file as it goes (decisions, tradeoffs, deviations from spec). skim after to check broad strokes + project conventions held. usually enough for me to be ok with it. github.com/galsapir/skills/tree/main/skills/long-horizon

skills/skills/long-horizon at main · galsapir/skills

Claude Code plugin: deep project interview command that produces actionable specs before implementation - galsapir/skills

github.com

medmarks v1.0 dropped last week, the largest open medical eval suite to date, and the verifiers framing is the right call imo. sitting with it for a few days pulled out a thread that hasn't resolved: the curation problem goes all the way down sparsethought.com/2026/05/16/curation-all-the-way-down/

curation all the way down: on clinical AI benchmarks

the curation regression, the openness trade-off, and what a substrate worth evaluating against would actually need: on Medmarks.

sparsethought.com

new post: small workflow changes that have started to add up. nothing deep, a few adjustments to how i work with agents that converged into something that feels more comfortable to me

week two with a second brain for agents (now on cloudeflare D1!) the part that surprised me: it works as connective tissue between agentic systems (cc, codex, other), so i can be largely agnostic to which one i'm in. +low-grade dread about losing threads is lighter. sparsethought.com/2026/05/01/s...

a second brain, week two

week two with a memory MCP as second brain: the background dread is lighter, reviews are cheaper but shallower, and why i’m not ready to let the system auto-fix itself.

sparsethought.com

context doesn't transfer between claude code, desktop, mobile, codex. (inspired partly by @austegard.com + @timkellogg.me ) i wrote a memory MCP so i could say "remember this" once. five days in, its mostly a log of things stuff i missed-hopefully just once sparsethought.com/2026/04/23/s...

a week with a second brain

notes from five days of running a memory MCP across Claude Code, desktop, mobile, and Codex: what’s in there, what’s already not working, and why the corpus is mostly corrections.

sparsethought.com