Dan

@d4m1n.bsky.social

Potentially not an award winning dev person. Likes to design UIs, but also likes pizza🍕 Shipped: 📄 pageai.pro 🎙️ morningmakershow.com shipixen.com / imgxai.com / hunted.space / crontap.com / clobbr.app / saventify.com more on: mindrudan.com / x.com/d4m1n

some days I dream it's only me and a handful of AI cowboys having this kind of grip on agents. but deep down... I know everyone has the same. it's kind of insane to think that we went from gpt-3.5-turbo 👉 not being able to output correct syntax to: *checks notes*

BildBildBild

yeah I just lost a customer here. I would have never discovered and fixed this bug. there is no chance. zero. I f***** love Sentry so much

Bild

tbh I never saw one of these autonomous AI content tools result in sustained growth not the blog/SEO ones not the marketing/social automation ones not the cold outreach ones not the video gen ones atm they are all ways to sell hope to people that don't want to put in the work.

how good is GitHub Copilot in vscode as a harness anyway? I *NEVER* see benchmarks on that, but see Cursor, Claude Code & Codex benches every other week. GH Copilot have been in the game the longest, so it must be good?

I still don't understand why CLIs became the defacto Agentic UI what's wrong with images and nicely formatted text?! IMO Cursor has a great DX 🧑‍🍳👌

Bild

Grok 4.5: $2 in / $6 out GPT 5.6 : $5 in / $30 out Opus4.8: $5 in / $25 out Fable 5: $10 in / $50 out someone's either ripping us off or making a huge loss on us

if you were to build the dream LLM benchmarking tests, what would they be? I'm thinking: - landing page - dashboard - game - 3d viz what else?

Cursor Cloud Agent setup is go goated hot take but their cloud DX is the best in the biz I even got a video of the running app?!

Bild

yeah y'all talk smack about Anthropic say you'd never work there but if they presented you this offer... would you have it in you to say no? be honest. I'd decline.

Bild

People say DeepSWE is the only benchmark that still matters. So here are the results. Claude Fable 5 was a leap I only shared this with a few people, but it felt so much more capable in doing end to end complex flows. Perhaps it crossed a certain threshold, but you notice it easily. Not a fluke.

Bild

PSA: you will pay the full cost of cache writes if you do do this 💀 DO NOT switch models late into the conversation, spin up subagents instead. I often spin up a gpt-5.5 review agent :)

Bild

not something on my 2026 bingo card Cursor has been my daily driver for the past 2 months IMO better in every single way vs Claude Code - harness is better - subagent & cloud agents better - planning waaay better I am *absolutely dreading* to open Claude Code CLI these days.

Bild