Any recommendations regarding people thinking deeply about how agents will affect organizational structure and process? There's a lot right now about "how does Joe and Jill get better at using LLMs" but not as much about "this is how an organization needs to change to maximize gains from LLMs".
Erik Istre
@eistre91.bsky.social
Software engineer talking about LLMs, video games and horror media. My opinions are my own, etc. etc.
Anthropic please release Haiku 5 in August. Pretty please with sugar on top? My token budget at work is dying over here.
If there's any settled rule for good software it's this: software should be the minimal implementation that satisfies a given set of desired outcomes. This has always been true, and agents have not changed this fundamental reality. Good software is as much about absence as it is about presence.
We should have a Big Walk and Talk digital meetup with all of the interesting folks in the AI space on bsky to discuss LLMs, agents, and generally Life, the Universe and Everything.
House House's Big Walk is equally smart as their Untitled Goose Game in the way it wields its silliness, similarly playful and exuberant, and stuffed with innumerable delightful little touches. Otherwise, it's not much like Untitled Goose Game at all. Our review: https://bit.ly/3RFnP49
The pace of progress for LLMs has really skewed my judgment for how fast anything else should move. Some of it is living in a constant state of urgency with regard to security and some of it is base lizard brain rewiring that now thinks any process which isn't exponentially improving is a laggard.
In-depth alignment sessions with LLMs are one of the most powerful techniques I'm currently aware of. (Matt Pocock's Grill Me, bidirectional prompting) And also the sessions are so damn cognitively exhausting. Making a compressed sequence of high level executive decisions burns the brain juice.
Some initial thoughts, and a complicated mix of feelings. Wow. I mean, Erdos problems are cool (I genuinely mean that), I didn't know about the Jacobian conjecture before it got disproved. But this newest batch from OpenAI hits home in a way the previous announcements did not.
I got a bit down yesterday with everything going on. So here's the optimism that I consistently choose to anchor myself to: humans are capable of exploring amazing possibilities. I commit to believing in and remain ever bullish on what humans can achieve.
Why Grogor use club when should smash food with hands? Grogor deskilling himself.
Man I've had a good run of optimism lately and I'm sure I'll back there but today feels bleak for a myriad of reasons. People virtue signaling by attacking Hank, everywhere is hot as hell, American politics, and my own personal annoyances with big companies moving slow on LLMs.
I've seen people complaining about Opus 5 performing in ways that I haven't encountered. This is probably workflow specific in a lot of ways. But also I wonder if people are using different effort levels like xhigh/max and underestimating the impact?
How did I miss that Sephiria exists till now?! The melee combat looks slick and satisfying, online co-op, interesting inventory management system for power, maybe some mild town building for meta-progression? Extremely my type of game.
Save 40% on Sephiria on Steam
Save Bunnyville from a doomed fate: a top-down action/RPG roguelite from the makers of Dungreed. Explore the depths of the tower, collect and upgrade artifacts, and defeat hordes of demons to keep you...
store.steampowered.com
My favorite part about coding with LLMs is the ability to devote more of my attention and energy every day to software design. I really love thinking in terms of abstractions and domain modeling. Biased but I think people who like thinking in this way take to LLMs better and get better outcomes.
Anyone who doesn't post while Claude is down is provably an AI agent.
must read slides from Terrance Tao for anyone that cares about math. i love the framing and this way of thinking about the issues. my only wish is there were more slides of discussion at the end in the "recommendations" section.
slides from Terrance Tao on "Mathematics in the age of AI" from a public lecture, July 24, 2026
Anyone have a way to get LLMs to stop preserving historical information? My current pet peeve with using LLMs is that they are loathe to leave information behind. I've struggled to get them to have a good sense of what it means to only include information that future humans or agents need.
What is the "way of thinking" that maximizes the leverage in working with agents? Learning science, math and software engineering have value even if you don't make a career out of those things because they teach new thought processes that disentangle the squishiness of natural human thought.
The year is 2030. A robed figure supplicates themself before the whirring drives and flashing lights as a whirlwind of sacred telemetry scrolls across the screen.
Hot takes on Opus 5. Promising release. - So Sonnet 5 is just completely pointless now isn't it? Opus 4.8 was already better on a cost per task/performance ratio and this looks better. - What's up with the notable quality decrease on coding benchmarks moving from xhigh -> max?
Most people ARE really bad at working with LLMs and even software engineers who use them regularly don't realize that they're bad at it. People are lost in a sea of "unknown unknowns".
Hot take: neurodivergent people benefit from skill transfer when using LLMs because they've spent years having to carefully build a mask that makes them agreeable to society. A large part of that is being overly thoughtful about every single word you use and how it might be interpreted.
People absolutely keep gilding thy lilies on how hard their problems are and they're mostly just not. Opus 4.5 could cook most things people throw at these models now but it feels like you're Doing More by roasting 3T-parameter tokens over an open fire Workflows continue to matter more
Some days with agentic coding tools it feels like we're on the verge of a qualitative shift in both the quantity and quality of software that's available. Other days it feels like that future is years away as we toil to change a mountain of cargo culted process which no longer serves us.
Yea the GPT 5.6 family is excellent. Right now using Sol xhigh or max for planning and as a design partner and Luna Max as implementer. I'm sure I'll find the edges in the coming weeks. But the models are aligning themselves with my intent super well and the behavior is exactly what I want.
If we assume things play out like they did with search (Google wasn't the first mover and arguably benefited from that), who are the first movers in AI that crash and burn (OpenAI?) and who wins in the end? Or are there reasons to believe that "trust me bro, this time is different"?
Our industry appears to have collectively forgotten that waterfall is a terrible way for making good software. We got stochastic code generation tools and then immediately regressed a few decades in software planning and implementation. I'm looking at you spec driven development.