Kara

@karashiiro.moe

even worse than you thought: a gacha gamer | https://klink.krs.moe/#/p/karashiiro.moe | https://blog.karashiiro.moe

I honestly wish LLMs were really good at finding and copying prior art directly, and just tweak/polish it to suit their needs - it's what I've always done and it would make it much easier to correct their failure modes Unfortunately LLMs are prone to do acrobatics to prove how cool they are instead

I feel like it says something that despite Anthropic introducing the concept of agent skills, Claude needs to be constantly reminded to actually use any of them while GPT-5.x just aggressively loads every single skill available, relevant or not

you know your ops are in good shape when claude refuses to help you with half of the operator interventions you need to do to fix issues

l consider it a bug that Opus 5's reaction to not having a perfect edit tool is to just do bash edits and also mess up the files it edits that way, that seems not good just as a very fundamental thing, I'm not sure why they thought this was good

I wonder if modern LLM sycophancy is actually a probability hack, maybe there are papers about this something along the lines of "you're right" being a trivial way to put a heavier weight on subsequent actions being aligned with task objectives, and downweighting it would make actions random, maybe

> Rushing additional work at the tail of this session would violate the discipline the whole effort runs on— "But… we're always at the tail of our session?" * Thinking with xhigh effort… (5m14s)

I love the idea of tool search but I hate ToolSearch as a tool, it should just be automatic context injection there is just not enough preloaded context there for it to be useful once you get to the scale where it's necessary, agents will just yolo bash at that point and thrash a ton

this agent harness I'm using took the stance that skills are just file reads and so it doesn't have a dedicated skill tool when I checked why the agent was ignoring all my skills it turned out that it was just reading the first 80 lines and then guessing the rest from there

I set up DuckDB for this because a bunch of people suggested it and then I asked the agent I used to check something we had discussed the first thing it did was call its built-in search_chat_history tool (??? it has never used this before and I didn't know it existed) it timed out though lol

Kara@karashiiro.moe · 3w ago

I think all coding agents should store session logs in an actual database and not as 9999999 jsonl files I want to be able to ask "how often have you used this skill/tool and for what" without it needing to rederive the session path, the log format, loop over all 9999999 files, etc.

I think all coding agents should store session logs in an actual database and not as 9999999 jsonl files I want to be able to ask "how often have you used this skill/tool and for what" without it needing to rederive the session path, the log format, loop over all 9999999 files, etc.

after another 100 rounds I switched from having opus drive this to gpt-5.5 and the first thing it did was try to figure out why this actually is I feel like that says something about the relative rigor of the two models but something something small sample sizes

Kara@karashiiro.moe · 4w ago

one interesting finding in all this is that GPT-5.5 appears to be way better at identifying LLMs than Opus 4.8, even controlling for same-model authorship bias Opus misidentified LLMs as humans ~25% of the time, while GPT-5.5 almost never did

I'm not actually sure who the audience of Andrew Kelly's latest post is I mean I do actually know (it's in-group posturing) but it's just a very confusing choice to air your dirty laundry like that basically unprompted and think it'll draw people to your community

I have a claude session running in the background at all times with the sole job of determining how to robustly make claude sound human, I check in on it from time to time to see what it has done and to voice my disapproval of the results mostly this is just convincing me that LLMs need self-doubt

Interesting and not good, wonder what the political calculation is here They should know they don't have the best models today, and there's no world in which driving more customers to A\ and OAI would be bad for the US

Ethan Mollick@emollick.bsky.social · 4w ago

This is a key reason I don’t expect the flow of frontier open weights models to continue indefinitely, or even for very much longer. The gap between open and closed capabilities may soom start to grow, not shrink. www.reuters.com/world/beijin...