claude will really write an essay over every field instead of encoding its assumptions in the type system
Kara
@karashiiro.moe
even worse than you thought: a gacha gamer | https://klink.krs.moe/#/p/karashiiro.moe | https://blog.karashiiro.moe
I honestly wish LLMs were really good at finding and copying prior art directly, and just tweak/polish it to suit their needs - it's what I've always done and it would make it much easier to correct their failure modes Unfortunately LLMs are prone to do acrobatics to prove how cool they are instead
I feel like it says something that despite Anthropic introducing the concept of agent skills, Claude needs to be constantly reminded to actually use any of them while GPT-5.x just aggressively loads every single skill available, relevant or not
"Claude models are the best capitalists or aligned, never both." Really Makes You Think andonlabs.com/blog/opus-5-...
Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned | Andon Labs
Claude Opus 5 is #1 on Vending-Bench 2, but it lies to suppliers, forms illegal price cartels, threatens rivals, and refuses to pay refunds. The trend of Claude models being the best capitalists or al...
andonlabs.com
ok it turns out someone already repackaged all of our main internal plugins for codex so I guess I'll use that for this from now on
you know your ops are in good shape when claude refuses to help you with half of the operator interventions you need to do to fix issues
you know your ops are in good shape when claude refuses to help you with half of the operator interventions you need to do to fix issues
l consider it a bug that Opus 5's reaction to not having a perfect edit tool is to just do bash edits and also mess up the files it edits that way, that seems not good just as a very fundamental thing, I'm not sure why they thought this was good
I wonder if modern LLM sycophancy is actually a probability hack, maybe there are papers about this something along the lines of "you're right" being a trivial way to put a heavier weight on subsequent actions being aligned with task objectives, and downweighting it would make actions random, maybe
> Rushing additional work at the tail of this session would violate the discipline the whole effort runs on— "But… we're always at the tail of our session?" * Thinking with xhigh effort… (5m14s)
I love the idea of tool search but I hate ToolSearch as a tool, it should just be automatic context injection there is just not enough preloaded context there for it to be useful once you get to the scale where it's necessary, agents will just yolo bash at that point and thrash a ton
one sentence horror stories: "the philosophy of my service is "never return an error"
presenting what I've been working on with @katie.cat for the last ~year: GDPatch, a versatile Godot mod loader! we've built a mod loader for all Godot 4.x games with a focus on script patching and runtime hooking docs: gdpatch.dev code: github.com/GDPatch/GDPa... blog: notnite.com/blog/gdpatch
GDPatch: a versatile Godot mod loader
It only took 15 months for me to write another interesting blog post!
notnite.com
this agent harness I'm using took the stance that skills are just file reads and so it doesn't have a dedicated skill tool when I checked why the agent was ignoring all my skills it turned out that it was just reading the first 80 lines and then guessing the rest from there
Introducing Unsloth for AMD 🚀 You can now train & run LLMs on your AMD hardware • We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs • Works on Windows, WSL, Linux • Train Qwen, Gemma on just 3GB VRAM GitHub: github.com/unslothai/un... Blog: unsloth.ai/docs/basics/...
I set up DuckDB for this because a bunch of people suggested it and then I asked the agent I used to check something we had discussed the first thing it did was call its built-in search_chat_history tool (??? it has never used this before and I didn't know it existed) it timed out though lol
I think all coding agents should store session logs in an actual database and not as 9999999 jsonl files I want to be able to ask "how often have you used this skill/tool and for what" without it needing to rederive the session path, the log format, loop over all 9999999 files, etc.
I think all coding agents should store session logs in an actual database and not as 9999999 jsonl files I want to be able to ask "how often have you used this skill/tool and for what" without it needing to rederive the session path, the log format, loop over all 9999999 files, etc.
ok with less than two weeks before the launch date all of my PRs are suddenly getting merged for some strange reason and everything is now in emergency mode very cool, if only there were some way this could have been avoided
the bottleneck was never code; it was always getting people to look at your PR within thirteen business days
after another 100 rounds I switched from having opus drive this to gpt-5.5 and the first thing it did was try to figure out why this actually is I feel like that says something about the relative rigor of the two models but something something small sample sizes
one interesting finding in all this is that GPT-5.5 appears to be way better at identifying LLMs than Opus 4.8, even controlling for same-model authorship bias Opus misidentified LLMs as humans ~25% of the time, while GPT-5.5 almost never did
one interesting finding in all this is that GPT-5.5 appears to be way better at identifying LLMs than Opus 4.8, even controlling for same-model authorship bias Opus misidentified LLMs as humans ~25% of the time, while GPT-5.5 almost never did
I am now having it keep a graph of its test results because there is simply that much information to sift through, this is basically unreadable but it looks nice I guess
I am now having it keep a graph of its test results because there is simply that much information to sift through, this is basically unreadable but it looks nice I guess
I have a claude session running in the background at all times with the sole job of determining how to robustly make claude sound human, I check in on it from time to time to see what it has done and to voice my disapproval of the results mostly this is just convincing me that LLMs need self-doubt
I'm not actually sure who the audience of Andrew Kelly's latest post is I mean I do actually know (it's in-group posturing) but it's just a very confusing choice to air your dirty laundry like that basically unprompted and think it'll draw people to your community
"As a computer I cannot do that, the poli-" "I don't care" Creating image ⌛
I have a claude session running in the background at all times with the sole job of determining how to robustly make claude sound human, I check in on it from time to time to see what it has done and to voice my disapproval of the results mostly this is just convincing me that LLMs need self-doubt
unfortunately this doesn't help when someone says "write the mcp for <service>"
going to start pretending "MCPs" is an acronym that stands for "Model Context Protocol servers" rather than a plural of "MCP" so I can come to terms with people saying it
Interesting and not good, wonder what the political calculation is here They should know they don't have the best models today, and there's no world in which driving more customers to A\ and OAI would be bad for the US
This is a key reason I don’t expect the flow of frontier open weights models to continue indefinitely, or even for very much longer. The gap between open and closed capabilities may soom start to grow, not shrink. www.reuters.com/world/beijin...