aria

@aurelium.me

she/her sw infra @ Arcee, opinions my own

That'll do it. aint go agent getting into my infra any time soon! Would like to see GPT-6 try to get past this...

Bild

"destroying books to digitize them is a necessary evil imposed directly by publishers" is true but I think we are maybe overstating how bad it actually is coming up with a valid use for pallets of damaged/used books, sold by the pound, is recycling. they would've been thrown away otherwise

how many of you actually check the hash when people post hashes to call things in advance and then share the apparent original string

if you want you can just try using a model post-trained mostly by just SFT'ing on hundreds of millions of frontier model replies. it's like 80% of the models on huggingface generally speaking, they suck and barely work. you need extensive RL to make modern models work, there's no way around it

SE Gyges@segyges.bsky.social · 2w ago

in case anyone doesn't know this the thing about open weights models being made by stealing outputs from big players is almost entirely made up to justify regulations locking openai, anthropic and for some reason google into a cartel position

congratulations to the Moonshot team for extending Claude Fable 5's inclusion in claude subscriptions for another few weeks

Coding benchmark comparison with Kimi K3 highlighted. Kimi ranks third on DeepSWE, behind GPT-5.6 Sol and Fable 5; second on FrontierSWE and Kimi Code Bench, behind Fable 5; second on Terminal Bench, narrowly behind GPT-5.6 Sol; and first on Program Bench and SWE Marathon, narrowly beating GPT-5.6 Sol and Opus-4.8 respectively.

2030: The Consortium has placed you under arrest for the crime of improving MFU a Consortium agent shouts in your face. "You make me sick. Fused kernels! Overlapped comms! How do you sleep at night!?". he winds up to strike you, but another agent holds him back. "They're not worth it, man!"

aria@aurelium.me · 4w ago

i am skimming "Plan A" and am a huge fan of there being some kind of global shadow government which, among other things, psyops all MLEs into never pursuing efficiency improvements ever again

In the past, companies have trained bigger and better AIs using both compute scaling (bigger training runs) and software progress (advances in AI algorithms—new paradigms, better training recipes, better data, etc.). Now, the Consortium tries to steer things so that the majority of improvement comes from increasing training compute.

I asked both gpt-5.5-pro and fable5-max about the same bulk-storage-schema problem. 5.5pro had an efficiency oversight but is overall sensible fable5-max's was batshit, and when I asked to clarify it began the response with the densest claudism I have ever seen

Bild

the dog appears to have temporarily detached its teeth from this car, and is now eagerly looking at a semi truck coming down the road

i am apparently expected to politely forget that all of the people who are convinced that zhipu is working entirely by distilling claude traces were also convinced that deepseek v3/r1 simply must have been primarily that because 10m was a completely infeasible budget

bluesky is the only place safe from deltarune spoilers because everyone here but me is too 30-50 years old to care

so what is the overlap between "my usecase is not ML, math, physics, biology, or cybersecurity" and "I would pay exorbitant per-token rates for a better model" on an enterprise level, who wants to use a model with a silent active-sabotage feature and which is trained to never talk about security?

not about anything in particular but some people on here are a bit too gullible w/r/t new AI papers no, this novel training method didn't make a 3B model better at all tasks than Claude, that paper didn't find a 10x efficiency gain, that new VRAM-saver is slow or degrades performance, etc.

I guess they're trying not to push their luck with already-tepid non-SF municipalities but I wonder how long it'll be until cities start making infrastructure explicitly more legible to AVs via short-range comms, including IR legibility in standards for signage, etc.

theory: the unifying principle between cranks, naïve people, and grifters is "using literary analysis in lieu of actually knowing what you're talking about"

unironically a beautiful and effective way to keep grifters out. if you get filtered by an anime catgirl popping up for 2 seconds because it's too cringe and gay for you was your heart really in the systems programming?

Bild