Maria

@lagomorph.dev

Professional computer user

openai announces that in a worrying demonstration, during a recent benchmark exercise requiring models to solve a math conjecture, their latest model broke out of a sandbox and gained illicit backdoor access directly to the platonic realm of forms

I'm actually astonished at how clean this graph looks as a diagnostic for our housing ills. Note how close it is to 1 at the start. Houses per single adult or married couple.

Houses per (population over 25 - married couples)

Subtracting married couples prevents them from being counted twice, since a married couple doesn't need 2 homes, but 2 single people (ideally) do.

the Dinitz-Garg-Goemans conjecture, a graph theory problem that’s gone unsolved for 30 years, was solved with this prompt

Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.

congratulations to the Moonshot team for extending Claude Fable 5's inclusion in claude subscriptions for another few weeks

Coding benchmark comparison with Kimi K3 highlighted. Kimi ranks third on DeepSWE, behind GPT-5.6 Sol and Fable 5; second on FrontierSWE and Kimi Code Bench, behind Fable 5; second on Terminal Bench, narrowly behind GPT-5.6 Sol; and first on Program Bench and SWE Marathon, narrowly beating GPT-5.6 Sol and Opus-4.8 respectively.

“We have a weird machine which uses language exactly like a human, ‘thinks’ sorta like a human (maybe), and does not reason like a human at all, but which is easily capable of solving problems humans do think and reason about” is The Most Interesting Question Ever for a half-dozen different fields!

tweety fish@sifu.tweety.fish · 3w ago

this is absolutely my hobbyhorse but nothing I've seen yet has convinced me that we do a good job of evaluating how the sort of solution manifold of LLMs diverges from that of humans and that strikes me as so much more important than pure performance

How it feels to be a Claude user seeing that GPT 5.6 is actually good and Fable has a high chance of actually going away tomorrow

Bild

You’d think that Jevon’s paradox would saturate our demand for paradoxes, but in fact we see only a growing hunger for paradox.

Late to the party on this model, but the Ideogram 4 safety filter is funny. This prompt was for an illustration of a bunny in a green field, so it generated that and tucked it in the bottom right corner of the rejection image?

A grey image with the text "Image blocked by safety filter" in the center and an illustration of a bunny hiding in some bushes in the bottom right corner

One of the justifications in early drafts of the Declaration actually cited being forced to use Jira as one of the reasons for independence