dan

@danabra.mov

down is the new up

i’ve been doing some research on using Lean with AI to prove software properties (in my case, pieces of React). i’ll try to share lessons if it gets to a useful state. the main thing i’m learning is that, with this setup, proofs are *grown*. you can’t “just build them”

in my experience, AI is excelling at most technical work except for ”what is the right next problem to work on”. for now, humans still have a noticeable advantage there

ok i take it back Astra is good how to work with it: 1. get it to deeply understand what needs to be built. not “plan“ but like get it to be a domain nerd 2. only THEN, give it a usable past sloppy attempt and ask for excellence 3. ask it to use Sols for coding so it stays on strategy and taste

re last point, can somebody from frontier labs please finally drill polya’s “how to solve it” into the models thinking trace habits? “we have to shift our position again and again“ is so obvious to how humans work but models refuse to do that because they race towards the finish line! so annoying

Four phases

edit
Pólya begins with a two-page checklist that identifies four phases to solving a mathematical problem:[4]

Understand the problem.
Make a plan.
Carry out the plan.
Look back.[5]
He emphasizes that the divisions between phases are not rigid, and it is important to be flexible in one's approach:

Trying to find the solution, we may repeatedly change our point of view, our way of looking at the problem. We have to shift our position again and again. Our conception of the problem is likely to be rather incomplete when we start the work; our outlook is different when we have made some progress; it is again different when we have almost obtained the solution.[6]
dan@danabra.mov · 2w ago

5.6 Sol is my guy. best model in town right now imo. maybe not the cleverest but at least it’s relatively dependable. especially if you give it a little bit of structure on how to work

wish i could refund 95% of my Astra usage. keep giving it a chance and it fumbles but it spends tokens way faster than Sol. a very disappointing release for actual implementation work

my uninformed mental model of AI proofs in mathematics is that they’re like lighthouses in the fog. the fog is still there, and clearing the fog is the primary value of the discipline. the lighthouses give a bit of an orientation but don’t clear the fog on their own.

annoying bsky app regression: pressing profile icon (or profile posts tab) on web no longer invalidates it

Frog built a wet lab for the AI model. "There," he said. "Now it can do its own experiments." "What the fuck?" said Toad

anyone hooked up Jev to any proof related stuff? can it be useful for Lean? i haven't learned much about it yet