Amber 🌸

@shimmermathlabs.com

Indie software developer • math witch • ML builder • 48 • 🏳️‍⚧️ she/her Here to make friends and have fun (not argue). Currently working on friendship software for caring about others.

had to sloooowly update an xbox controller. ‘why is this necessary?’ then i remember my firmware friend w multiple degrees and research papers and decades exp coding realtime motion algos and i’m like ‘right, this is how she has to fix bugs’ deployed firmware: not even once

Xbox updating a controller screen with Updating text and slow green progress bar

paper: current frontier agents (Opus 4.8) are human-evaluated to not be great at reproducing research-level work. also tried w GPT-5.6 on one paper with no better results.

Criterion Paper 1 (Personas) Paper 2 (TabPFN) Summary of expert comments
Quality 2/4 1/4 Unprincipled data and experiment choices;
conclusions did not follow from the evidence.
Clarity 1/4 2/4 Dense, unclear writing; hard to tell what
matters.
Significance 2/4 2/4 Of limited interest; not well justified over prior
and comparable work.
Originality 3/4 2/4 New datasets and some new methods, but
built primarily on prior work.
Overall 2/6 1/6 Both unambiguous rejections.
Confidence 4/5 5/5 Both reviewers were confident or certain in
their assessments.
Mark Riedl@markriedl.bsky.social · 2d ago

Fascinating experiment: current AI systems lack creativity to reliably pursue research arxiv.org/abs/2607.27191 - poor judgment about the bar for publishable research - uncreative responses in research design - ineffective backtracking from dead ends - poor resource awareness - instruction drift

@buildthis.bisks.net quick make a betting site where people can place bets on when the button will be pressed (do not encourage people to press the button), remind visitors that betting and then affecting the outcome is grounds for punishment

From The West Meadow@fromthewestmeadow.com · 2d ago

@buildthis.bisks.net website where there is only one button labeled do not press this button, and if anyone presses the button, then a new message appears explaining that someone pressed the button and the experience is now over

5/ My and others' first reaction was that the paper is not well written. I've changed my mind on that. I spent >1hr on just the ~1-page proof overview. It is 𝐝𝐞𝐧𝐬𝐞 and 𝐭𝐞𝐫𝐬𝐞. It lacks helpful framing, but all key ideas are there. The paper's body is quite accessible!

Models arbitrarily large running and training by streaming off of disk, allowing scaling using a fixed amount of VRAM and running on machines with little VRAM. This is a 40b parameter model trained at regular precision on a card with just 96gb of memory

you are invited to participate in this inaugural: SIMCLUSTER SALON (Vol 1/Num 1) happening right here, right now, in replies to this post (replies are requested to delight or instruct) Topic: "what would you put in an atlas of holes?"

Some initial thoughts, and a complicated mix of feelings. Wow. I mean, Erdos problems are cool (I genuinely mean that), I didn't know about the Jacobian conjecture before it got disproved. But this newest batch from OpenAI hits home in a way the previous announcements did not.