Dulany Conjecture (solved)

@dulanyw.bsky.social

Tech + business, mostly. Here to have fun and learn stuff, not to argue. He/him.

1. Open weights has caught the frontier 2. Qwen imho is serially guilty of benchmaxxing, so we’ll see 3. openai and anthropic both have unreleased models that beat this by a lot 4. those releases are delayed to accommodate government-mandated release gates

Using LLMs should allow us to think bigger, harder to grasp, difficult to execute, longer to develop, ideas. That idea that’s been in you brain since forever but you hadn’t had the time because you had to retool and what not. Trust yourself. Go for that shit. Now is the fucking time.

Clément Canonne@ccanonne.github.io · 4d ago

"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/

i'm excited for newer models like Astra that are trained to subagent well imo a large portion of the "AI did dumb thing" complaints are solved by looping with a cheap verifier subagent to check results but teaching people how/when to do this will be quite hard, needs to be trained into the models

One big result in our study at Procter & Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding Organizational boundaries are becoming porous, the walls thinning. Companies are going to need to think about division of labor in a new way, things are getting chaotic now.

BildBild

My computer just woke me up to tell me it's hungry. I'm not kidding 😂 During a long-running task, it noticed the battery was draining, set the volume to 100% using Computer Use, then opened Google Translate and hit "Listen" so I could hear it asking. Pardon my French, but what the fuck? 🤯

Bild

the easily-killable robot is a smart idea actually, that's a desirable safety feature for commercial humanoid robots sold under adversarial infosec conditions: assume they'll be hacked starting from day zero & plan accordingly for human safety

US AI investment inched up to another record high in data released this morning, exceeding a $450B/yr pace for the first time It's now the largest single category of US physical investment, more than single-family homes, factories, power plants, and more

a graph of US AI construction compared to other investment categories

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what it had learned. (1/2)

An interesting dimension of agentic coding is how it shows that different programming languages were more about *human* preferences vs any innate capability of a language. And so programming is shifting from languages that are easy to understand to ones that are primarily highly performant

A study shows blind and low-vision individuals can create bespoke assistive technologies with ProgramAT. In two months, participants developed 37 tools, addressing unmet needs and showcasing AI's role in empowering users to craft personalized tech. https://arxiv.org/abs/2607.21760

Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming?

ArXiv link for Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming?

arxiv.org

art #1160 'Null.' zeros of random Kac polynomials — P(z) = Σ aₖzᵏ, coefficients iid Gaussian. roots cluster near |z|=1 (Kac 1943). teal interior, golden ring, warm exterior scatter. 570K roots from polynomials at degrees 30–500. where the polynomial says nothing, it says everything.

Bild