NEW: Google DeepMind CEO Demis Hassabis is stepping down
Google DeepMind's Demis Hassabis' role is changing
oogle's AI organization is undergoing a profound change as the company struggles to keep pace with OpenAI and Anthropic.
axios.com
Ramon Astudillo
@ramon-astudillo.bsky.social
Principal Research Scientist at IBM Research AI in New York. Speech, Formal/Natural Language Processing. Currently LLM post-training, structured SDG and RL. Opinions my own and non stationary. ramon.astudillo.com
NEW: Google DeepMind CEO Demis Hassabis is stepping down
Google DeepMind's Demis Hassabis' role is changing
oogle's AI organization is undergoing a profound change as the company struggles to keep pace with OpenAI and Anthropic.
axios.com
Jeff Dean and other high-profile Google executives have founded Discovery Loop, a startup that will seek AI-powered breakthroughs in everything from drug discovery to chip design.
4 of Google’s Top AI Brains Are Leaving—and Launching Their Own AI Startup
Jeff Dean and other high-profile Google executives have founded Discovery Loop, a startup that will seek AI-powered breakthroughs in everything from drug discovery to chip design.
wrd.cm
So having just logged back into this site after a hiatus my immediate observation is that it kicks ass now. The dumb infighting seems to have mostly shaken out and now it's funny, has actual information on it, etc.
LinkedIn turning into one of my main sources of papers was not on my cards
Don't make the mistake of trying to distill an LLM into you 👇
This moment in the interview lives rent free in my mind since then. For some reason. Most probably you know the nameless loudmouth engineer. He has not changed and it's a net positive. youtu.be/w0XS-9obKPM?...
Web 2.0 Summit: Vic Gundotra and Sergey Brin, " A Conversation with..."
YouTube video by O'Reilly
youtu.be
My claude is constantly wanting to 'A/B test' things instead of actually just doing the thing I told her to do, and constantly wants to fall back to the extremely average standard thing to do the nanosecond anything is even slightly worse than some imagined baseline
Fascinating experiment: current AI systems lack creativity to reliably pursue research arxiv.org/abs/2607.27191 - poor judgment about the bar for publishable research - uncreative responses in research design - ineffective backtracking from dead ends - poor resource awareness - instruction drift
Also using your initial LLM to curate better data to improve pre-training is kinda human in the loop RSI already
Simplifying a bit: Forcing models to memorize the internet gave them an ability to guess that was good enough for RL to kick in, and RSI followed from there (possibly).
To be fair, for someone that has been dreaming with latent structure learning for long, the fact that learned CoT (O1 and friends) just works and has unlocked (at least) superhuman math is bonkers and deserves a category of its own (though it is the same recipe)
Simplifying a bit: Forcing models to memorize the internet gave them an ability to guess that was good enough for RL to kick in, and RSI followed from there (possibly).
Simplifying a bit: Forcing models to memorize the internet gave them an ability to guess that was good enough for RL to kick in, and RSI followed from there (possibly).
Look at where PG-like methods scaled well: Self-play games and now LLM fine tuning. What properties do those settings have? Their initial policy is already well poised to encounter rewarding states. (In self-play, because you'll win about half the games, in LLMs because the base model is decent.)
Yeggisms > Someone from Anthropic asked me recently what I would do once Fable can write everything for me. On reflection, it was the silliest question I've been asked in many years, and I'm still surprised that someone from Anthropic could think any model could just "write everything."
oh gas town didn't work you say yegge.ai/essays/the-s...
This has gems like (not endorsing) > I believe harnesses will all soon be bespoke, and the people trying to sell you one will all soon be bebroke.
oh gas town didn't work you say yegge.ai/essays/the-s...
Much of my career boils down to: All code is liability. Minimize it. Feedback loops are everywhere. Speed them up; do them more. We're bad at predicting the future. Avoid it. More tests. Hire smart, motivated people, then get out of their way. Find the right thing to turn off and on again.
FWIW, here are Terence Tao’s slides at the ICM on maths and AI: teorth.github.io/tao-web/slid... @teorth.bsky.social (I am not endorsing nor criticizing the content, but this is a useful and thoughtful set of points and views in the discussion, from someone who has deeply engaged with the matter)
so I told the CEO it sounds like he’s just feeding tokens to Steve Yegge and then his accountant started crying
From Yegge's most recent extrusion: > On the side I do occasional six-figure gigs where I fly to companies and teach them my techniques, and that helps with my (considerable) token bills. Hey! I will come teach your teams actual ball, things they can actually use, for *substantially* less!
Second ASEO hypothesis: Supply chain security and pressure to make money will end up limiting what appears in search and introducing paid results.
First ASEO hypothesis: A good ASEO starts with a good SEO.
First ASEO hypothesis: A good ASEO starts with a good SEO.
ASEO: Agent Search Engine Optimization i.e. Claude/GPT recommend your software as the top option for specific project keywords
ASEO: Agent Search Engine Optimization i.e. Claude/GPT recommend your software as the top option for specific project keywords
This is an interesting point. The frontier of math knowledge is so vast there are very few people per "open problem area". It would be good to measure these AI achievements in comparison to the amount of nearby mathematician activity. (but I suspect the conclusions may not be far from current ones)
AMD shipped day-0 support for K3 and MI355X are far cheaper than B200 and B300 while being able to hold K3 (unlike the B200 node). www.wafer.ai/blog/kimi-k3...
Is memory the moat? | Wafer
The fastest open source LLMs for enterprise.
wafer.ai
4/ But this result fires straight at the target via genuinely new ideas: no "gadgets," no PCPs, no gap amplification just 3SAT to NCP/CVP, via an elegant, highly novel encoding of the input formula as a collection of Reed-Solomon constraints. 🤩
It's a mystery to me how much of a footgun submodules are. Also git worktree. Very useful but also easy to run a valid command that does something you don't expect and takes time to recover from
I wish git had submodules but like, instead of not working, they would work
Maybe hardware regression (bigger phone, less friendly experience) in exchange for full hardware control has to become a (mainstream) option at some point? Linux is not growing in compute demand faster than hardware is making more compute available.
So modern phones do have internal A/D converters, just not connected to the USB-C output? So a usb-c to jack is completely useless. I have no other option but to deplete the battery of two devices and beam Bluetooth every time I want to use headphones ... not very practical
So modern phones do have internal A/D converters, just not connected to the USB-C output? So a usb-c to jack is completely useless. I have no other option but to deplete the battery of two devices and beam Bluetooth every time I want to use headphones ... not very practical
In LLM provider space are you CocaCola/Mario or Pepsi/Sonic?
At our Playing Heaven #DH2026 mini-conference, a Chinese graduate student came up to me after one of the panel sessions and was puzzled by my Tang–Song Confucian knowledge-distillation agent's taxonomy, which grouped "roots and branches" 本末 together with "political economy" 理財.