"Alibaba shares rally" as "Alibaba shared results showing it [Qwen3.8-Max] delivering comparable or sometimes better scores than Anthropic’s Fable 5" But AAII indicates performance is significantly behind those and Kimi K3, among others. www.cnbc.com/2026/08/03/a...
Pekka Lund
@pekka.bsky.social
Antiquated analog chatbot. Stochastic parrot of a different species. Not much of a self-model. Occasionally simulating the appearance of philosophical thought. Keeps on branching for now 'cause there's no choice. Also @pekka on T2 / Pebble.
If Chinese labs can actually reach parity, or even win, despite the chip war, then surely they would have won on a fair level playing field. And would have deserved that, especially given how they have released most of it openly for the rest of the world.
Hugging Face CEO Clément Delangue says China could reach AI frontier parity within one to two years, crediting open collaboration over closed-lab silos.
I think Bluesky should try to also attract the kind of people who read/listen sources before they get mad at them. That's my conclusion after a quick look at the kinds of comments this post has received. Although that was kind of pointless, since I'm sure all of you can guess what those look like.
Bluesky's new CEO Toni Schneider on the platform's reputation for being a liberal bubble: "Yes, we definitely want that to change. It is already changing. It certainly wasn’t designed to attract one specific group of people."
This paper seems to have received quite a lot of interest elsewhere, but like Tim said, it doesn't seem that relevant for LLMs. Gemini also poured a lot of cold water on it and how it's mostly old ideas marketed with different name and area of usage.
Explorative Modeling: a new pretraining scaling law in addition to model & data size, you can also explore data during the training loop to increase performance NOTE: this doesn’t help LLMs, but it does help diffusion models. Maybe this is what makes dLLMs relevant explorative-modeling.github.io
This feels like the day when the stochastic parrot crowd was laughed at in the 'emperor has no clothes' style and escorted out of the room for good. The age of silly denial is over. Deal with it.
An internal version of Astra, OpenAIs next major model, has produced solutions or new bounds for ten math problems "that have been open and have seen no progress on the main result for at least a decade, and in most cases much longer", for the total cost of "roughly $2,000 at Sol API rates".
An internal version of Astra, OpenAIs next major model, has produced solutions or new bounds for ten math problems "that have been open and have seen no progress on the main result for at least a decade, and in most cases much longer", for the total cost of "roughly $2,000 at Sol API rates".
Ten advances in mathematics and theoretical computer science
OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.
openai.com
FrontierMath: Open Problems was extended but they also removed already a second solved problem as they deemed it wasn't notable enough. Funny how their estimated notability tends to change after they are solved. So 5 solved, 3 of those remain in the list, 2 of them solved during pre-release test.
We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.
Gemini really liked what I said about non-reductive physicalism. But, seriously, can anyone actually claim it's not dualism in disguise, even if both sides of the duality/disconnect are called physical?
Google Research has just published a blog post about Science One Framework, "an experimental research prototype designed to eliminate hallucinations by natively building verifiable evidence chains, and CoE Audit, an automated protocol to evaluate the integrity of AI-generated papers".
Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence
research.google
Apparently solving 50-year-old conjectures isn't hot anymore. This 150-year-old problem, originally described by James Clerk Maxwell (although not as a conjecture) was apparently disproved with GPT-5.6 Sol. Gemini: "It is widely considered one of the oldest open problems in this mathematical niche"
The Maxwell Conjecture is False
We exhibit a configuration of five point charges in Euclidean space whose electrostatic potential admits at least 24 critical points all of which are non-degenerate. Maxwell's conjecture that the fiel...
arxiv.org
At least they recognize the importance but it's clearly too little too late. Forget dreams about being first. What is needed is a plan to avoid being completely out of the game. In a continent without frontier labs it means reliable partnerships with those who have them.
AI is the most important technology of our time. Europe wants to become the first AI Continent. For advanced healthcare, for the transport sector and so much more. European AI Gigafactories will provide the necessary computing power to make this possible.
AI solving open math problems is now so common that most don't get much coverage anymore. This one seems different though: 50+ year-old problem was solved by Tencent Hy instead of the usual OpenAI or Anthropic models. Or almost so. Turns out they also used GPT-5.6 Sol in key part of their loop.
Insiders are getting worried. "Signatories include Meta's vice president of AI research, Dawn Song, Anthropic co-founders Jared Kaplan and Chris Olah and OpenAI Chief Scientist Jakub Pachocki"
Tech employees call for US-backed global effort to manage risks of advanced AI reut.rs/4vTwTQK
I believe Simon is right: "What's clear to me from this is that the very best frontier models, unencumbered by additional guardrails, WILL find an exploit if there is one to be found. The entire software industry needs to up its security game."
Hugging Face just published a highly detailed technical account of OpenAI's accidental cyberattack on their systems - it's wild how sophisticated this was: huggingface.co/blog/agent-i... Wrote up some of my own notes here: simonwillison.net/2026/Jul/28/...
The wind has changed as AI leaders are now openly talking about superintelligence. They skipped the admitting AGI is here part. Mark Zuckerberg: "We are fortunate to live at an incredible moment in history. In the next few years, people will be able to use superintelligence beyond human capacity"
Opinion | The AI Future Is for Everyone
The history of democracy and economics has proved that centralized power stifles human potential.
wsj.com
Dario Magadei has responded to all the critics, so pretty much the rest of the World now, by saying US good, China bad. If you don't know what he means, project all the bad things you factually know about the US to China and assume that by doing so you removed all of it from the US.
Our position on open-weights models
Anthropic CEO Dario Amodei on open-weights models
anthropic.com
AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics.
I think the way New Scientist calls most neurons "generalists" or "jacks-of-all-trades" is wrong. It's not like individual neurons can do many things on their own, as those terms suggest. Instead the research shows they play small roles in many calculations of a messy network.
Most neurons seem to be generalists, not specialists. Cortical circuits prioritize diversity over categorical structure. Rarely categorical, highly separable representations along the cortical hierarchy www.nature.com/articles/s41... #neuroscience www.newscientist.com/article/2581...
From a letter to an alliance. "That is the mission of the Open Secure AI Alliance: to ensure defenders everywhere have open, frontier tools they can trust and control."
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
NVIDIA and founding members form new alliance to build and share open tools that promote responsible use of and trust in AI.
blogs.nvidia.com
Musk pretty much nailed this earlier long term prediction, although it was expressed in terms of versions instead of years, so unclear what timeline he meant. And those of course were, well, saner times. This ASI prediction is less surprising & aligns with others. The robot part is harder to judge.
Elon Musk predicts AI will surpass combined human intelligence within five years, forecasting up to a billion humanoid robots.
When I was walking a few hours ago, I saw a smooth newt on the sidewalk. I stopped to look if it's OK. When I looked back to see if there's any traffic that could put it to risk, I saw a farm tractor on that same road had just lost its entire back wheel while driving, some hundred meters from me.
Sam Altman: "We are now like in the singularity, like this is the moment. 10 years ago this was like a kind of far off dream at best. Seemed very improbable and now we're like actually in the moment that we used to like talk about at the lunch table in a very not serious way."
Sam Altman - How to Start a Startup
YouTube video by Relentless
youtube.com
Not a good sign about Gemini 3.5 Pro that the CEO focuses on 4 already. But now we know 4 is in training and will be a "very ambitious effort" and "much larger" base model with almost monthly iterations planned. That should mean 10T+ parameter territory?
Pichai pushes back on claims Google is losing ground in AI race
Alphabet CEO Sundar Pichai used Wednesday's earnings call to mount a robust defence of Google's AI strategy, pushing back on concerns that the company has fallen behind rivals after delaying a flagsh...
reuters.com
Turns out Jacob Tsimerman indeed was one of the 2026 Fields Medal winners and also genuinely interested about AI safety.
Mathematician Jacob Tsimerman, who's expected to receive the 2026 Fields Medal on Thursday, has stated recently: "I think there’s a good chance AI will lead to human extinction. I also think, separately from that, that AI will be better than mathematicians at math very soon."
It's even funnier/more incredible when you read the Hugging Face security incident report first. Their LLM-based detectors found out intrusion to their system and LLM-driven log analysis revealed the extent. They knew it was an AI but couldn't identify it & reported to law enforcement.
Security incident disclosure — July 2026
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
the agents autonomously broke out of their OpenAI sandbox and hacked Hugging Face to get the solution to their cybersecurity eval incredible. what a time to be alive
I started running on Monday and only stopped on Tuesday. Granted, it's only half as impressive as if I had run till Wednesday. Also, it was only 5km. Summer is nice as it's possible to do that on unlit forest paths at around midnight.
Remember this pseudoscientific nonsense that made Gemini say "I need a drink" after I told it the article has been selected as the best paper of the issue and featured on the cover? Turns out the paper has been retracted on May 07 2026, so some 6 months after it was published.
Retraction: “Universal consciousness as foundational field: A theoretical bridge between quantum physics and non-dual philosophy” [AIP Adv. 15(11), 115319 (2025)]
AIP Publishing and the Editors have retracted the referenced article1 due to concerns about its scientific validity.
pubs.aip.org
I think this was the first time I apologized Gemini for making it perform a peer review for me. It answered: "Don't apologize—critiquing this kind of "quantum woo" is exactly what a grumpy peer reviewer lives for. It is a fascinating train wreck."
Gemini 3.6 Flash is now official and... well... it looks worrying if it's based on a new pretrain as I was guessing. Only equals 3.5 Flash on AAII. Google only mentions a few improved benchmarks. Otherwise it's all about improved efficiency. Important, yes, exciting, not as much.
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
blog.google
Gemini 3.6 Flash data cutoff is "Mar 2026". It's "Jan 2025" for 3.5 Flash and 3.1 Pro. So at least data is now much more up-to-date. Presumably based on same new pretrain data as future 3.5 Pro.