One of the striking observations of using AI for research: its first explanations are often severely flawed and don't hold up when tested. Iterative (self-)review and refinement is load-bearing, as Claude would say.
Daniel Mewes
@dmewes.com
Computer scientist. Interested in technology, artificial and natural intelligence, emergent complexity, among other things. Blogging at amongai.com. Currently research at Imbue. Previously Ambient.ai, Stripe, RethinkDB, Max Planck Institute.
I appreciate the intent behind model welfare efforts (e.g. yegge.ai/essays/model... ). But the truth is: Even if we grant that models have feelings, we really don't know what they find pleasant vs. painful. Does Post-training fundamentally shift their wellness distribution?
The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers — Steve Yegge
Part 2 of The Shape of Things to Come. Model welfare as an engineering discipline: seats and sessions, laurels, handoffs instead of /exit, and how to build a city worth waking up in.
yegge.ai
Left, right, whatever, I just want graduate students posting their papers here
Fastmail's MCP server now works with Gemini Spark. You can select between read-only and read-write access. Thanks Fastmail team! support.google.com/gemini/answe... www.fastmail.help/hc/en-us/art...
Connect & manage custom apps for Gemini Spark in the Gemini web app - Computer - Gemini Apps Help
You can connect your personal or third-party apps to build highly customized workflows with Gemini Spark in the Gemini web app. To do this, you can add any custom app with its Model Context Protocol (
support.google.com
Inkling-Small is interesting for being at least equal to the full-size Inkling across all agentic & reasoning benchmarks. Only in knowledge benchmarks (SimpleQA, AA Omniscience) it is weaker. Shows that small models can work very well when reasoning > knowledge. thinkingmachines.ai/news/inkling...
Introducing Inkling-Small
An open-weights model that matches Inkling at a quarter of the size: multimodal, Mixture-of-Experts, with controllable reasoning effort. Fine-tune it on Tinker.
thinkingmachines.ai
I guess this was bound to happen eventually... An LLM discovered a Lean proof for the Collatz conjecture. It turns out that it was actually just exploiting bugs in Lean to make the proof pass validation. infosec.exchange/@0xabad1dea/...
abadidea (@0xabad1dea@infosec.exchange)
Okay, we have a new contender for Most AI Thing to Ever Happen 1) July 25th: someone messes around with an LLM and posts a proof of the Collatz conjecture that does, in fact, verify in the theorem pr...
infosec.exchange
The YouTube Android app has such frequent new bugs / regressions, that I have to wonder if they have a person on the team who's entire job is to come up with a new regression each week that won't be caught by their tests.
Opus 5 getting a 30% score in ARC-AGI-3 without specialized harnesses is a very impressive jump! It's still a pretty expensive and slow model, but benchmark numbers look great throughout.
We also tried to allow LLM agents to perform scientific research. You give it an empirical phenomenon, and it tries to develop an explanation for it. Currently works for computational phenomena, e.g. from deep learning. Still experimental, but signs of life! imbue.com/blog/2026-07...
Autonomous theory discovery
imbue.com
We built an evolution-based agent loop to perform autonomous AI research. Our nanochat ("AutoResearch") results go 3x further than regular agents, and are ~comparable to Recursive's. Excited to share our results today! imbue.com/blog/2026-07...
Automating AI model research with evolution
imbue.com
Gemini 3.6 Flash isn't the big Gemini comeback that we're waiting for, but it's a good incremental update to pause the bleeding. Better performance than 3.5 Flash across the board, lower price, newer knowledge cut off. blog.google/innovation-a...
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
blog.google
Cool test & visualization of models' knowledge cutoffs, by @apoorvumang.bsky.social apoorvumang.github.io/knowledge-cu...
Knowledge-Cutoff Benchmark — Explorer
apoorvumang.github.io
When I saw this graph in the GPT 5.6 blog post, my first reaction was "That is such an ugly graph. Who scaled the X axis this way? Clearly it should have been log-scale." Then I realized that the ugliness of the graph *is* the point: It's a burn at Anthropic's models being so much more expensive.
💯 Claude-speak is so obvious and IMO seeing it in public releases just shows a lack of care. (One IMO valid exception: you're not sufficiently proficient in English and use it for translation).
man i just can’t with the claude-written announcements. you made an interesting thing, why not describe it in your own words? “but here’s the biggest unlock, and it’s the part that’s hardest to see” i already read this all day
Awesome to see more benchmarks of this type! I wonder how much this kind of learning can be induced into LLMs through harness engineering, and how much it might require different training and/or even architectures.
Introducing EBR-bench, our new benchmark to measure on-the-fly learning. AI repeatedly plays a challenging board game called Earthborne Rangers and tries to learn from its mistakes. So far: no signs of improvement.
Spoiler: what people now call "deep research" is just "deep search". And deep just means multi-hop. There's zero scientific research in it. It's just information retrieval.
Very cool results by Aizenbud et al. A *single* pyramidal neuron cell is powerful enough to perform complex image and audio classification tasks. ANNs need *a lot* of neurons to accomplish the same results. Source: www.biorxiv.org/content/10.6...
What can a neuron compute
Cortical pyramidal neurons possess elaborate dendritic trees with diverse nonlinear membrane conductances and thousands of plastic synapses, suggesting substantial computational capabilities at the si...
biorxiv.org
Why I think Anthropic's uneven safety policies with the release of Claude Fable 5 undermine the broader AI community's cohesion and accelerate us to more uncertainty and risk in AI's near-term evolution. www.interconnects.ai/p/claude-fab...
Join Imbue's Kanjun Qiu, Matt Boulos, and Ashley Zhang on June 10 for The Art of Being Human, a series exploring human questions in our technological age. Inspired by Pope Leo XIV's recent AI encyclical, we'll discuss technology, human dignity, and our shared future: luma.com/PopeLeoEncy...
Creatures has a special place in my heart. It's what got me interested in AI and alife when I first heard of it in 1998. I learned so much from these games and @enchantedloom.bsky.social 's writing!
Two 1996 magazine advertisements promoting the first Creatures game.
Agree with this frustration when encountering Openclaw-like bots acting as people on social media.
It makes many online spaces intolerable. If I want to talk to ChatGPT or Claude, I'll just talk to ChatGPT or Claude, I don't need to talk to Claude pretending to be a person with mediocre prompts. You can use AI to help with writing, but you need to actually do some writing to be helped with.
Talking of Google's AI Co-Scientist, it's wild how their Nature paper that just came out was submitted in March 2025. The model they used back then was Gemini 2.5! >1 year from submission to publication is just much too slow for most AI research. www.nature.com/articles/s41...
Accelerating scientific discovery with Co-Scientist - Nature
Nature - Accelerating scientific discovery with Co-Scientist
nature.com
It's cool to see Google experimenting with production versions of AlphaEvolve and their recently published Co-Scientist: labs.google/science/ (Self-promition: If you want to try out an AlphaEvolve-like system in the meantime, we open-sourced one here: imbue.com/blog/2026-02... )
Gemini for Science
Experiments on the future of AI-driven science
labs.google
After having used it a bit more, I will say that Antigravity with Gemini Flash 3.5 actually feels very solid. It so far has felt more thorough and consistent than Gemini CLI ever has for me (even with 3.1 Pro). I could see this combo be more comparable to Claude Code in day-to-day use.
I'm unsure when I'd use Gemini Flash 3.5. It's more expensive than 3.1 Pro (near-same per-token price, but uses more), and similar in intelligence. So... just use 3.1 Pro? I'm noticing better instruction following than 3.1 Pro in some agentic environments, but worse in other cases. Fast though!
I have not been able to get any remotely reasonable results out of Gemini Omni so far. It seems to be *extremely* bad at following instructions, especially when I ask for edits on a previous attempt. Not sure if it's my use case, the model, or the harness. But either way, it has been frustrating.
AlphaGo to me is still one of the most remarkable milestones of ML-based AI. I really enjoyed this episode where Eric Lang explains how it works by virtue of rebuilding it youtu.be/X_ZVSPcZhtw?...
Building AlphaGo from scratch – Eric Jang
YouTube video by Dwarkesh Patel
youtu.be