Looking forward to LLM skeptics publicly updating their views after the Erdos problems and the Jacobian counter-example any moment now.
Federico
@federicovaggi.bsky.social
F_vaggi on Twitter. Senior staff scientist at Google X, previously Amazon.
Well, the Jacobean Conjecture appears to have just been proven false by Fable.
I just fulfilled my annual pledge today by donating 21% of my income this year to clean water projects in DRC and Somalia. The biggest aid cuts took place over a year ago, but because of runways & emergency bridge funding, some of the worst effects are still to come. We can help reduce them.
We don't have to sit back and just watch the horror unfold
This month, I signed a pledge to donate at least 10% of my lifetime income to effective charities.
scientificdiscovery.dev
I wrote a quick political rant on what people mean by "democratic resistance to AI," which I sometimes strongly agree with and sometimes think is actually a depressingly anti-democratic way of steamrolling the debate blog.andymasley.com/p/what-does-...
What does it mean for AI to be democratic?
Some pushback on a specific fuzzy idea with some ominous implications
blog.andymasley.com
"If I am to flower into something other than myself, I would rather rot into nothingness as I am." - The Reflecting Pool
Breaking news: Indiana University plant microbiologist Roger Innes has been locked out of his laboratory by the school in response to a request by one of his federal funders. The move comes after Innes complained about the government’s prosecution of Chinese postdocs. https://scim.ag/4tsDqRr
After USDA request, Indiana plant biologist locked out of lab by school
Move comes after Roger Innes complained about the government’s prosecution of Chinese postdocs
scim.ag
FYI: I am going to stop summarizing Supreme Court decisions on here as they come down. One comment has been plucked out of context of all my reporting, misread, and used as the basis of a mean-spirited pile-on. I am not going to subject myself to this. If this was your goal, then congratulations.
pmarca reading Scanlon: "oh man, they got this all screwed up!"
Where to begin with this
Today, new guidelines have expanded recommendations to take cholesterol-reducing medicines early! If you missed it, I'd highly recommend our episode on it:
New episode of HARD DRUGS! Should everyone be taking statins? Statins have revolutionised heart disease and they're one of many reasons for the long-term decline in cardiovascular mortality.
Very excited to have finally found a reason to actually watch a TED talk!!!
This still sounds kind of crazy to me, but I'll be speaking at the TED Conference this year. They have written a very nice bio of me.
Aside from me personally, EAs co-organized the only real-life protest against the cuts, phone-banked to get Congress to revoke the cuts, and donated between seven and eight figures to help African clinics affected by the moves (for example, CTRL+F "GiveWell" in www.nytimes.com/2025/12/28/h... )
How Cameroon Fought to Save Its Malaria Program After the U.S. Cut Critical Funding
nytimes.com
There is a cool paper showing that modern LLMs “know” that p-hacking is wrong: however they can be coaxed with some creative prompting: andrewbenjaminhall.com/asher_et_al_...
I don't think I've ever seen gears shift this fast before.
Bluesky in a nutshell. Some context: BLS released its jobs numbers today, which were better than expected, and resistance libs would rather stay mad/conspiratorial than listen to actual on-the-ground experts that the data is still reliable.
New preprint! So, what's a multiverse analysis good for anyway?> With @jessicahullman.bsky.social and @statmodeling.bsky.social juliarohrer.com/wp-content/u...
An interview by Chotineer with the cinematographer behind Melania, it went exactly as you'd expect. www.newyorker.com/culture/q-an...
A video of Alex Pretti reading out the final salute of an unnamed veteran he cared for until the end of his life in the ICU, posted to Facebook by his son.
Amazingly balanced and down-to-earth discussion on the future of AGI. Very good, knowledgable, informative host with great questions driving it, referring back to relevant past comments. However you feel about the topic, this is worth watching. 1/4 www.youtube.com/watch?v=02YL...
FULL DISCUSSION: Google's Demis Hassabis, Anthropic's Dario Amodei Debate the World After AGI | AI1G
YouTube video by DRM News
youtube.com
This is a point I desperately wish more people (especially here) would grasp. There are very few low hanging fruits or pareto improvements when it comes to big policy questions, almost all policy decisions involve pissing at least some people off.
Governing means making trade offs and disappointing people. People tend to see the glass half empty rather than being grateful for what they get. So you get a lot of disappointed people.
I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.
The AI Researcher's Guide to a Non-Boring Bluesky Feed | Naomi Saphra
How to migrate to bsky without a boring feed.
nsaphra.net
This is a good write up, but, I think it's missing an important angle. The challenge is that people *want* to believe it's true, and that post was perfectly calibrated to appeal to people's priors, in spite of having a lot of obvious "tells" that had nothing to do with AI.
The author of a viral Reddit thread alleging fraud at a food delivery company tried to back up his claim by sending me AI-generated documents. Today I'm publishing those documents in the hopes that it helps other reporter see what we're up against in the age of AI www.platformer.news/fake-uber-ea...
Of course it would probably help if I pasted the right url: federicov.github.io/uncertainty-...
Semantic Entropy as a Regularizer for LLM Calibration
This post explores using semantic entropy as a training signal for calibrating confidence in language models. Here is what I found: Training on semantic entropy alone does not converge and leads to un...
federicov.github.io
Following @nsaphra.bsky.social - will try to do more science posting here as well. Here is a little project I worked on during the Christmas break. Using semantic entropy as a regularizer to calibrate LLM uncertainty: federicov.github.io/uncertainty-...
My first foray into systems biology -- this work was led by Boya Hou, a postdoctoral researcher at UIUC, now on the academic job market. She presented an early version of these results at the @nitmb.bsky.social MathBio Convergence Conference in August of 2025.
Spatially-Coupled Network RNA Velocities: A Control-Theoretic Perspective
RNA velocity is an important model that combines cellular spliced and unspliced RNA counts to infer dynamical properties of various regulatory functions. Despite its wide applicability and many varian...
arxiv.org
We're thrilled to welcome everyone to the first day of the inaugural NITMB MathBio Convergence Conference! Check out the schedule and book of abstracts to learn more about today's presentations and posters at www.nitmb.org/nitmb-mathbi...
Following @nsaphra.bsky.social - will try to do more science posting here as well. Here is a little project I worked on during the Christmas break. Using semantic entropy as a regularizer to calibrate LLM uncertainty: federicov.github.io/uncertainty-...
federicov.github
I find it hard to square the preciousness about "why would someone ever stay on Twitter?" when everyone on this site knows that the community here intentionally chased off lots of people.
What if you could train agents on a 𝗱𝗲𝗰𝗮𝗱𝗲 of driving experience in 𝘂𝗻𝗱𝗲𝗿 𝗮𝗻 𝗵𝗼𝘂𝗿, on a single GPU? Excited to share 𝙋𝙪𝙛𝙛𝙚𝙧𝘿𝙧𝙞𝙫𝙚 2.0: A fast, friendly driving simulator with RL training via PufferLib at 𝟯𝟬𝟬𝗞 𝘀𝘁𝗲𝗽𝘀/𝘀𝗲𝗰 🐡 + 🚗 youtu.be/LfQ324R-cbE?...
PufferDrive 2.0 release
YouTube video by Daphne Cornelisse
youtu.be
I appreciate the forthright acknowledgement of the mistake from the author, and look forward to decisive action by @princetonupress.bsky.social to correct this error, and any others, that made it past their robust editorial process.
Requested a correction from @princetonupress.bsky.social. Let's see what happens.
I actually think people are earnest when they believe that <claim> is false. The process of justification is something like: - <claim> is upheld by people I don't like, which is very strong evidence it's false. - <claim> undermines something I think is important or implies a trade off
Person A: <claim> Person B: Did you honestly say <claim>? You idiot. You moron. Person A: What is incorrect about <claim>? Person B: I can't tell you, for Secret Reasons. *** I don't understand why Person B thinks this is effective rhetoric.