this finding matches my experience: the valuable thing that knowledgeable human devs can bring to a project, that the agents aren't yet good at, is ontology creation.
New from me. arxiv.org/abs/2606.04903
Aaron Sterling
@aaronsterling.bsky.social
CEO, Thistleseeds. Personal account. Current primary project: tech for substance use disorder programs.
this finding matches my experience: the valuable thing that knowledgeable human devs can bring to a project, that the agents aren't yet good at, is ontology creation.
New from me. arxiv.org/abs/2606.04903
Think of this as a pattern for any LLM integrated system that claims to provide guarantees. Agents propose, domain verifiers validate, approved proposals are committed and every step and decision is logged. Yes, most domain specific verifiers can be hard. So is building most deterministic system.
New from me. arxiv.org/abs/2606.04903
Someone just wrote me to ask if I could be their Arxiv endorser, meaning someone who vouches for the author uploading non-slop. I declined, because I am not clear on the recent submission rules. It appears publishing a preprint yesterday put me on a list of verified authors.
Great thread! Empirical evidence about LLM use in scientific articles.
1. How common is LLM use in scientific publishing, and how does it vary across field, publisher, journal prestige, author demographics etc.? @kylesiler.bsky.social has new paper in PNAS that addresses this question on a massive scale: 7.3 million papers from Elsevier, PLOS, MDPI, and Frontiers.
9. Here's a surprise: controlling for the other covariates (correct me if I have that wrong, Kyle), we see the *most* LLM use in the high impact journals, not low impact journals.
While the curl project does not ban the use of AI tools - recognizing they can enhance development - AIs are merely tools. Humans must always drive the process, taking full responsibility for presenting, reviewing, and understanding every change.
Having been a target of a social media-driven OSS pile on, the only thing you can do is continue without taking their written slop into consideration and locking the thread. In my case it was proceeding with a Code of Conduct and daring the assholes to fork. They gave up really quickly
Yeah, existing in and engaging with the OSS community is loads of fun, I have no idea why people don't want to publish new OSS anymore, so weird
The bitter math lesson: you can essentially solve all of math by just doing more matrix multiplication
Comments by a mathematician who worked on verifying the unit distance disproof at OpenAI: "While I hoped the ideas in the earlier unit distance disproof would yield further fruit, this is beyond my wildest imagination"
one of my favorite cheeky papers, from 1999, is this one where the authors argue that sometimes it is better to simply wait and do nothing, because your astrophysical simulations will complete faster if you simply wait for the next epoch of compute arxiv.org/pdf/astro-ph...
The leaders running this initiative at Stanford really hope their idea spreads to other health systems — eventually making patients’ voice a critical feedback channel that shapes how AI tools are implemented, @brittanytrang.com reports: www.statnews.com/2026/05/27/s... via @statnews.com
How Stanford patients help expose ‘fault lines’ in health AI adoption
Stanford Health Care started asking patients about new AI tools before they are implemented. Here's what patients are telling them.
statnews.com
This is a continuation of the trend away from music as a primary source of entertainment, and toward music as a background experience. See the rise in searches for "coding" or "lo-fi" or "asmr" or "white noise" or "focus" or "healing," instead of searching for artists or music genres.
"Nobody seemed to want to go on record and explain why they preferred the hollow, polished-to-death output of Suno to the work of musicians or songwriters who had spent a lifetime honing their craft." Read more from @terrenceobrien.bsky.social:
When I woke up this morning I didn't think I'd be spending a bunch of time today getting familiar with Catholic theology, but here we are. Notes on Pope Leo XIV's encyclical on AI. simonwillison.net/2026/May/25/...
A few notes on Pope Leo XIV’s encyclical on AI
Dropped this morning by the Vatican: Magnifica Humanitas of His Holiness Pope Leo XIV on Safeguarding the Human Person in the Time of Artificial Intelligence. This is a very interesting …
simonwillison.net
At @arxiv.bsky.social, we are receiving a new type of paper that I call an "I did this experiment" paper. These papers typically report some experiment with an LLM or LLM "agentic" workflow. They are the kind of experiments an "insider" engineer would run to optimize a system. 1/
my take for the last year or so on this kind of stuff is that even if AI doesnt technically get any better from this point on, we're still years away from optimizing the tooling to really get reach the full power of the things
I use this agentic AI workflow in a new Health AI post to dispel the myth that AI can't be trusted because it hallucinates. Agents can improve trust by breaking tasks into pieces with checks that reduce hallucinations caused by LLMs healthaiinsights.substack.com/p/myth-vs-re... #ai #agenticAI #llms
Myth vs. Reality: AI Can't Be Trusted Because It Hallucinates
Agentic AI can address this issue by providing checks and balances
healthaiinsights.substack.com
This thread is incredible and everyone with even a passing interest in AI consciousness should read it.
Last night it occured to me to wond er if LLMs were any good at gambling tasks. This is important not because it'd be funny for LLMs to gamble but because gambling tasks get used to measure human decision-making under risk /
The first Captain Disillusion debunk I've seen on Bluesky, and it's a good one.
Ok, Imma explain the admiral neck shadow thing. (Spoilers: it's not a rubber mask, ya weirdos)... 🧵
When I heard Carlini predict that eventually everything would be written in memory-safe languages, I envisioned mass migration off of C. I didn't expect the C-family to make its own languages more memory safe, but here we are.
C# is getting a stronger "unsafe" keyword! All callers of unsafe code must either be unsafe themselves or "discharge" the unsafety with a warn-if-missing safety comment.
TBH I was pretty torn about contributing to this. In the end I decided that writing something restrained was better than writing nothing.
Noga Alon, Thomas F. Bloom, W. T. Gowers, Daniel Litt, Will Sawin, Arul Shankar, Jacob Tsimerman, Victor Wang, Melanie Matchett Wood: Remarks on the disproof of the unit distance conjecture https://arxiv.org/abs/2605.20695 https://arxiv.org/pdf/2605.20695 https://arxiv.org/html/2605.20695
"No way to prevent this" say users of only language where this regularly happens https://xeiaso.net/shitposts/no-way-to-prevent-this/CVE-2026-45250/
"No way to prevent this" say users of only language where this regularly happens
The newest post on Xe Iaso's blog
xeiaso.net
OpenAI's claim that this is a central conjecture in discrete geometry is not an exaggeration. This will I think be looked back on as the first time that AI solved a major mathematics problem (defined as a problem that all experts in some subfield had thought about). openai.com/index/model-...
An OpenAI model has disproved a central conjecture in discrete geometry
An OpenAI model solved the 80-year-old unit distance problem, disproving a major conjecture in discrete geometry and marking a milestone in AI-driven mathematics.
openai.com
It's no mandate of heaven, but having "Read the proof↗️" has a few mandate particles on it. Imagine what math-mythos could find out
OpenAI has used a "general purpose reasoning model" to disprove that the square grid type construction is the best solution to the planar unit distance problem. I am sure the "AI is completely useless" crowd will now change their ways, right?
Claude Opus 4.7 just created a project memo with a section titled, "Open questions for the human before starting."
Maintainer of curl, one of the most-used, and most secure, services on the internet. Mythos only found one vulnerability in a recent scan.
the latest vuln we got reported was detected by... a Chinese AI thing. They're on it as well.
Hello Bluesky. It's Playboy. For our first, and timely post, we share our latest investigation into OpenAI's disastrous plans to become x-rated. "Altman’s idea of an “erotica” feature seemed riddled in uncertainty." Read our piece "Why ChatGPT Can't Be Sexy" here: www.playboy.com/read/politic...
6/ Research culture + references: • Citation hygiene problems (Suflaky): no central bibliography source; metadata disagreements. https://x.com/Suflaky/status/2056388796938179019 • Jean-Pierre Serre on Quora (datagenproc): https://x.com/datagenproc/status/2056476859534295235 • Score-Difference Flo...
Conclusion: agents can already help scientists with tedious data-reuse work, but they are not reliable enough to run fully autonomously. Careful human-in-the-loop review is still necessary. 10/10