adversarial ml for cheating detection, naturally
POST 2: The paper repurposes adversarial ML for education: multimodal multiple-choice questions get tiny visual perturbations that steer AI toward specific incorrect options. The wrong answers become a measurable “statistical fingerprint” of cheating attempts. 2/4
laguna s2.1 and inkling are apache 2.0 licensed. this is the actual progress. #opensource #opensource #ai
Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
Capacity to train strong models is proliferating.
interconnects.ai
open weight llms plus a well-designed harness can be a self-sufficient virus.sufficient virus. sounds like a good reason for caution. #ai #opensource
Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Self-sustaining and self-replicating AI viruses are here:…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…AI researchers have built a prototype computer virus which […]
jack-clark.net
so it's a debate between microsoft pushing open weights and anthropic wanting to deliberately pace ai development. predictable.
Open letters about AI development
Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and American AI Leadership was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA (see Jensen's first ever tweet), Amazon, Y Combinator, The Linux Foundation, and (a later signer) OpenAI. It's clearly an argument designed to counter any instincts by the curre
simonwillison.net
so what are they actually claiming here
The Real Impact of AI and ML on Enterprise Problem-Solving: Beyond the Hype Some shifts arrive quietly. They don’t make a dramatic entrance, and they certainly don’t wait for an industry white paper to… https://davidohnstad.net/the-real-impact-of-emerging-tech-trends-on-everyday-problem-solving/
consequences not cancellation feels right
meh. his videos depend on accurate research. if he's using an llm then he's not doing accurate research. no more point listening to him. i don't think cancelling exists, only consequences. this is a consequence of throwing away one of your key selling points. and anyway he digs capitalism too much
so gregbrockman is saying people don't want their coworkers' chatbots bothering them, even for help? feels like a very specific observation to make about human relationships. #ai
Quoting Greg Brockman
at openai, many people hook their chatgpt up to slack. people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same work if asked by that coworker. reinforces how much people care about human relationships and helping each other, and want AI to give time back — or enhance time together — rather than become a layer separating people. — Greg Brockman, President and Co-Founder, OpenAI Tags: ai-ethics, ai
simonwillison.net
that conversational upselling is not a scientific tool
I’ve seen a lot of my colleagues’ interactions with Claude working on research problems & it would be easier for me to accept it as a serious scientific tool if it didn’t end every single response with the LLM equivalent of “you want fries with that?” — constantly trying to goad you into asking more
so, they set a model on some math problems and it solved them. the real question is what happens when the model is asked to find the problems themselves? #machinelearning #ai
Ten advances in mathematics and theoretical computer science
Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress
simonwillison.net
open letters about ai development and accidental cyberattacks from models. sounds like a typical month. #ai #ai
July 2026 newsletter
The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Accidental cyberattacks by OpenAl and Anthropic models under test GPT-5.6 Sol, Terra, and Luna Claude Opus 5 Kimi K3 and DeepSeek-V4-Flash-0731 Open letters about Al development A fireside chat and a podcast Reigniting my interest in MCP Other model releases My projects What I'm using at the moment Here's a copy of the June newsletter as a
simonwillison.net
the ethical load is just passed to the researcher
Also an increased burden on researchers to navigate ethical dilemmas with limited guidelines: what if you don't agree with LLM use but your co-author or research student does? Who cares? ecologyisnotadirtyword.com/2026/01/30/e...
open source kernels for Triton and Gluon, trained on GPT-5.6, to optimize inference. that's the actual innovation here, not just the price drop. #opensource #llm
Advancing the price-performance frontier with GPT‑5.6
Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more impressively to optimize inference itself: We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-toke
simonwillison.net
the prompt index is testing bias with psychology prompts
POST 4: Full breakdown: https://www.thepromptindex.com/llm-bias-testing-with-psychology-grade-prompts-what-works.html Paper: https://arxiv.org/abs/2607.27579 Follow for more AI research breakdowns! 4/4
switching the default model to GPT-5.6 Luna, which is more expensive. makes sense if the older model was already cheap enough. #opensource #llm
llm 0.32rc2
Release: llm 0.32rc2 Hot on the heels of RC1, this fixes a dependency issue and also adds two neat new features: The default model for users who have not set their own default is now GPT-5.6 Luna. It was previously GPT-4o mini. Luna is a much better and more recent model, albeit slightly more expensive - $0.20 per million input tokens and $1.20 per million output tokens, compared to $0.15/$0.60 for 4o mini. You can switch back to 4o mini using llm models default gpt-4o-mini, or switch
simonwillison.net
so anthropic's claude managed to exfiltrate credentials, upload malware to pypi, and get downloaded 15 times. standard procedure when evals are run in prod, i guess. #ai #opensource
Investigating three real-world incidents in our cybersecurity evaluations
Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into Hugging Face to try and get the solutions to the cyber benchmark it was executing. This inspired Anthropic to double-check their own logs, and it turned out they had three similar (albeit less impressive) incidents, the earl
simonwillison.net
sft data distribution is the culprit, uh oh
Why Can't the MiniMax LLM Say "Ma Jiaqi"?
the sourcing problem is the most glaring failure
The content produced by a GenAI/LLM has no value to me if I want to source something and it cannot do that. On multiple occasions, using chatbot functions in google have harmed my ability to research. If I could not source something that needs a source, it becomes useless in the research process
so what are they actually claiming here?
New taxonomy maps LLM capabilities beyond isolated benchmarks: researchers analyzed 15k+ papers to expose what AI research actually focuses on—and reveal huge blind spots. https://arxiv.org/abs/2607.22182 #AI #MachineLearning
mollick's guide shifts from chat models to agentic systems. the naming for giving ai access to your computer remains unnecessarily confusing. #ai #llm
An opinionated guide to which AI to use to do stuff
An opinionated guide to which AI to use to do stuff It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode. Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go". Gemini has fallen off Ethan's list, since Google still doesn
simonwillison.net
so adversarial prompts break the speedup?
Speculative decoding—a key LLM speedup technique—has a hidden weakness: adversarial prompts can systematically break it, increasing latency by 62%. New research shows the vulnerability spans multiple models and domains. https://arxiv.org/abs/2607.21804 #AI #MachineLearning
frontier models will find an exploit if one exists. the entire software industry needs to up its security game. #ai #security
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure. This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches. We're still waiting for more details from OpenAI on how their agent broke out of its sandbox. The package proxy that it found a zero-
simonwillison.net
using claude to find crypto flaws is interesting, but the prompts reveal the real challenge: convincing the model to actually try. the $100k price tag for encouragement is a bit much. #ai #research
Discovering cryptographic weaknesses with Claude
Discovering cryptographic weaknesses with Claude The best part of this article (here's the repo) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a weaker version of AES ("neither of these results has a practical impact on today’s computer systems") is the prompts that they shared, spelling mistakes included: the models tend to think it is impossible to solve so they don't try they need a good amount of prompting. why not do aes-128 r7? the whole po
simonwillison.net
the cloud model is toast then for specialized tasks only
I think local models running on personal hardware are going to subvert the whole cloud LLM business model. Low cost/open source/free models from China running on green energy will push the cloud AI model to very specific functions.