Avery Yen

@averyyen.bsky.social

Empirical AI Researcher. Once a Pivot, always a Pivot. Leave me anonymous feedback: https://www.admonymous.co/avery-yen

This is basically true, but ignores the fact that back when I used to review/work with human slop code for a living, it was way harder to pass it off as plausibly good. That's the superpower of the LLM. (Reading LLM code also hurts my brain personally)

Sung Kim@sungkim.bsky.social · 2w ago

When coding with AI agent, using either CLI or agentic UI, do you review the code? or just review the functionalities? Me. Personally, I do not review the code. People may say AI-generated slop, but have you seen human-generated slop?

I have so many other random benchmark questions and so little time. How do K3 and Qwen 3.8 bench on unseen or purely continuous tasks? We have claims that they're distilled and benchmaxed but does it matter if they generalize? E.g. stuff like post cutoff evals and optimization benchmarks

Avery Yen@averyyen.bsky.social · 2w ago

We're pretty sure that swapping harnesses actually changes effective agent capability. But recently everyone and their mother has a new coding agent not to mention the open/agnostic ones. Someone needs to run a Harness Bench across like three different models and ten different harnesses or something

We're pretty sure that swapping harnesses actually changes effective agent capability. But recently everyone and their mother has a new coding agent not to mention the open/agnostic ones. Someone needs to run a Harness Bench across like three different models and ten different harnesses or something

OpenAI's internally deployed models hacking Hugging Face does not seem to have been unpredictable or inevitable. We talked about the root of the problem & what policymakers can do about it back in February. Props to Joe for hitting the nail on the head.

BildBildBild

Look, I know social media isn't everything, but it certainly measures SOMETHING. "Meet Kimi K3" has 16M views in 4 days versus Fable 5 at 773k views (topped by a few others including just barely 1M views for Claude Code 1 year ago).

The channel page sorted by most popular Kimi AI videos on YouTube. "Meet Kimi K3" has 16M views in 4 days.The Anthropic channel page sorted by most popular videos on YouTube. "Introducing Claude Fable 5" has 773k views in 1 month.

It seems like agentic development benchmarks are strongly driving competition right now. And I hear anecdotal reports of regression on other, less RLVR-able tasks. How long before the idea of a single "frontier" breaks, and we see differentiation emerge between dev models and chat models?

It feels like the Anthropic/Pentagon scuffle was in the distant past, but these kinds of events will keep reverberating in the safe and beneficial deployment of emerging tech for a long time to come. Thank you Alex for your writeup on this incident.

Alex Turner@turntrout.bsky.social · 3w ago

I think many hope that when things get bad enough, someone powerful will say "no." I tested that for months. Anthropic defended its red lines, Google did not. Pledges of conscience often vaporize on contact with power. My full account: turntrout.com/why-i-left-g...

New paper: LLMs encode harmful content generation in a distinct, unified mechanism Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities. 🧵

Bild

Sometimes, I look at the agentic AI work out there, and I think about how airbags and now things like blind spot detection and automatic emergency braking in cars doesn't require decision making of any "intelligence" level beyond if this then that. Agentic AI needs airbags, too.