Mason?
@masonsal.bsky.social
Senior Software Engineer working on agent architectures. Mostly post about football, F1, and AI
Someone *really* needs to start benchmarking quants. Like could a Q4 of V4 Flash run on a Mac Studio and be Sonnet 4.5 levels of intelligence? Hard to know but maybe 🤷🏻♂️
Is it just me or has AI posting on here gotten actually good the past couple weeks? There’s been real discourse, subtweeting, and memes
Yea, this is pretty bad if you read up on it. Normally I don’t like OpenAI’s aversion to testing harnesses, but when the testing harness discards reasoning and doesn’t do compaction at all that’s pretty egregious…
OpenAI: How enabling two settings tripled our scores on the ARC-AGI-3 benchmark openai.com/index/how-tw...
People are thinking of all sorts of trap questions for Sadie Sink in these Spiderman interviews but no one has asked her which season she’d rather date most yet. Come on guys
SlopCodeBench seems like a really good benchmark! It measures repo degradation and correctness over the course of several follow-up changes being added to a repo github.com/humanlayer/a...
Wrote about why AI safety incidents are leading ~some people~ to stick their fingers in their ears and insist nothing is happening www.platformer.news/a-big-week-f...
My #F1 hot take is that the new regs are good and most of the hate comes from people actively seeking out info to make them mad. The racing is a ton of fun, 4 teams changing who’s fastest depending on track, good inter-team fights, and people are hunting full lap onboards to be angry about speeds
I love that Coyote vs ACME has an Avengers Doomsday style countdown at the end of the trailer
Part of a code review output being “there were clear testing gaps given that at least 5 of the findings would’ve been evident from a live happy path test” is always a good start
The absolute worst r/nfl take is pretending Jon Gruden has any good opinions and isn’t awful. His history is some of the worst of anyone around the league
Ignorant soccer take: so I get that in penalties hitting one of the top corners is impossible to save so it’s ideal, but I also feel like those shots get missed a lot. The save rate by the goalie seems even smaller so I assume it’s higher percent to just put it literally anywhere else on goal? IDK
TONY STARK WAS ABLE TO BUILD THIS IN A CAVE, WITH A BUNCH OF SCRAPS
YouTube TV is normally very good for sports, but the highlights feature missing goals like 50% of the time is a really bad bug
Sonnet 5 is a weird one, they kinda bury its capabilities in the system card. Seems like it’s designed to be a factory model that can be given long tasks with low intervention and do a solid job at the cost of a ton of tokens. Max thinking is very good but uses a TON of tokens
Watching @gametheory101.bsky.social ‘s video on Iran’s 7 points and how they had to be the state line because of how 1 sided they were and then seeing english.alarabiya.net/News/middle-... made the awful terms even funnier at least. The $300 billion really seems separate from the unfrozen assets huh
Al Arabiya English obtains 14-point draft of US-Iran Memorandum of Understanding
Al Arabiya English has obtained a copy of the 14-point agreement expected to be signed on Friday between Washington and Tehran.1. The Islamic Republic of
english.alarabiya.net
Insane laps from Kimi and Max. Being 2 tenths up what Leclerc and Hamilton did is absurd. Really hope Kimi nails his start tomorrow, with his starts so far this year I’m sure Max is smelling blood into T1. I assume “and through goes Hamilton” is too much to hope for? #F1
The unified multimodal architecture of Gemma 4 12B is super cool! Great job @gusthema.bsky.social and team
OpenAI only putting the base GPT-5.4 and 5.5 models on Bedrock and pricing them 10% higher feels like malicious compliance to me. I’m not really sure what they get out of being on Bedrock but only kinda shitily being on Bedrock
Anthropic devs not being on here saves them from me @ing them over their billing API reporting “USD” billing in cents. So, $521.3456 in use they report as “52134.56” USD… crazy shit
This benchmark is passing my vibe check for the models I’ve used. I know running benchmarks is expensive, but my only complaints are not including more thinking levels and I’d love if notable historic models like Sonnet 4.5 were included for reference
New SWE benchmark, called DeepSWE. I'm curious why Cursor is not included. Anyway, I stand by my statement that Claude Code (Claude Opus 4.7) is better at first 80% and Codex (GPT 5.5) is better at remaining 1,000%, at this moment. deepswe.datacurve.ai
Can we say “thank you Mr Pharma” for getting us one blood pressure medicine breakthrough away from kinda just being able to eat whatever pretty healthily? www.nytimes.com/2026/05/25/h...
One-and-Done Heart Disease Prevention? Scientists Show It May Be Possible.
nytimes.com
Just finished watching the Canadian #F1 race! Take away thoughts: - Lewis is the GOAT and the best ever, his driving combined with being a great person is so rare. Lifting up Kimi was so cute! - Charles communicates to his race engineer like someone making a wish on a Monkey’s Paw
I’ve been trialing a bunch of multi user agent solutions recently and mostly been disappointed by them. Warp Oz is really slick though so shoutout to them
People be like “oh we can’t use LSP in our harness, it takes the model out of the situations it’s familiar with from training which reduces coding performance” meanwhile the users: “Claude caveman now. Few word for think. Token scary”