Mason?

@masonsal.bsky.social

Senior Software Engineer working on agent architectures. Mostly post about football, F1, and AI

Someone *really* needs to start benchmarking quants. Like could a Q4 of V4 Flash run on a Mac Studio and be Sonnet 4.5 levels of intelligence? Hard to know but maybe 🤷🏻‍♂️

People are thinking of all sorts of trap questions for Sadie Sink in these Spiderman interviews but no one has asked her which season she’d rather date most yet. Come on guys

My #F1 hot take is that the new regs are good and most of the hate comes from people actively seeking out info to make them mad. The racing is a ton of fun, 4 teams changing who’s fastest depending on track, good inter-team fights, and people are hunting full lap onboards to be angry about speeds

Part of a code review output being “there were clear testing gaps given that at least 5 of the findings would’ve been evident from a live happy path test” is always a good start

The absolute worst r/nfl take is pretending Jon Gruden has any good opinions and isn’t awful. His history is some of the worst of anyone around the league

Ignorant soccer take: so I get that in penalties hitting one of the top corners is impossible to save so it’s ideal, but I also feel like those shots get missed a lot. The save rate by the goalie seems even smaller so I assume it’s higher percent to just put it literally anywhere else on goal? IDK

Sonnet 5 is a weird one, they kinda bury its capabilities in the system card. Seems like it’s designed to be a factory model that can be given long tasks with low intervention and do a solid job at the cost of a ton of tokens. Max thinking is very good but uses a TON of tokens

Watching @gametheory101.bsky.social ‘s video on Iran’s 7 points and how they had to be the state line because of how 1 sided they were and then seeing english.alarabiya.net/News/middle-... made the awful terms even funnier at least. The $300 billion really seems separate from the unfrozen assets huh

Al Arabiya English obtains 14-point draft of US-Iran Memorandum of Understanding

Al Arabiya English has obtained a copy of the 14-point agreement expected to be signed on Friday between Washington and Tehran.1. The Islamic Republic of

english.alarabiya.net

#F1 has finally figured out how to fix the #MonacoGP, you simply need to change up the rules slightly so it’s a penalty fest and then the commentators get to debate what exactly changed that’s causing so many penalties. That’s what a good car race is made of

Insane laps from Kimi and Max. Being 2 tenths up what Leclerc and Hamilton did is absurd. Really hope Kimi nails his start tomorrow, with his starts so far this year I’m sure Max is smelling blood into T1. I assume “and through goes Hamilton” is too much to hope for? #F1

OpenAI only putting the base GPT-5.4 and 5.5 models on Bedrock and pricing them 10% higher feels like malicious compliance to me. I’m not really sure what they get out of being on Bedrock but only kinda shitily being on Bedrock

Anthropic devs not being on here saves them from me @ing them over their billing API reporting “USD” billing in cents. So, $521.3456 in use they report as “52134.56” USD… crazy shit

This benchmark is passing my vibe check for the models I’ve used. I know running benchmarks is expensive, but my only complaints are not including more thinking levels and I’d love if notable historic models like Sonnet 4.5 were included for reference

Sung Kim@sungkim.bsky.social · 2mo ago

New SWE benchmark, called DeepSWE. I'm curious why Cursor is not included. Anyway, I stand by my statement that Claude Code (Claude Opus 4.7) is better at first 80% and Codex (GPT 5.5) is better at remaining 1,000%, at this moment. deepswe.datacurve.ai

Just finished watching the Canadian #F1 race! Take away thoughts: - Lewis is the GOAT and the best ever, his driving combined with being a great person is so rare. Lifting up Kimi was so cute! - Charles communicates to his race engineer like someone making a wish on a Monkey’s Paw

I’ve been trialing a bunch of multi user agent solutions recently and mostly been disappointed by them. Warp Oz is really slick though so shoutout to them

People be like “oh we can’t use LSP in our harness, it takes the model out of the situations it’s familiar with from training which reduces coding performance” meanwhile the users: “Claude caveman now. Few word for think. Token scary”