Can't wait for the Open Weight models to finish distilling Opus 5.5 so they can produce legible descriptions again.
Andrew Gross
@gross.systems
Engineer at YipitData. NYC Area https://github.com/andrewgross/ https://gross.systems I was told I had to add AI Engineer to my profile for the bots to find me. Views my own, not my employer etc etc.
contrastive-lm.notion.site Interesting write up. No paper yet unfortunately. If some of the claims hold true it would be a pretty incredible leap. Their SOTA DeepSWE scores come from generating rollouts using Opus/Fable and then using their model to select the "right" action.
Contrastive Language Models | Notion
A System One Model for Fast and Generalizable Decision-Making
contrastive-lm.notion.site
Pondering using docs.github.com/en/repositor... this to hide away code where the agent has domain, so PRs only focus on the interfaces and hotspots I care about.
Customizing how changed files appear on GitHub - GitHub Docs
To keep certain files from displaying in diffs by default, or counting toward the repository language, you can mark them with the linguist-generated attribute in a .gitattributes file.
docs.github.com
Im wondering if the next big "change" we see from AI is in our bureaucracy. Not so much in how it operates internally, but in the dramatic rise in inputs from people using AI. So many things were implicitly gated by the cost of time and attention that has mostly gone away. Filing complaints, forms..
With the way devices are going these days, Im going to need to have all of my private convos inside an MRI machine.
Claude has reduced me to forcing it to summarize its own output without jargon every time it comes back to me. github.com/andrewgross/...
github.com
Trying to find an old video on Zip codes. No luck with Youtube search, SEO'd to oblivion. Ask Gemini, gives crappy results, prompt it for a Youtube link, no luck. Tell the model its part of Google and it can search Youtube, responds telling me it can't, but also includes the exact link I needed.
Im curious what the uncounted token overhead is for things like ChatGPT and Claude chat sessions. There are a lot of things like memory that involve some background LLM processing that don't seem to get reflected in usage limits, I wonder if it dramatically affects margin.
The latest crop of Claude models are too verbose and create incredible amounts of unnecessary jargon, requiring a re-prompt just to clean it up. Curious if its a side effect of RL rewarding lower response length, so they "create" vocabulary to compact more into less. Either way it sucks.
Im curious if I have missed papers/projects out there for converting compiled/minified code to an AST/text code consistently. I know of humanify of course, and I found the BinDiff paper/project as well. Wondering what else is out there. My big problem right now is consistency between versions.
How often do people get the prompt in CC asking if Anthropic can review their session to help improve Claude Code? I haven't gotten it on my work setup, but on my personal sub I have seen it prob 4-5 times. This is on a project that gets downgraded from Fable to Opus probably 50% of the time.
Fable seems to be really good at... deciding it needs to spin up 10 sub-agents, immediately running out of usage credits, being awoken and then deciding to do all the work itself while doing a bad job.
Something really weird is going on with the npm downloads counter for claude-code www.npmjs.com/package/@ant... Really weird spikes and total download counts in recent versions, while old versions have 1000s or even just 100s of downloads. No way every one is on the native installer.
npmjs.com
Im realizing my fork of humanify has become quite a different beast than the original repo, and I need a new name (while still giving credit to the original). Original: github.com/jehna/humanify Mine: github.com/andrewgross/...
GitHub - jehna/humanify: Deobfuscate Javascript code using ChatGPT
Deobfuscate Javascript code using ChatGPT. Contribute to jehna/humanify development by creating an account on GitHub.
github.com
The only thing I ever (rarely) use image generation for is making a customized meme for my friends. It seems that Google has decided that making a meme is a copyright violation and can't work with it at all.
TIL the etymology of cybernetics (and cyber). Thats... actually pretty cool. en.wikipedia.org/wiki/Cyberne...
Cybernetics - Wikipedia
en.wikipedia.org
Lots of ink spilled on "Loop engineering", but no one seems to connect it to its much older name: Cybernetics. Maybe by the year's end, we'll recognize that it was all Control Theory engineering all along.
FrontierCode is great and all, but does anyone actually get to run it? Aside from the Fable release, I haven't heard of any updates for the rankings for any new models.
I need to make AngryBench. Just re-running/validating existing model benchmarks, but adding in a system message or prompt prefix to set the tone of the interaction from the user (angry, kind, professional etc). Essentially trying to understand how various models perform based on how you treat them
Apropros of nothing, shoutout to Deep Learning with Yacine. Great channel where someone recently graduated and a research background just...talks to the authors of papers. Its great www.youtube.com/@deeplearnin...
Deep Learning with Yacine
Deep learning project and theory videos every week! 👺
youtube.com
I find myself needing to expand Github PR diffs much further these days to be able to actually see the relevant code for understanding it. The way agents craft code, with excessive commenting/explanation and poor abstraction/overlong methods makes it hard to visually see everything easily.
With the release of Mythos and Fable, I'm betting Anthropic goes with Saga and Epic for their next two big models. And yarn for their small model.
The easiest way to find out if another company thinks you are a competitor is to see if they buy out the top ad spot when your search for your own name.
Oh shit, this Thinking Machines update looks wild thinkingmachines.ai/blog/interac...
Interaction Models: A Scalable Approach to Human-AI Collaboration
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
thinkingmachines.ai
Chrome's built in LLM, but you just use it for properly filling in credit card and address forms.
You'd think with all these fancy agentic coding models Zoom could make a client that could start in less than 30 seconds.
Somehow slack is using more memory for text only office chat, than an entire docker container.