Andrew Gross

@gross.systems

Engineer at YipitData. NYC Area https://github.com/andrewgross/ https://gross.systems I was told I had to add AI Engineer to my profile for the bots to find me. Views my own, not my employer etc etc.

Im wondering if the next big "change" we see from AI is in our bureaucracy. Not so much in how it operates internally, but in the dramatic rise in inputs from people using AI. So many things were implicitly gated by the cost of time and attention that has mostly gone away. Filing complaints, forms..

Trying to find an old video on Zip codes. No luck with Youtube search, SEO'd to oblivion. Ask Gemini, gives crappy results, prompt it for a Youtube link, no luck. Tell the model its part of Google and it can search Youtube, responds telling me it can't, but also includes the exact link I needed.

Screenshot of a conversation with Google Gemini about trying to find a video. The model responds that it can't search Youtube indexes directly and for me to give it more info on what to look for.  However, it also includes a link to the exact Youtube video I was searching for.  Its not linked in the text, or a reference, just the last thing in the message, completely unrelated to the text.

Im curious what the uncounted token overhead is for things like ChatGPT and Claude chat sessions. There are a lot of things like memory that involve some background LLM processing that don't seem to get reflected in usage limits, I wonder if it dramatically affects margin.

The latest crop of Claude models are too verbose and create incredible amounts of unnecessary jargon, requiring a re-prompt just to clean it up. Curious if its a side effect of RL rewarding lower response length, so they "create" vocabulary to compact more into less. Either way it sucks.

Im curious if I have missed papers/projects out there for converting compiled/minified code to an AST/text code consistently. I know of humanify of course, and I found the BinDiff paper/project as well. Wondering what else is out there. My big problem right now is consistency between versions.

How often do people get the prompt in CC asking if Anthropic can review their session to help improve Claude Code? I haven't gotten it on my work setup, but on my personal sub I have seen it prob 4-5 times. This is on a project that gets downgraded from Fable to Opus probably 50% of the time.

Fable seems to be really good at... deciding it needs to spin up 10 sub-agents, immediately running out of usage credits, being awoken and then deciding to do all the work itself while doing a bad job.

The only thing I ever (rarely) use image generation for is making a customized meme for my friends. It seems that Google has decided that making a meme is a copyright violation and can't work with it at all.

FrontierCode is great and all, but does anyone actually get to run it? Aside from the Fable release, I haven't heard of any updates for the rankings for any new models.

I need to make AngryBench. Just re-running/validating existing model benchmarks, but adding in a system message or prompt prefix to set the tone of the interaction from the user (angry, kind, professional etc). Essentially trying to understand how various models perform based on how you treat them

I find myself needing to expand Github PR diffs much further these days to be able to actually see the relevant code for understanding it. The way agents craft code, with excessive commenting/explanation and poor abstraction/overlong methods makes it hard to visually see everything easily.

With the release of Mythos and Fable, I'm betting Anthropic goes with Saga and Epic for their next two big models. And yarn for their small model.

The easiest way to find out if another company thinks you are a competitor is to see if they buy out the top ad spot when your search for your own name.