Hrishi

@olickel.com

Previously CTO, Greywing (YC W21). Building something new at the moment. Writes at https://olickel.com

Jev models are incredible for data pipeline work. Correctly integrating @typesafeai models for entity resolution cut costs by 99% while coming within a percentage point of Fable and increasing throughput > 7x. Prompts, details and results: southbridge.ai/blog/jev-en...

Using system-one models inside high-throughput data pipelines

Jev (the system-one model) used with hank-based fallbacks can cut entity resolution costs and time by incredible amounts. Here's how.

southbridge.ai

Data engineering will follow the same path as software engineering. Work will increasingly move from expert senior practitioners to non-data folks with coding agents, the same way software has.

My suspicion is that the recent string of problems - sandbox escapes, 11-figure model companies not realising that the last two releases had wrong measurements - are happening because everyone is vibing their benchmarks.

This new generation of models are different beasts - I've had to change my workflow top to bottom to actually use them to build and review complex software. Here's what I know:

This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole thing obsolete. It also pushed forward the state-of-the-art for multi-stage media-to-transcript, with offmute-v2. I come with receipts: www.southbridge.ai/blog/offmut...

Bild

You know Fable is the real deal when it calmly picks up 4 threads (2 pure research + 1 code + 1 hank) that had been stuck for a month and just moves forward genuinely clever ideas and solutions Deepseek pushing cost, Mimo pushing speed and now Fable - new frontiers already

Agent-first systems (for agent emails, interfaces, etc) still have a big problem: What is the difference between an agent and a bot? How can this be made explicit and provably easy? Most of the internet is designed to monetize human attention while fighting a never ending war against bots.

What is a harness? I've been asked far too many times this week alone, so here's my simple working definition: A harness is a system prompt with basic tool definitions for read, write, exec and external calls. That's it. Optionally, a harness may or may not concern itself with:

Context graphs feel like mindmaps for agents. Similar kind of fun-to-look-at, wish-I-had-one thing that is nearly impossible to build well, or efficiently. I say this as someone who wasted A LOT of time on mindmaps.

I'll just leave this here for future hrishis and friends Complete non sequitur: sprites.dev is awesome - they feel like the first prototype of that elegant weapon from a more civilized age, if it had been weathering in a pyramid for 50 years. Crazy rough around the edges, but feels like the future.

Bild

bunx hankweave-trace can now directly upload (real-time or after the run) hankweave traces to braintrust or langfuse!

Bild

Skills will likely fail the same way that MCPs did. Why did MCPs fail? They were a wonderful idea, but the protocol was too open. Too many ways to do things means no one's in charge. Who's responsible for an MCP? Is it the service? the author? You? Who's running the MCP?

Fun little case in point about real-time connectors: @Calclavia tells me about Cursor Agent over lunch, go home, run Clausetta and point it to cursor agent, and now we can use it in hankweave!

New release! Harness engineering made easy: You can now switch harnesses in hankweave with a few characters. "sonnet" will run your prompts and code inside the Agents SDK. "gpt-5.3-codex" will run it in Codex. "pi/google/gemini-3.1-flash", "opencode/cerebras/glm4.7" - exactly what you'd expect.

Bild

GPT-5.4 is a very capable visual model. It's now part of a TINY club (with opus 4.6) as the only two models I've tested that can actually see. It's also way cheaper than opus!

This is the hardest it will ever be to write Hanks Is what I tell myself as I sit down to write a new one Here's some hank building to vibe to for Saturday

Bild

I'll be honest: we didn't necessarily want to build hankweave - we didn't have the time to build a runtime. (Who does?) The problem is that hankweave becomes an inescapable requirement once you add up some self-evident truths about AI today.

You're all sleeping on Haiku. Haiku 4.5 is a new beast - it isn't the 'glorified autocomplete' we've come to think of it as - it's a capable, strong agentic model that's crazy fast and dirt cheap.

My personal most awaited feature in Hankweave is out: budgets! With budgets, users at runtime can say • Don't spend more than $100 • Don't take longer than 15 hours or 750K tokens With budgets, builders at build-time can say • Give this step 20% of the budget • Loop steps 1 and 2 until we hit $5

It's fun if you make a game out of it - I try and see how many flashes it'll take me to read what Claude Code is trying to communicate to me in winks

Does instruction tuning create individuated sentience? This is something I've been wondering lately. Having interacted with LLMs since GPT-J, the difference between base models and instruction-tuned ones has always been crazy. We someone take autocomplete-level models and turn it into some*ones*.

Bild

Asked Opus for a reading recommendation at one point and thought nothing of it Picked it up recently and its Blindsight by Peter Watts Almost entirely about Chinese rooms not being Chinese rooms AGI has already happened trust me

BildBild

If we use agents without reading the output or having any interaction with the actual resulting code, are we the Chinese room, or is the agent? Which of us can claim to "understand"?

Bild