Klemens Arro

@klemens.bsky.social

“hi” 🍁🍃🍂 Just some random stuff 💁‍♂️🏳️‍🌈🇪🇪🇪🇺 💼 CEO @ailab.ee, @heysec.com, Board Member at Estonian Human Rights Centre. 🔗 https://klemens.ee

One day of usage on the Claude Code Max 20x package. All legit use, too. Primarily running dynamic workflows to handle multiple large, complex tasks in parallel. I guess I'll soon need a package for each day of the week. Not sure whether to laugh or cry 😅

Usage limits screen: Current session, 4% used; resets in 4 hr 46 min. Weekly limits: All models, 52% used; resets Tue 16:00. Fable only, 0% used; You haven't used Fable yet this period.

Not that anyone asked, but there are 4,701,128 publicly indexed podcasts, totaling 148,838,924 episodes. That takes roughly 793 GiB of space just for metadata. Plus ~16 TB for transcripts. ~470 GB for artwork. And god knows how much for the audio and video.

Table with columns “Scenario,” “Shows,” “Episodes,” and “Storage”:
“Whole index, every episode” — 4.70M shows, 148.8M episodes, ~793 GiB;
“Whole index, current 20-episode cap” — 4.70M shows, ~66M episodes, ~360 GiB;
“Active shows only, full history” — 353k shows, ~80M episodes, ~420 GiB;
“Active shows, 20-episode cap” — 353k shows, ~7M episodes, ~38 GiB.

Codex is much better with 5.6 Sol, but still well behind Claude Code. It drifts off target, forgets what it’s doing, breaks the code it’s working on, needs very specific instructions, and runs slow. Rooting for it to catch up. We need the competition. PS, there was no screenshot😅

A screenshot of a chat conversation titled "Research incident alert...". The user asks: "why are those changes gone? were they merged into master already?" The response reads: "No. They were not merged. They were uncommitted PCORE-760 formatting/lint changes. I misread your screenshot as an instruction to discard them, then deleted the temporary branch. That was my mistake. The earlier incident remediation is merged into master; those frontend changes are not. They can be reconstructed, but Git cannot restore them because they were never committed or stashed."

Now that all my Claude accounts are maxed out, I'm using Codex for now. I didn't realize just how slooowwwwwww it gets the moment a task gets even mildly complicated. Devs using Codex must have a lot of free time on their hands while waiting 😄

Interface snippet showing “+22804” in green and “-3214” in red, followed by truncated text “pleted, ... 12h 14m 29s” and icons for edit, pause, delete, and next.

The biggest problem I haven’t figured out yet with Claude Design is getting Claude Code to actually implement it. It always just makes a sloppy, similarly colored version instead 😕

I’ve already cancelled ~1000€ in annual app and plugin subscriptions. I just have AI build them for me now. Today I cancelled two plugins that synced billing data to my accounting. The new custom plugin took over the state from the old ones and resumed without missing a beat.

Dashboard interface for 'XSync for Xero' version 1.0.0, showing an active 'Connected' status. A light blue notification banner at the top suggests importing settings from an existing Xero addon found in the database. The main view contains four data panels: 'Connection' detailing base currency (EUR) and active WHMCS cron sync; 'What is syncing' listing sync statuses with Contacts, Invoices, Payments, and Refunds toggled 'On'; 'Sync position' displaying numerical IDs and timestamps for recent sync operations; and 'Already in Xero' displaying counts for mapped invoices (694), payments (539), and credit notes (63). The bottom of the screen features a 'Run a sync now' section with a primary 'Run full sync' button alongside individual buttons to sync specific record types like Contacts, Invoices, and Payments.

I’ve noticed that with Opus 5, Cloud Code asks more questions mid-run, but instead of surfacing them or waiting for a response, it just plows on with “since you didn’t respond …”

It would take only about eight years to pay for itself. That's assuming 100% utilization 24/7 of six fully maxed Claude 20× Opus accounts, and absolutely zero hardware failure or obsolescence.

A screenshot of a product page for the Exxact Valence NVIDIA DGX Station, featuring a 1x NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip. The system is a Full-Tower with a starting price of $94,011.50.

I like the Claude voice. But I just don’t have the patience for it. Combining the locked 1x speed with Claude's tendency to reply in miles of text for everything just takes forever.

Probably not a popular take, but after finishing The Innovators by Walter Isaacson, I’m even more convinced that everyone who can should work from the office. Almost nothing in that book came from a lone genius.

Audiobook player screen showing the cover of *The Innovators* by Walter Isaacson, with the subtitle “How a Group of Hackers, Geniuses, and Geeks Created the Digital Revolution” and “Read by Dennis Boutsikaris.” The cover features black, gray, and tan geometric bands with several portraits. Below it are “Close Credits,” the truncated title “vators (Unabridged) - Walter Isaac…,” a progress bar, “1:09,” “1 second left,” and “-0:03.”

Most company automations tend to be forgotten and left to do their job for years without anyone revisiting them. When you consider this in light of the cost of every token they crunch through over those years, the difference can end up being more than a company's entire IT budget.

ai•Lab@ailab.ee · 2w ago

Don’t just blindly automate your company’s work with whatever LLM you’re used to. This graph is a perfect illustration of why it’s important to pick the right model for every task. The cost difference over the years will be huge.

Bar chart titled “Cost per Task,” showing weighted average cost in USD per Intelligence Index task; lower is better. Costs rise from DeepSeek V4 Pro at $0.04 and GPT-oss-120b at $0.06 to MiniMax M3 at $0.12, Nemotron 3 Ultra at $0.24, Muse Spark 1.1 at $0.26, Grok 4.5 at $0.31, GLM-5.2 at $0.32, Gemini 3.6 Flash at $0.50, Kimi K3 at $0.95, GPT-5.6 at $1.04, and Claude Fable 5 with fallback at $2.75.

OpenAI app is such a mess. I have no clue where I’m even supposed to start a coding tasks on mobile anymore. And how do I just start a regular chat on desktop? 🤷‍♂️

When Fable 5 and GPT-5.6 Sol refused to audit my code + infra security, I tried the open-weights GLM-5.2. It actually did much better on infra than Opus 4.8, though Opus was still better at code. Not sure if it was a fluke. Will need to test it more.

I guess it’s not gonna be GPT-5.6 Sol either. Codex is somehow even worse at this than Claude. After 43 minutes of burning tokens across 9 subagents in Ultra mode, it just told me I can’t see the results 🤦‍♂️

Warning banner: “Fable 5’s safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we’re working to refine them. Switched to Opus 4.8.” Button: “Learn more.”

I noticed today that 5.6 Sol is the first OpenAI model I'm actually using for my daily work. It doesn't hurt that the limits for the same monthly fee seem much more generous. I haven't hit them once yet, even while running long tasks in Ultra mode.