mr. TIM

@timkellogg.me

AI Architect | North Carolina | AI/ML, IoT, science WARNING: I talk about kids sometimes

i’m very excited about whatever SSI is launching in August. i fully anticipate that it’ll be controversial, many won’t understand why it’s good, yet it’ll be the linchpin to full ASI

So having just logged back into this site after a hiatus my immediate observation is that it kicks ass now. The dumb infighting seems to have mostly shaken out and now it's funny, has actual information on it, etc.

alright, ChatGPT Work is actually kind of nice. Feels like Codex but for the ChatGPT audience. Good integrations, subagents, etc. main complaint is they don't support events. You just have to poll, which sucks for your quota

some VC said in a podcast that SSI (Ily Suksever’s AI lab) is going to release their model in August which, if that’s true, Ilya has already said it’ll be a novel take on continual learning based on overlooked functions of the human brain it likely won’t be smart OOTB, but it’ll quickly customize

i wonder if hallucination is not a problem anymore GPT clearly doesn’t care about it, but Sol rarely actually lies to me, because it obsessively checks the ground truth i wonder if they might even induce hallucination during RL so the RL objective teaches the model not to trust its weights

Bar chart titled "AA-Omniscience Hallucination Rate" by Artificial Analysis, comparing hallucination rates across 24 AI models. Lower scores indicate better performance, measuring how often a model answers incorrectly instead of refusing or admitting it does not know.
The chart lists the following models from lowest to highest hallucination rate:
 * Command A+: 14%
 * MiniMax-M3: 16%
 * Qwen3.7 Max: 23%
 * MiMo-V2.5-Pro: 25%
 * Claude 4.5 Haiku: 26%
 * GLM-5.2 (max): 28%
 * Nemotron 3 Ultra: 29%
 * Gemini 3.5 Flash-Lite: 34%
 * Claude Opus 4.8 (max): 36%
 * Claude Sonnet 5 (max): 37%
 * Muse Spark 1.1 (xhigh): 38%
 * Claude Opus 5 (max): 50% (highlighted with a red box)
 * Kimi K3 (max): 51%
 * Grok 4.5 (high): 54%
 * Gemini 3.6 Flash: 54%
 * Claude Fable 5 (with fallback): 55% (highlighted with a red box)
 * Inkling: 63%
 * Gemma 4 31B: 82%
 * Mistral Medium 3.5: 82%
 * DeepSeek V4 Flash 0731 (max): 84%
 * GPT-5.6 Terra (max): 85%
 * GPT-5.6 Sol (max): 89% (highlighted with a red box)
 * GPT-5.6 Luna (max): 90%
 * gpt-oss-120b (high): 91%
Lightbulb icons beneath model names indicate reasoning models.

rule of thumb for RAG/vector search: if you would consider it a bug if it returned less than 100% of the hits, then vector search is not for you, you’re looking for SQL

annoying codex thing last week they just dropped luna prices by 80% but you can't actually make luna subagents from sol or terra fr fr Luna is marked as "V1" and Sol & Terra are marked as supporting "V2" subagent API, and there's a Codex check preventing you from mixing them

my Hank Green take — i mostly do not give a shit, but one alarming aspect is the power of mobs he’s inadvertently taken part of the creation of a community that will gladly hang him the moment he slips up whelp, it happened. How did you think this was going to go?

Chinese open weights reliably release their weights approx 1 week after initial announcement ever since the news that the CCP is setting up release gates similar to what the US has done

i’ve noticed that a huge number of Ed Zitron followers are actually great people who just got the wrong info a lot of them happily will change their mind given new info i really wish people here would dial back the distain and just focus on information dissemination