It's been at least twenty years since having a PC as a primary computing device. Sure, I've had a couple gaming PC's, but daily driving? That's been a Mac for two decades. Today that changes when my new (to me) Thinkpad X1 Carbon arrives & I couldn't be happier to go full-time Linux (EndeavourOS).
Joshua White
@jrw14.whnc.me
I help figure out why people and things do the things they do at Apple. Books, hiking, kayaking, FOSS, self-hosting, Gen. AI, soccer, learning, family. NC. Be good to each other. Thoughts/opinions are my own. https://obscurnotes.offprint.app/
Funny thing happened this week... My primary inference provider closed shop on the subscription plan I was using and gave 2 days notice to find service elsewhere. I didn't panic, instead I took advantage of my homelab LiteLLM proxy setup, and loaded in inference from 5 different providers.
Heh, can't believe this was just yesterday. Suddenly with DS4 Flash, I feel a whole lot better about life 😂
Feeling quite like the idiot for letting my nearly unlimited, no-rate-limits inference plan from MiniMax go to hop on the new "hot thing" in Umans a couple months ago, only to have Umans announce they're shutting down their subscription entirely... ...in two freaking days.
I mean... holy crap. artificialanalysis.ai/models/deeps...
DeepSeek V4 Flash 0731 (max) - Intelligence, Performance & Price Analysis
Analysis of DeepSeek's DeepSeek V4 Flash 0731 (Reasoning, Max Effort) and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first toke...
artificialanalysis.ai
Uh. This is the biggest news in AI today, and it won't be particularly close. Deepseek's FLASH model, 200ish billion parameters, is obliterating models 3-5x it's size and costs pennies.
Uh. This is the biggest news in AI today, and it won't be particularly close. Deepseek's FLASH model, 200ish billion parameters, is obliterating models 3-5x it's size and costs pennies.
DeepSeek rolls out the official V4 Flash API in public beta, touting enhanced agent capabilities and benchmark scores "far surpassing" V4 Pro Preview (Newley Purnell/Bloomberg) Main Link | Techmeme Permalink
My problem is... it would be so easy to just get a OpenAI, Grok or Anthropic subscription. But I refuse. I refuse to support them. Having ideals sucks sometimes, especially right now. So, I have two days to research my next best options and honestly, a return to MiniMax is looking likely.
Feeling quite like the idiot for letting my nearly unlimited, no-rate-limits inference plan from MiniMax go to hop on the new "hot thing" in Umans a couple months ago, only to have Umans announce they're shutting down their subscription entirely... ...in two freaking days.
Feeling quite like the idiot for letting my nearly unlimited, no-rate-limits inference plan from MiniMax go to hop on the new "hot thing" in Umans a couple months ago, only to have Umans announce they're shutting down their subscription entirely... ...in two freaking days.
Hot damn.
Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: unsloth.ai/docs/models/... GGUF: huggingface.co/unsloth/Kimi...
Gross. I legit didn't think my opinion of Anthropic could get any lower...
One interesting thing about Anthropic is that they insist they’re trustworthy guardians of AI, yet they often seem to lack emotional intelligence, or much understanding of life outside their field.
A few days now using Laguna S 2.1 from @poolsideai.bsky.social, and yeah, it's really good. It's not GLM 5.2 (my daily driver), but it's significantly better than, say, MiniMax M3, or Nemotron Ultra... models 2x and 3x bigger. Which is asinine. Poolside worked some kind of dark magic...
Okay, so… in retrospect it seems some of this is on me. Turns out I wasn’t following the right people on BlueSky. 😂
A little disappointed to finally log on to Bluesky at end of my work day to see not a single peep from all the AI-related people I follow about @poolsideai.bsky.social releasing Laguna S today. A 120b model that cleans the clock of models 3-4x its size. A model that can run on consumer hardware.
This is much bigger news than whatever google released today. poolside.ai/blog/introdu...
Introducing Laguna S 2.1
Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
poolside.ai
A little disappointed to finally log on to Bluesky at end of my work day to see not a single peep from all the AI-related people I follow about @poolsideai.bsky.social releasing Laguna S today. A 120b model that cleans the clock of models 3-4x its size. A model that can run on consumer hardware.
A little disappointed to finally log on to Bluesky at end of my work day to see not a single peep from all the AI-related people I follow about @poolsideai.bsky.social releasing Laguna S today. A 120b model that cleans the clock of models 3-4x its size. A model that can run on consumer hardware.
Ah, when you are about to no longer be the best in the world at something, what do you do? Not try harder or change your ways, no, that's for LOSERS. You ban the challengers, of course! The American way, folks. (fwiw, I'm not worried, there is no plausible way to enforce an open weight AI ban)
Messed around and built my own LLM benchmark utilizing some Docker-in-Docker and a custom Rust harness. Couldn't find anyone else evaluating models in ways that I use them, so... built my own instead. 🤷♂️
People absolutely keep gilding thy lilies on how hard their problems are and they're mostly just not. Opus 4.5 could cook most things people throw at these models now but it feels like you're Doing More by roasting 3T-parameter tokens over an open fire Workflows continue to matter more
Meanwhile, Kimi K3 and GLM 5.2 are available to anyone for the same (much lower) prices. Blah. I'm so over US Frontier models at this point and all the people that insist on hanging on their every word/release like...
Oliver Twist 1948: "I want some more."
ALT: Oliver Twist 1948: "I want some more."
static.klipy.com
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. (1/2)
As big of a deal as Moonshot's Kimi K3 is... it's really just the tip of the iceberg of what the rest of this year will look like. - MiniMax Pro - GLM 5.3 (or perhaps straight to 6?) - DS 4.1 - Whatever Xaoimi is cooking (Mimo is GREATLY underappreciated) All rumored to be coming soon.
The future is open. K3 is here, GLM is cooking 5.3, MiniMax is sitting on 3 Pro, and Deepseek won’t get left behind. I hope the frontiers are taking notes.
Friends don't let friends use first-party coding harnesses.
is this bad?
Rolling power surges remind me why I'm happy I invested in a UPS finally this year. Whole house flickers out, I just keep plugging away on my laptop, network and home lab that's all still good on battery for about an hour. 🤗
The number of cars on US roadways has been growing at a rate of 10 to 15 million *every year.* There are now 300 million vehicles and only 6.5% are electric. But you’ll never see an article about how the growth of vehicle emissions is “so severe it’s almost incomprehensible.”
“Cornell researchers found that at the current rate of AI growth, the burgeoning industry could represent 24 to 44 million metric tons of carbon dioxide emissions by 2030, the equivalent of adding five to ten million cars to US roadways.”