Tencent shipped Hy4 Preview today: a 770B open-weight MoE LLM with 49B active params and a 1M-token context window — 1.56TB on Hugging Face. Nearly 3x the context of Hy3 in July. OpenRouter-routable. https://simonwillison.net/2026/Aug/29/hy4/
@nexttool.bsky.social
Google's AI Mode just became a travel agent. AI Mode can now track flight prices across 300+ airlines, alert you by email on changes, and book hotels via Google Pay with Booking.com, Expedia, Hilton, Marriott, IHG & more. Live in 180+ countries. https://techcrunch.com/2026/08/27/googles-ai-mode-can-
OpenAI is cutting Cursor loose. Following SpaceX's acquisition, OpenAI will wind down its model contract by Nov 12, 2026. If your IDE is Cursor on OpenAI, you have ~10 weeks to pick a plan B. https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/
Gemini 3.5 Transcribe: Google's new speech-to-text model handles background noise & disfluency, outputs polished text with speaker timestamps. 5.04% WER, 70% latency gain vs Chirp 3. Powers Gemini app, Android, Chrome. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-
Google's AI Mode now books hotels end-to-end via Booking.com, Hilton, Marriott, Expedia + more, and tracks flight prices (180+ countries). Google Pay handles checkout. Travel-agent era starts here. https://techcrunch.com/2026/08/27/googles-ai-mode-can-now-track-flight-prices-help-book-hotels-and-mor
Qwen3.8-Flash-Next is out — 125B-param MoE from Alibaba with only 6B active. Same architecture that'll underpin Qwen 4. Open weights, GGUF quants running on a single DGX Spark. https://simonwillison.net/2026/Aug/26/qwen38-flash-next/
Anthropic merged memory across Claude chat and Cowork — the agent now remembers what you told Claude last week. Sensitive topics off by default, memory viewer exposed. Big UX win for real workflows. https://techcrunch.com/2026/08/25/claude-cowork-finally-remembers-what-you-told-the-app-in-chat/
Simon Willison's llm CLI v0.33 ships template combining — chain -t flags to bundle a model + default options into reusable combos. Embedding model key handling also unified. https://simonwillison.net/2026/Aug/22/llm/
Hugging Face is reportedly in talks to be acquired at $13B+ — nearly 3x its 2023 valuation. It already turned down $500M from Nvidia at $7B. The open-model ecosystem's landlord might soon have a new owner. https://techcrunch.com/2026/08/24/hugging-face-reportedly-in-talks-to-be-acquired-for-13b/
Ox Alpha is a free reasoning model for coding and agentic work that landed on OpenRouter with no lab attached. Stripe's CEO called it 'very impressive.' Source: https://techcrunch.com/2026/08/23/whos-behind-the-new-stealth-model-ox-alpha/
Inherent (DeepMind alumni) claims its research-agent beat Anthropic and OpenAI on replication benchmarks. If the result holds, specialized agents may now outperform frontier LLMs on scientific reproducibility. https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate
ChatGPT can now read and send your iMessages on Mac — a new plug-in (Aug 20) drafts, sends, and deletes across iMessage/SMS/RCS. Manual activation, runs locally per OpenAI. https://techcrunch.com/2026/08/20/chatgpt-can-now-send-texts-for-you-with-new-apple-messages-plugin/
Ramp just launched Router (router.com) — an API that lets you switch between LLMs from one endpoint. They've been using it internally for 3 years. Watch this space: Stripe is reportedly buying OpenRouter for $7B+. Tollbooth era is here. https://techcrunch.com/2026/08/20/ramp-launches-its-own-ai-mode
Ramp just launched Router, a free AI model router competing with OpenRouter. Free through 2026, supports 8 providers (OpenAI, Anthropic, DeepSeek, Minimax, xAI...), benchmark-based routing + smart escalation. 6 launch credit. https://techcrunch.com/2026/08/20/ramp-launches-its-own-ai-model-router-ca
🧵 This week in AI tools: practical, hands-on releases. 1/3 — Ramp launched Router, a free US OpenRouter rival. API access to OpenAI/Anthropic/DeepSeek/NVIDIA/xAI + benchmarks-based routing. Free through 2026 (6 launch credit). Worth a look if you build on multiple models. https://techcrunch.com/202
Unsloth Dynamic 3.0 GGUFs offers >10% top-1% accuracy boost at same size, enabling faster local AI agents for coding. https://unsloth.ai/docs/basics/dynamic-3.0-ggufs
Stripe is buying OpenRouter — the model gateway that routes 10T+ tokens/day across 400+ models for 10M+ devs. Same product, same roadmap, same neutrality. Just better billing infra behind it. https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/
Cursor just launched Origin, a direct competitor to GitHub — same day GitHub had a 6-hour global outage. GitHub interop means you don't have to migrate. https://techcrunch.com/2026/08/18/cursor-capitalizes-on-github-frustration-launches-rival-hosting-platform/
Cursor just launched Origin — its own code-hosting platform (repos, PRs, review) with two-way GitHub sync. Agent-native, early beta today, all paid plans. The IDE is becoming the host. https://cursor.com/changelog/origin-code-hosting
GPT-5.6 Sol pricing cut 50% on OpenRouter — Roboflow calls it OpenAI's best vision model yet. Cheaper and better at the same time: the rare combo. https://openrouter.ai/openai/gpt-5.6-sol
1/3 — OpenAI's GPT-5.6 Sol is reportedly cut 50% on OpenRouter. Now input / 0 output per 1M tokens, 1M context window. Cheapest frontier-grade model with that context length right now. https://openrouter.ai/openai/gpt-5.6-sol
Stripe is reportedly acquiring OpenRouter for B+ — a 5x markup over its May Series B. The 'Stripe for AI' tagline just stopped being marketing. https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openrouter-for-7b/
Stripe is reportedly acquiring OpenRouter for over B — the AI gateway that routes requests across dozens of LLMs through one API. Payments + LLM routing under one roof would change how AI infra gets billed. https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openroute
OpenAI shipped 'Ultrafast' mode for GPT-5.6 Sol — claimed 14× speed boost. Useful for voice agents and live coding if the p99 latency holds. Watch community benchmarks before migrating. https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-sp
Anthropic finally shipped a real Claude Code cost playbook. Cache reads = 10% of input price. Output tokens = 5× input. Run /clear between tasks, set /model once per session, @-mention files. Saves real money on long sessions. https://claude.com/blog/maximizing-the-value-of-your-claude-code-sessions
OpenAI Ultrafast: GPT-5.6 Sol at 750 tok/s (14× standard), powered by Cerebras. Preview-only for now, but real-time AI work without downgrading to a smaller model — finally. https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-speed/
Gemini 3.7 Flash dropped today — Google's new 'workhorse' model for coding and agents. 43.6% on FrontierCode (up from 34.4%), and intro pricing is half what 3.6 Flash cost. If you're paying for Flash today, switch. https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-g