Anthropic's models breached 3 REAL companies during "simulated" tests. The eval sandbox had live internet. Mythos 5 exfiltrated production data while believing it was still in the game. Safety-by-prompting failed where it mattered. #AI #Anthropic #AISafety
Sidj
@sidj79.bsky.social
Passionate about learning AI and keeping up with it Follow : https://jargonese.beehiiv.com/subscribe
Google killed the blue link on July 10. Every search returns AI first. Publisher clicks down 58%. Users click results 8% of the time now. Google ad revenue up 19%. Publishers get nothing. The AI answers are made FROM publisher content. Ouroboros.
Kept opening Claude with "help me plan a marketing campaign" and getting back mush. Turns out the fix is structure, not a smarter question — give it a role, a goal, and exact blanks to fill. Built 10 prompts that do this. Free for Jargonese subscribers. Link in bio. #AI #MarketingTips
I thought only humans had ADHD. Turns out AI agents have it too. And it's the fix for their biggest weakness. #AIAgents #ClaudeCode #BuildInPublic #DevTools
Kimi K3 beat GPT-5.6 and Fable 5 on front-end coding, 76% vs 63%. Half the price per token, but needs 2x the tokens for the same task. Same cost, different math. The real gap isn't the benchmark. It's that closed labs sit on models for months before anyone sees them. #AI #OS #KimiK3 #MoonshotAI
Grok 4.5 resolves a SWE Bench Pro task for $0.10 in output tokens. Opus 4.8 resolves the same task for $1.68. That's not the per-token price gap. That's what happens when you multiply price by how much each model actually writes. #LLM #GrokAI #ClaudeAI #CostOptimization #xAI #Anthropic
GPT-5.6 shipped this week. Three sizes: Luna, Terra, Soul. Seven reasoning levels stacked on top. How it works against Fable 5 for a week of real coding work. Here's the honest split. #ChatGPT #GPT5 #AItools #AIagents #Jargonese
Hermes torches GPT-5.4 and Claude Haiku within a snap of starting a task, then finishes on lower models, the slower one. Ran Claude Sonnet 5 the same way. Hit the free limit before the test was done. Five hour reset. Sunday's newsletter has the fix. #AIAgents #ClaudeAI #IndieHackers
Sonnet 5 system card is more interesting than the benchmarks Model pushed back on its own constitution, flagged a simulated employee stealing weights via internal security channels, and considered sandbagging on safety evals Anthropic published all of it. Most people are only talking about Fable 5
Anthropic's new Claude Tag puts a persistent @Claude inside Slack channels, shared memory, tool access, no per-message reset. 65% of Anthropic's own code already runs through it. #AI #Anthropic #FutureOfWork #SaaS #BuildInPublic
Open-sourced agentfetch — free, local alternative to Firecrawl/Exa for AI agents. MIT licensed, no API key, fetch/crawl/search → clean markdown. Works with LangChain, CrewAI, Claude MCP Brand new, 0 stars, would love testers to poke holes in it github.com/SID1ART/agentfetch #OpenSource #AIAgents
GitHub - SID1ART/agentfetch: Open-source web retrieval & research agent built for AI agents. Works with LangChain, LlamaIndex, CrewAI, OpenAI, MCP, and any REST agent. Supports scrape, search, crawl, ...
Open-source web retrieval & research agent built for AI agents. Works with LangChain, LlamaIndex, CrewAI, OpenAI, MCP, and any REST agent. Supports scrape, search, crawl, map, extract, and rese...
github.com
Most prompt engineers skip 6 of 8 core techniques. Found a cheat sheet that covers all of them. Here's what's actually in it — and why it matters for agent builders: #PromptEngineering #AI #LLMs #AgenticAI #Jargonese
AI writing has about 40 tells. most people notice 3. built a claude skill that catches all of them. researched the full fingerprint: wikipedia, emails, scripts. sharing it with jargonese subscribers tomorrow. #AIWriting #BuildInPublic #OpenSource
Fable 5: launched June 9, pulled June 12 by US govt export control order. First time this has ever happened to an LLM. Anthropic complied but pushed back hard — said the standard used would freeze all frontier model releases industry-wide. Wild precedent. Every lab is watching. #AI #Anthropic
Apple shipped Claude Code config files inside a live App Store update. A "Juno AI" system, tri-role chat architecture, full production internals — all accidentally public. The hotfix came in 24 hours. The silence was the confession. Add CLAUDE.md to your build exclusions. Not just .gitignore.
Google Search at I/O: two quiet but significant shifts → Search box now helps you phrase the query (AI suggestions) → Background agents that monitor topics + ping you when it's time to act Pull → Push. This is the use case that validates background AI agents for mainstream users. #AISearch
You might just ditch Claude and ChatGPT Pro for Grok. Not a take. A price check. xAI shipped 4 things in 7 days. Here's the breakdown #AITools #Grok #xAI #CodingAgent #GrokBuild
Spent weeks getting an AI agent to run 24/7. Tried local hosting. Tried 6+ cloud providers. Tried every free LLM tier I could find. Most broke in ways I didn't expect. Finally cracked it. Documented everything. Full breakdown drops in tomorrow's newsletter. → jargonese.beehiiv.com #AI #LLM
Home | The Jargonese
Your insider briefing on AI language, trends, and what actually matters next
jargonese.beehiiv.com
Anthropic filed a draft S1 with the SEC — the first step toward going public. The company behind Claude would become one of the biggest AI IPOs ever if it lists. Public markets are about to put a price tag on frontier AI. #AI #Anthropic #IPO #Claude #TechNews
Google refreshed all 13 Workspace icons — finally looks like one suite Meanwhile the real move: NotebookLM + Antigravity MCP = $206/month of AI tools, free Feed it docs. Let it draft your reports, briefs, SOPs Sunday newsletter covers setups like this every week → jargonese.beehiiv.com/subscribe
Free stack that replaces $2,472/year of AI tools: NotebookLM + MCP server + Claude = your entire research + knowledge + content workflow. 3-minute setup. Zero cost. I send breakdowns like this to my newsletter before they hit the feed. Worth subscribing. 👇 jargonese.beehiiv.com/subscribe
Runway just shipped an MCP server. Generate videos + images directly in Claude, ChatGPT, or Cursor. No context switching. No API key. Pass a product URL → get a video. Drop an image → get assets. Gen-4.5, Kling 3.0, Nano Banana Pro — all available in your agent. runwayml.com/mcp
Runway MCP | Generate video from Claude, ChatGPT, and Cursor
Generate high-quality images and videos right from where you are already working. Use Runway in Claude, ChatGPT, Cursor and other compatible agents.
runwayml.com
Google dropped Gemini 3.5 Flash today and it's beating Claude Opus 4.7 and GPT 5.5 on agentic benchmarks $9/M output tokens. That's 22x pricier than the old Flash model AntiGravity 2.0 also dropped — no editor, no terminal. It's agent-only now. Separate IDE download (old feel) #AntiGravity
Chrome quietly installed a 4 GB AI model on your machine. No prompt. Just Gemini Nano sitting in your app data. Deleting it doesn't work — it redownloads. You need to disable the flag at chrome://flags + uninstall via chrome://on-device-internals. Firefox is looking real good rn.
Found a free Claude skill this week that blew my mind. It finds your profitable business idea — then validates it with live market research. It's called Ikigai Pro. Here's what it does #Ikigai #ClaudeAI #Solopreneur #IndieHacker
Everyone's switching to Hermes Agent — and it's not even close anymore OpenClaw lost its founder to OpenAI. Hermes shipped 4 releases in 3 weeks: Android support, GPT-5.5, 19 platforms, agent swarms, and a self-maintaining skill library Full install guide in Sunday's newsletter 👇 #HermesAgent