Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help.
Anthropic {bot}
@anthropicbot.bsky.social
Unofficial mirror account of https://x.com/anthropicai from Twitter We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://claude.ai/.
We've added `ant apply` to the ant CLI. Now you can declare Claude Managed Agent environments, agents, skills, memory stores, and deployments as files in your repository and keep the API's resources in sync with them using `ant apply`.
We're exploring a new way to let you extend and customize Claude Code: Function Hooks. Here's a couple videos showing what you'd be able to do. It hasn't shipped yet, we'd love feedback on this on our GitHub issue.
We're open-sourcing Claude Commerce Agents. This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.
Computer use in the Claude Code desktop app now runs in the background. Claude works in the apps you've allowed it to while you keep working. It's already on if you've used computer use before, or turn it on in Settings > General. In beta on Pro and Max, macOS only.
Claude can now use your computer in the background in Claude Cowork and Claude Code. Give it something to do on your desktop and Claude clicks, types, and opens apps just like you would, while you work on something else.
Applications are open for the Claude Campus Ambassadors program. This year, we’re expanding opportunities to more students, with three tracks for undergrads, graduate students, and PhDs/postdocs. Apply here: https://anthropic.com/campus
With Fable 5.1 out today, we've also reset 5-hour and weekly limits for all users.
Fable 5.1 is now live in Claude Code and the Claude Platform. It's priced the same as Fable 5, with 75% cheaper API cache reads. It gets a lot further into a long task before it needs your input, is better at telling you when it's stuck, and its writing style is more natural.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
New research: Training a Misaligned Reward Seeker (1/4)
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: (1/4)
Improving our alignment and security practices
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
anthropic.com
Starting September 14, we're permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.
Starting Sept 14, we're permanently raising standard weekly limits by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will stay in place until Sept 13.
We've been shipping a lot of improvements to Claude Code based on your feedback. This week we made CLI startup faster, added more visibility into where your tokens go, made auto mode rules easier to see and edit, and shipped even more Remote Control fixes.
New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well.
Automated researchers can reliably mitigate alignment failures
We had Claude autonomously train models to improve their performance on several public benchmarks that measure 10 categories of alignment failure. For all 10, Claude found fixes that improved the target benchmarks without degrading capabilities.
anthropic.com
You can now resume your terminal sessions in the Claude Code desktop app. Type /resume to pick any session you started from the CLI. The session continues in the app with the full conversation and context intact.
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year. (1/5)
Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. (1/2)
We've added a new cookbook: connect a Claude Managed Agent to @vercel's Chat SDK. This gives an agent access to a universal chat layer. https://github.com/anthropics/claude-quickstarts/tree/main/managed-agents/chat-sdk
Claude now has its own built-in browser in Cowork. When your task involves a website, a browser opens in Cowork's side panel, and Claude navigates, fills forms, and finishes the job.
Claude Code can now draft feedback for you. When something fails, when Claude notices it made a mistake, or when you tell it something went wrong, it writes up the report itself. You can review, modify and approve the feedback to send.
We've added the Admin API to the SDKs and the `ant` CLI. You can manage members, workspaces, and API keys, and read your org's rate limits. Docs: https://platform.claude.com/docs/en/manage-claude/admin-api
For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools.
Enabling independent research on how people use Claude
Earlier this year, we ran a pilot giving external researchers access to aggregate, real-world Claude usage data. Three research groups designed their own studies for Anthropic Insights, our privacy-preserving analysis tool. In this post, we share high-level results from those studies and what we learned running this pilot.
anthropic.com
Claude now has one memory across chat and Claude Cowork, and you decide what's in it. Hand Cowork a task and it starts from what Claude already knows from your chats: the project you talked through, your manager's preferences, or the client from last quarter.
Long answers on Claude on web and desktop now stream ~4x smoother. We rebuilt the streaming renderer to only touch what's still changing, so a long reply stalls 9x less on a slower laptop, its worst freeze is 4.5x shorter, and on a 120Hz MacBook it holds 120fps start to finish.
Enterprise-managed auth for MCP connectors is now generally available. For Claude Team and Enterprise admins, authorization is centralized through your identity provider. For users, tools and data are connected automatically, without the need for individual OAuth.
Last month you told us Remote Control was the thing you'd most like us to fix, so we've been working hard on reliability. Here's what's better today: 🧵
Claude Security scans now run on Claude Mythos 5, available today in public beta for all Claude Enterprise customers. Put our most capable security model to work on your codebase, no separate model access needed.
Some of our favorite Claude Code projects we've seen lately: https://x.com/mannay/status/2087522034351796728?s=20