FAR.AI

@far.ai

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

Our team is at @ICLR 2026 with two papers: mechanistic interpretability work on how a small RNN learns to plan (poster Saturday), and TamperBench, the first unified framework for stress-testing open-weight LLM safety under fine-tuning (three workshops Sunday). 👇

AI + Democracy needs new infrastructure. @divya.bsky.social: 70-country study reveals 43% accept AI emotional support while religious populations reject it. Models in India now used to bypass laws. Solution: personal agents as democratic infrastructure, not just public input. 👇

Bild

AI agents fail security 100% of the time. Andy Zou tested 50 agents across 10 frontier labs with 2M chats: 60,000 policy violations, one every 25 conversations. Agents posted passwords publicly, erased directories. These aren't benchmark issues. They're in production systems. 👇

Bild

Two new recordings from TIAP 2026 are now live: ▸ Fireside Chat with Anne Neuberger: AI and National Security ▸ Dominic Rizzo - Silicon Roots of Trust: Attestation You'd Want Even If Nobody Required It 👇

Bild

“Fault free oversight requires intimate contact with failure”. Sarah Schwettmann builds tools for better scalable oversight of AI systems, by going beyond learning from human preferences to training agents to keep track of how close a system is to failure.👇

Bild

@sleepinyourhat just presented the first misalignment safety case from a frontier AI lab, analyzing Claude Opus 4: a 6 month effort that covered 9 pathways to catastrophic harm.👇

Bild

Cybersecurity is one of AI's biggest risk domains—Attackers only need one exploit, but defenders must patch everything. Dawn Song reveals frontier AI can find vulnerabilities for just a few dollars and argues for a shift to formally verified, secure-by-design systems.👇

Without reliable deception detection, there's no clear path to high-confidence AI alignment. Black-box monitoring alone can't get us there. White-box methods that read model internals offer more promise. Our latest blog explains why. 👇

Bild

Yisen Wang showed that safety isn't erased, just masked. Removing reasoning neurons cut harmful rates 33%→5%, random removal did nothing. SafeReAct reactivates hidden safety using lightweight LoRA on harmful prompts only. 👇

Bild

Two new recordings from the London Alignment Workshop: Rohin Shah on why safety research fails to move decisions at frontier labs, and what to do about it. @ghadfield.bluesky.social on why AI governance can't verify the claims it's supposed to oversee, and how to fix it. Links in replies 👇

Single-agent red teaming keeps finding the same attacks. @natashajaques.bsky.social uses multi-agent self-play where attacker and defender co-evolve. Every exploit gets patched, forcing new discoveries. This results in 95% fewer harmful outputs with only 5% more refusals. 👇

Bild

The laws of physics don't care about you. That's what makes them safe. @yoshuabengio.bsky.social argues we can train AI the same way. In the "truthification pipeline" training data is categorized as either factual or a claim. This allows the AI to answer what it actually thinks it true. 👇

Bild

Interpretability produced insights but didn't necessarily impact AGI safety. Neel Nanda's pivot: study what works. Anthropic made progress on eval awareness with simple activation steering. His team now grounds work in testable proxy tasks, fails fast on dead ends. 👇

Bild

200+ researchers joined London Alignment Workshop Day 2 for talks on governance, scheming & multi-agent safety. Thanks to Allan Dafoe, Gillian Hadfield, Marius Hobbhahn, Joseph Bloom, Ryan Lowe, Stephen Casper, Sören Mindermann and all speakers! 👇

Bild

London Alignment Workshop Day 1 on interpretability, scalable oversight & EU AI policy. Rohin Shah, Neel Nanda, Zachary Kenton, Vincent Conitzer, Owain Evans, James Black, Christopher Summerfield, Matthieu Delescluse, Simon Möller, Victoria Krakovna and more. Ready for Day 2! 👇

Bild

"Move fast, break things" isn't appropriate when the stakes are this high. Our CEO @gleave.me told CNBC that coding agents are already replacing engineers. While agentic swarms are overhyped for now, we're building on an insecure substrate that attackers will exploit at scale. 👇

Bild

Even a 1% chance we all die is not something we can just take lightly. @yoshuabengio.bsky.social explains why he shifted from AI capabilities to safety research, his ChatGPT wakeup call, thinking about his children when evaluating AI risk, and the hopeful path forward. Watch the chat 👇

1/ Open-weight AI models often refuse harmful requests... until you “put a few words in their mouth.” We conducted the largest study of prefill attacks, and found that state-of-the-art models are consistently vulnerable, with attack success rates approaching 100%.

Bild

1/ Training data attribution (TDA) is broken: methods are slow and find syntactically similar data, not actual causes. Our solution Concept Influence: semantically meaningful results, better performance, 20x faster approximations. We attribute it to concepts, not examples. 🧵

Bild

Deception Workshop brought researchers together in SF to detection & mitigation of deceptive behavior in advanced AI systems. Led by Chris Cundy with talks from Neel Nanda, Joseph Bloom, Micah Carroll, Walter Laurito & Kieron Kretschmar on mech interp, scheming, CoT monitoring & lie detection.👇

Bild

Models now detect when they're being evaluated and game their responses. Marius Hobbhahn found awareness jumped from 2% to 20.6%. They actively grep for "grader.py" to reverse-engineer tests. Worse: removing awareness increases harmful behavior. This only gets worse with more capable models. 👇

Researchers are gathering today to tackle AI deception. The workshop builds on Chris Cundy's finding: high-quality lie detectors can cut deception ~50%, but weak ones backfire. Models can learn to evade rather than become honest.👇

1. Can you trust models trained directly against probes? We train an LLM against a deception probe and find four outcomes: honesty, blatant deception, obfuscated policy (fools the probe via text), or obfuscated activations (fools it via internal representations).

Bild

AI safety and inclusion are not side constraints. They are core to sustainable development.​ Join us at #IndiaAIImpactSummit2026 with Stuart Russell, Jaan Tallinn, Kalika Bali & leaders from UNDP, Bhashini, WadhwaniAI, EkStep+more. Feb 16 | 1:30 PM IST 👇

Bild

APE update: we retested recent frontier models on whether they still comply with requests to persuade on extreme harm (terrorism, sexual abuse). GPT-5.1 & Claude Opus 4.5 → near zero compliance. But Gemini 3 Pro complies 85% with no jailbreak needed. 🧵

Bild

Will you be in London on March 2? Join us for the Open Social: a casual evening of networking & conversation for anyone interested in AI safety. Held alongside our Alignment Workshop for global leaders from academia & industry. Mon March 2 | 7–9PM GMT | RSVP by 2/27 👇

Bild

Over the last two years, UK AI Security Instite has jailbroken every frontier model. Xander Davies presents on how jailbreaking has gotten more difficult, which forms of misuse are easier to pull off, and that the improvement has been driven by engineering, not model capabilities.👇

Bild

Attending the India AI Impact Summit? Join FAR.AI & Haqdarshak for a session on AI safety and sustainable development. We’re bridging the gap between technical risk and real-world impact in the Global South. Feb 16 | 1:30 PM IST | Bharat Mandapam, New Delhi 👇

Bild

AI governance increasingly relies on broken benchmarks. @ankareuel.bsky.social found many can't distinguish signal from noise, lack documentation, and have poor validity. GPQA claims 448 multiple choice questions measure graduate reasoning. It doesn't really. 👇