Fazl Barez

@fbarez.bsky.social

Let's build AI's we can trust! https://fazlbarez.com

A glimpse into another successful Oxford Connected Life Summit, focused on what it means to live in an increasingly connected world. This year's theme was "New Intelligence, Old Questions," and featured notable speakers and organisations. Huge thanks to the student committee for making this happen!

Heading to #ICML2026 in Seoul next week with the TSG Lab and Martian 🇰🇷 10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights Giving an invited talk at the EIML WS Grateful to the students, collaborators, mentors and Claude who made it happen!

Bild

Excited to be debating at the Oxford Union this evening Motion: This House Believes that AI is the Great Equalizer Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point! We'll find out which way the House votes

Bild

In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵

Bild

How can we ensure AI-powered robots remain safe when operating in the real world? 🤖 A recent article co-authored by @aigioxfordmartin.bsky.social researcher Fazl Barez, explores safety challenges and the importance of being context-aware 🔗Read the research in full: www.science.org/doi/10.1126/...

Beyond alignment: Why robotic foundation models need context-aware safety

Because AI-enabled robots can be tricked into taking unsafe actions, they require layered, context-aware safety guardrails.

science.org

Incredibly excited to announce $1 Million prize pool to solve the world’s most important scientific problem in Interpretability. The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.

Evaluating the Infinite 🧵 My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities. Let me explain a bit about what causes the problem and how my solution avoids it. 1/N arxiv.org/abs/2509.19389

Evaluating the Infinite

I present a novel mathematical technique for dealing with the infinities arising from divergent sums and integrals. It assigns them fine-grained infinite values from the set of hyperreal numbers in a ...

arxiv.org

🚀 Excited to have 2 papers accepted at #NeurIP2025! 🎉 congrats to my amazing co-authors! More details (and more bragging) soon! and maybe even more news on sep 25 👀 See you all in… Mexico? San Diego? Copenhagen? Who knows! 🌍✈️

Other works have highlighted that CoTs ≠ explainability alphaxiv.org/abs/2025.02 (@fbarez.bsky.social), and that intermediate (CoT) tokens ≠ reasoning traces arxiv.org/abs/2504.09762 (@rao2z.bsky.social). Here, FUR offers a fine-grained test if LMs latently used information from CoTs for answers!

Chain-of-Thought Is Not Explainability | alphaXiv

View 3 comments: There should be a balance of both subjective and observable methodologies. Adhering to just one is a fools errand.

alphaxiv.org

Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵

Bild

Technology = power. AI is reshaping power — fast. Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight. Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.

Bild

Come work with me at Oxford this summer! Paid research opportunity to: White-box LLMs & model security Safe RL & reward hacking Interpretability & governance tools Remote or Oxford. Apply by 30 May 23:59 UTC. DM with questions.

Come work with me at Oxford! We’re hiring a Postdoc in Causal Systems Modelling to: - Build causal & white-box models that make frontier AI safer and more transparent - Turn technical insights into safety cases, policy briefs, and governance tools ] DM if you have any questions.

First-time Area Chair seeking advice! What helped you most when evaluating papers beyond just averaging scores? After suffering through unhelpful reviews as an author, I want to do right by papers in my track.

Technical AI Governance (TAIG) at #ICML2025 this July in Vancouver! Credit to Ben and Lisa for all the work! We have a new centre at Oxford working on technical AI governance with Robert Trager and @maosbot.bsky.social many other great minds. We are hiring - please reach out! Quote

Technical AI Governance @ ICML 2025@taig-icml.bsky.social · last yr.

📣We’re thrilled to announce the first workshop on Technical AI Governance (TAIG) at #ICML2025 this July in Vancouver! Join us (& this stellar list of speakers) in bringing together technical & policy experts to shape the future of AI governance! www.taig-icml.com

New paper alert! Curious how small prompt tweaks impact LLM accuracy but don’t want to run endless inferences? We got you. Meet DOVE - a dataset built to uncover these sensitivities. Use DOVE for your analysis or contribute samples -we're growing and welcome you aboard!

Eliya Habba@eliyahabba.bsky.social · last yr.

Care about LLM evaluation? 🤖 🤔 We bring you ️️🕊️ DOVE a massive (250M!) collection of LLMs outputs  On different prompts, domains, tokens, models... Join our community effort to expand it with YOUR model predictions & become a co-author!

What happens once AI can design better AI, which can itself design better AI? Will we get an "intelligence explosion" where AI capabilities increase very rapidly? Tom Davidson, Rose Hadshar and I have a new paper out with analysis of these dynamics.

1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇

Bild