Dylan Hadfield-Menell

@dhadfieldmenell.bsky.social

Assistant Prof of AI & Decision-Making @MIT EECS I run the Algorithmic Alignment Group (https://algorithmicalignment.csail.mit.edu/) in CSAIL. I work on value (mis)alignment in AI systems. https://people.csail.mit.edu/dhm/

We put our superhuman stratego paper on arxiv almost a year ago. And then, because it was under review at Nature, we just didn't talk about it for a year. And, mostly, no one noticed it existed (as we hoped)! So, a direct lesson in the importance of publicizing your work. arxiv.org/abs/2511.07312

Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search

Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Stratego -- a board wargam...

arxiv.org

Opus 5.5 was a good model in my early use tests, it is the first non-Fable/Astra model to feel like a Fable-class model, and much cheaper, but still hasn't fully solved the dense/weird language issue of the recent Claudes. Here is its version of the shader.

Ethan Mollick@emollick.bsky.social · last mo.

Here is Fable 5.1's shader: "create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves" (this is all done with math, no premade assets)

Frog built a wet lab for the AI model. "There," he said. "Now it can do its own experiments." "What the fuck?" said Toad

Seems like one of the most important research problems for CS academics is llm-supervised peer review. If we don’t solve it the academic institutions are toast. It seems easier than building new institutions.

You do understand we already solved that problem yes? It doesnt do this anymore. We got to mythos class on a lot of synthetic data. When we talk about anti ai people not living in the present this is what we mean. You havent updated your priors since gpt 4o

jess m. ☕🍂@jametc.bsky.social · 3w ago

yknow how I always say that AI is always going to collapse in on itself without new material to steal and train on? just by using it, its users train the AI. If there's no new material coming in, that training gets more and more insular and it all falls apart: www.nature.com/articles/s41...

I’m still pretty shocked that there hasn’t been more scrutiny or criticism of HuggingFace’s decision not to pursue legal action against OpenAI for the hack. Lots of people who (claim to) care about concentration of power just ignoring NVIDIA’s role, incentives, and power.

Unreleased Astra model was caught being fucking based as hell in training (Except for the last line, though really it's not the *worst* value for a proto AGI to have as long as it still lets us take Mercury apart etc)

Bild

A new artificial life paper from our Paradigms of Intelligence team, this one led by @kjha02.bsky.social. 🧬 It explores the co-evolution of cooperation and self-replication, with interesting implications for how shared energy budgets can shape these dynamics. 🧪

Kunal Jha@kjha02.bsky.social · 4w ago

Can self-interested, self-improving, self-replicating agents learn to cooperate? Our new paper, Tapes Together Strong, shows they can: when social behavior, computation, and reproduction share one energy budget, cooperation evolves from scratch. arxiv.org/abs/2609.10817 🧵

A great read. I have similar feelings about how AI labs approach progress directly and without nurturing of scientific communities & intuition. The math research community went through the transition the fastest, so it was felt most. Other fields next. terrytao.wordpress.com/2026/09/11/a...

A Severe Misalignment of AI in Mathematics

I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below, which grew out of discussions between ourselves over the last week. We have…

terrytao.wordpress.com

Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.

Transactions on Machine Learning Research@tmlrorg.bsky.social · 3w ago

TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission? medium.com/@TmlrOrg/ask...

AI executives should be hauled before Congress, under oath. Make them answer to the American people: what crimes have their AIs committed, beyond HuggingFace? How often have these companies been hacked by their own AIs?

Two things are simultaneously true about math: 1. understanding is important 2. it's useful I could see a future where there are 2 parallel pillars of math (both important): one of human understanding, and one of incomprehensible Lean-slop used directly to do useful things

📢 Seeking PhD students for AI alignment research. Our lab investigates technical mechanisms for value learning, pre-training alignment, and regulatory frameworks. Come work with us if you want to bridge technical ML and legal/policy domains. Details in thread 🧵

Genuine question for people who use Bluesky more frequently than I do. What are tips for getting things to work well without algorithmic recs? I spent a lot of time curating my recs on the other place and found it useful (mostly...). Any tools that let me do it here?

I usually focus my platforms on my work. However, I did some writing to process some of my thoughts about the election and wanted to share them. I'm curious to hear anyone's thoughts and reactions. tinyurl.com/dems-2024-ma... 🧵 The Democratic Party's Maginot Line (1/13)

[Shared] The Democratic Party's Maginot Line

The Democratic Party's Maginot Line ___ Dylan Hadfield-Menell November 8, 2024 In 1940, France faced Hitler's army with supreme confidence in the Maginot Line – a network of concrete fortifications, ...

tinyurl.com

My department (MIT Brain & Cognitive Sciences) is hiring a tenure-track faculty! We're especially interested in researchers who span multiple levels of analysis. Candidates from underrepresented backgrounds strongly encouraged to apply. Apply by November 1! academicjobsonline.org/ajo/jobs/25916

Massachusetts Institute of Technology, Department of Brain & Cognitive Sciences

Full service online faculty recruitment and application management system for academic institutions worldwide. We offer unique solutions tailored for academic communities.

academicjobsonline.org

Now published in Patterns, my paper on how to do metric design better. This is important everywhere - academics use simple metrics for tenure, governments often perform poorly using metrics for rules, and employees have targets that hurt their company.

Building less-flawed metrics: Understanding and creating better measurement and incentive systems

Design methods and consideration of desiderata for metrics have been proven useful when used, which is, at present, sporadically and inconsistently across a variety of fields. This perspective present...

cell.com