Sebastian Farquhar

@sebfar.bsky.social

Senior Research Scientist at Google DeepMind. AGI Alignment researcher. Views my dog's.

By default, LLM agents with long action sequences use early steps to undermine your evaluation of later steps; a big alignment risk. Our new paper mitigates this, keeps the ability for long-term planning, and doesnt assume you can detect the undermining strategy. 👇

David Lindner@davidlindner.bsky.social · 2y ago

New Google DeepMind safety paper! LLM agents are coming – how do we stop them finding complex plans to hack the reward? Our method, MONA, prevents many such hacks, *even if* humans are unable to detect them! Inspired by myopic optimization but better performance – details in🧵

Something I loved most about the internet in the 2000s was the idiosyncratic personal webpages that some people had put a crazy amount of time and effort into. These pages must still exist right? What are the best ones you know of?