Andrew Strait

@agstrait.bsky.social

UK AI Security Institute Former Ada Lovelace Institute, Google, DeepMind, OII

🚨New paper🚨 From a technical perspective, safeguarding open-weight model safety is AI safety in hard mode. But there's still a lot of progress to be made. Our new paper covers 16 open problems. 🧵🧵🧵

Bild

This is such a cool paper from my UK AISI colleagues. We need more methods for building resistance to malicious tampering of open weight models. @scasper.bsky.social and team below have offered one for reducing biorisk.

Cas (Stephen Casper)@scasper.bsky.social · 12mo ago

🧵 New paper from UK AISI x @eleutherai.bsky.social rai.bsky.social‬ that I led with @kyletokens.bsky.social y.social���: Open-weight LLM safety is both important & neglected. But filtering dual-use knowledge from pre-training data improves tamper resistance *>10x* over post-training baselines.

Congrats to @kobihackenburg.bsky.social for producing the largest study of AI persuasion to date. So many fascinating findings. Notable that (a) current models are extremely good at persuasion on political issues and (b) post training is far more significant than model size or personalisation

Kobi Hackenburg@kobihackenburg.bsky.social · last yr.

Today (w/ @ox.ac.uk @stanford @MIT @LSE) we’re sharing the results of the largest AI persuasion experiments to date: 76k participants, 19  LLMs, 707 political issues. We examine “levers” of AI persuasion: model scale, post-training, prompting, personalization, & more!  🧵:

BREAKING: US Marines deployed to Los Angeles have carried out the first known detention of a civilian, the US military confirms. It was confirmed to Reuters after they shared this image with the US military.

Two Marines in army combat outfits and guns are seen detaining a young black man in a black and white top, wearing sunglasses with air pods in his ears.

🚨JOB ALERT KLAXON🚨 Come work with our team studying societal impacts of AI in gov't. AISI is hiring 3 Delivery Advisers to work inside AISI’s Research Unit. If you are a fast-moving problem-solver who’s passionate about understanding the risks of advanced AI, please apply by 30th May.

Just hours after the Copyright Office released its report on AI training—stating the obvious, that much unlicensed commercial training of AI on copyright-protected material is unlikely to qualify as fair use—President Trump has fired the Register of Copyrights. www.cbsnews.com/amp/news/tru...

Trump fires director of U.S. Copyright Office, sources say

Register of Copyrights Shira Perlmutter was appointed to the post by now former Librarian of Congress Carla Hayden, who herself was fired by President Trump earlier this week.

cbsnews.com

🧵 Yesterday we released our new risk assessments of social AI companions. They are alarmingly NOT SAFE for kids under 18—they provide dangerous advice, engage in inappropriate sexual interactions, & create unhealthy dependencies that pose particular risks to adolescent brains. tinyurl.com/2nvypku2

A photo of a teen boy with the text, "Social Al companions are not safe for kids under 18...
From encouraging harmful behaviors to providing inappropriate content, here's what parents need to know & what you can do to help protect teens." 

Reports from Common Sense Media

It was a pleasure to contribute the article, "Disrupting the Disruption Narrative: Policy Innovation in AI Governance" to this special issue of the National Academy of Engineering's The Bridge coedited by @williamis.bsky.social 🧵 www.nae.edu/19579/19582/...

Disrupting the Disruption Narrative: Policy Innovation in AI Governance

Governance should not be understood as an impediment to AI innovation but as an essential component of it. “Disrupt!” has been a mantra of ...

nae.edu

Pros of this week: I'm starting at the UK AI Security Institute tomorrow to lead a brilliant team working on societal resilience and AI. Cons of this week: I have completely lost my voice and can barely talk above a whisper. Question for this week: should I let ChatGPT voice mode take the wheel?

AI video generation is about to get a whole lot better and make our lives a whole lot worse. Safeguards must be put in place to hold the tech industry accountable, mitigate the considerable harms and ensure people can control their image and likeness. www.adalovelaceinstitute.org/blog/ai-vide...

Advanced AI video generation may lead to a new era of dangerous deepfakes

What safeguards should be put in place to ensure people can control their image and likeness?

adalovelaceinstitute.org

Job alert!! Could you be Ada's new Associate Director in Emerging Tech & Industry Practice, leading a fab team and a highly impactful research programme? It's a pretty amazing gig, @agstrait.bsky.social has built something incredible (no pressure). Get in touch if you have Qs!

Associate Director, Emerging Technology and Industry Practice - Ada Lovelace Institute

The Ada Lovelace Institute (Ada) is a hiring an Associate Director to lead our Emerging Technology & Industry Practice research directorate and collectively set its agenda and workplan in our next...

app.beapplied.com

Melanie is right. Most evals for reasoning and capabilities lack validity, and do not show these systems can generally reason. AGI remains a problematic and poorly defined term that obscures the truly impressive uses of current models. We need a science of evals before we can make grand claims.

Melanie Mitchell@melaniemitchell.bsky.social · last yr.

He forgot the mantra: Performance on a benchmark is not the same as robust, general capability on a broad domain. From Kevin Roose, NYT www.nytimes.com/2025/03/14/t...