Yarin

@yaringal.bsky.social

Associate Professor of Machine Learning, University of Oxford; OATML Group Leader; Director of Research at the UK government's AI Safety Institute (formerly UK Taskforce on Frontier AI)

Hot take: I think we just demonstrated the first AI agent computer worm 🤔 When an agent sees a trigger image it's instructed to execute malicious code and then share the image on social media to trigger other users' agents This is a chance to talk about agent security 👇

Lukas Aichberger@aichberger.bsky.social · last yr.

⚠️ Beware: Your AI assistant could be hijacked just by encountering a malicious image online! Our latest research exposes critical security risks in AI assistants. An attacker can hijack them by simply posting an image on social media and waiting for it to be captured. [1/6] 🧵

I'm looking for PhD applicants who have expertise in Gaussian processes and/or Transformers for an exciting PhD project If this sounds interesting, application deadline for funding is 3/12 Please share with people you think this might be relevant to! oatml.cs.ox.ac.uk/apply.html

Bild

I don’t need to go on social media to have my worldview challenged I am in theoretical physics I have a new existential crisis daily

Since this platform is finally attracting a critical mass of ML researchers, here's our recent work on prompt-based vulnerabilities of coding assistants: arxiv.org/abs/2407.11072 TL;DR — An attacker can convince your favorite LLM to suggest vulnerable code with just a minor change to the prompt!

MaPPing Your Model: Assessing the Impact of Adversarial Attacks on LLM-based Programming Assistants

LLM-based programming assistants offer the promise of programming faster but with the risk of introducing more security vulnerabilities. Prior work has studied how LLMs could be maliciously fine-tuned...

arxiv.org

I've created an initial Grumpy Machine Learners starter park. If you think you're grumpy and you "do machine learning", nominate yourself. If you're on the list, but don't think you are grumpy, then take a look in the mirror. go.bsky.app/6ddpivr

Post nicht verfügbar.

I’m keen to dig more into safety cases, there’s something ‘proving a negative’ about them but equally it’s good to see a really concrete attempt to tether speculation. Here’s a new piece from UK AISI @girving.bsky.social and gov AI attempting to provide a template arxiv.org/abs/2411.08088

Safety case template for frontier AI: A cyber inability argument

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offering such assurances is...

arxiv.org