Javier Rando

@javirandor.com

Red-Teaming LLMs / PhD student at ETH Zurich / Prev. research intern at Meta / People call me Javi / Vegan 🌱 Website: javirando.com

Adversarial ML research is evolving, but not necessarily for the better. In our new paper, we argue that LLMs have made problems harder to solve, and even tougher to evaluate. Here’s why another decade of work might still leave us without meaningful progress. 👇

Bild

Recent LLM forecasters are getting better at predicting the future. But there's a challenge: How can we evaluate and compare AI forecasters without waiting years to see which predictions were right? (1/11)

Bild

A new competition on LLM-agents prompt injection is out! Send malicious emails and get agents to perform unauthorised actions. The competition is hosted at SaTML 2025 and has a pool of $10k in prizes! What are you waiting for?

Santiago Zanella-Beguelin@xefffffff.bsky.social · 2y ago

📢Have experience jailbreaking LLMs? Want to learn how an indirect / cross prompt injection attack works? Want to try something different to an advent of code? Then, I have a challenge for you! The LLMail-Inject competition (llmailinject.azurewebsites.net) starts at 11am UTC (that's in 5min!)

I will be at #NeurIPS2024 in Vancouver. I am excited to meet people working on AI Safety and Security. Drop a DM if you want to meet. I will be presenting two (spotlight!) works. Come say hi to our posters.

🚨Unlearned hazardous knowledge can be retrieved from LLMs 🚨 Our results show that current unlearning methods for AI safety only obfuscate dangerous knowledge, just like standard safety training. Here's what we found👇

Bild

Zurich is a great place to live and do research. It became a slightly better one overnight! Excited to see OAI opening an office here with such a great starting team 🎉

Alexander Kolesnikov@handle.invalid · 2y ago

Ok, it is yesterdays news already, but good night sleep is important. After 7 amazing years at Google Brain/DM, I am joining OpenAI. Together with @xzhai.bsky.social and @giffmana.ai, we will establish OpenAI Zurich office. Proud of our past work and looking forward to the future.

Anyone may be able to compromise LLMs with malicious content posted online. With just a small amount of data, adversaries can backdoor chatbots to become unusable for RAG, or bias their outputs towards specific beliefs. Check our latest work! 👇🧵

Ensemble Everything Everywhere is a defense against adversarial examples that people got quite exited about a few months ago (in particular, the defense causes "perceptually aligned" gradients just like adversarial training) Unfortunately, we show it's not robust... arxiv.org/abs/2411.14834

Gradient Masking All-at-Once: Ensemble Everything Everywhere Is Not Robust

Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations...

arxiv.org