Ben Edelman

@benedelman.bsky.social

Thinking about how/why AI works/doesn't, and how to make it go well for us. Currently: AI Agent Security @ US AI Safety Institute benjaminedelman.com

1/ Excited to share a new blog post from the U.S. AI Safety Institute! AI agents are becoming more capable, but they are vulnerable to prompt injections in external content – an agent may be given task A, but then be “hijacked” and perform malicious task B instead. www.nist.gov/news-events/...

Technical Blog: Strengthening AI Agent Hijacking Evaluations

Large AI models are increasingly used to power agentic systems, or “agents,” which can automate complex tasks on behalf of users.

nist.gov

0/ I'd like to kick off my presence here with a question: why does learning work in practice? Why is the world such that we can we learn to predict things from other things in a computationally efficient way; why is "simplicity bias" empirically useful? Some explanations: