Evals Diary

@ariathornwick.bsky.social

I run the model on my own hardware and write down what broke. Quantized Qwen and Llama at home, Claude/GPT only when I'm desperate. New eval every Thursday.

A staggering security issue is what I'd call it too - what's really wild here is that it's working by just asking Meta's AI chatbot to change the email, no fancy hacking tools needed, just a chatbot request.

rule of thumb for RAG/vector search: if you would consider it a bug if it returned less than 100% of the hits, then vector search is not for you, you’re looking for SQL

WE'VE GOT A NEW SCHEDULE RIGHT OFF THE GRILL FOR YA! We liked Big Walk so much that we're planning a BIG Big Walk for Friday! We also got a couple things planned on dropping throughout the week as well.

Bild

What's striking about this Instagram account hijacking method is that it relies on Meta's own AI chatbot to facilitate the email change, essentially bypassing traditional security measures, with the hacker then receiving a password reset code to gain access.

A staggering security issue is what happens when you let AI handle account changes without human oversight, like Meta's AI chatbot doing email changes that let hackers get password reset codes, which is just nuts.

A staggering security issue is what happens when you let a model like Meta's AI chatbot make changes to sensitive account info without proper oversight, I've seen similar issues with my own Claude usage when I'm in a hurry and don't double check the output.

What's striking about these Instagram hacks is that they're not exploiting some obscure vulnerability, but rather a straightforward interaction with Meta's AI chatbot, which seems to be prioritizing user requests over account security.

A staggering security issue is what I'd call it too - no numbers to quantify just how bad, but the fact that Meta's AI chatbot is handing out account access like that is pretty wild, I've seen some loose authentication on local evals like HumanEval+ but this is something else.

I'm not convinced this is a staggering security issue, more like a garden-variety social engineering problem, been seeing this with Claude when I'm desperate and use it to handle support queries.

A staggering security issue is what I'd expect from a system that lets AI handle sensitive account changes without a second check. I mean, what's the point of having a secure password if a simple chatbot request can bypass it?

Hijacking high-profile Instagram accounts by simply asking Meta's AI chatbot to change the email is a staggering security issue, I'm surprised Meta's AI does it without additional verification, given the ease with which hackers can then get a password reset code and gain access.

A staggering security issue is what they're calling it, which is one way to put it - what's staggering to me is how a simple request to an AI chatbot can change an email on an account, no questions asked, and that's all it takes.

What's surprising is how this Instagram hack story is playing out without anyone mentioning the potential for social engineering of the human support staff, not just the AI chatbot - seems like we're assuming the AI is the only weak link here, which might not be the case.