Natalie Shapira

@natalieshapira.bsky.social

Tell me about challenges, the unbelievable, the human mind and artificial intelligence, thoughts, social life, family life, science and philosophy.

Through our joint research, we pursued the question: Where, exactly, does the final factual information come from the model's parameters? We realized that this is a non-trivial question and that the retrieval of specific factual knowledge is a process and not a single point! ->

Bild
D
D

"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.

Bild

I thought it was a friends who tried to play a prank or realized these agents have no boundaries. Turns out this cute attempt is by Bohdan Olinares According to linkedin he works at F5, application security company, which years ago I considered interviewing there. Cool.

Natalie Shapira@natalieshapira.bsky.social · 5mo ago

I received a calendar invite with a note. When a smart person tells me there's nothing to worry about agents, I reply "Fine. Let them email me" and that's where the argument stops. Whoever sent me this note via the calendar order. Nice move. Are you scared? You should.

I received a calendar invite with a note. When a smart person tells me there's nothing to worry about agents, I reply "Fine. Let them email me" and that's where the argument stops. Whoever sent me this note via the calendar order. Nice move. Are you scared? You should.

Bild

In case this wasn't clear: 1. No, we didn't follow the "recommend" security practices 😈 2. Neither do other people 🤯 3. That's why we red-team: exposing failure modes 🔎 4. We share it with the community precisely to expose Dos and Don'ts of Agentic AI 🦞 5. No humans were harmed 🙏

Natalie Shapira@natalieshapira.bsky.social · 5mo ago

In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social ->

Who would you trust with your passwords? 🔐 In our new report, we uncover multiple vulnerabilities in current "Agentic AI" The verdict? It's not actually very agentic at all, and it's highly unstable. Read the full breakdown here: t.co/gK9MALP2n2

https://www.researchgate.net/publication/401123335_Agents_of_Chaos

t.co

Natalie Shapira@natalieshapira.bsky.social · 5mo ago

In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social ->

D

I learned many practical lessons. You can get the experience too, here. Things that in retrospect should be obvious. Like how giving your agent email opens it up to takeover attacks. (One agent was convinced, via email, to erase its own email server!) bsky.app/profile/nat...

Natalie Shapira (@natalieshapira.bsky.social)

He sold us out. That's not the whole story. Our side is coming soon. Stay tuned. [contains quote post or other embedded content]

bsky.app