I had a strange moment while coding with an AI agent this week, running on Kimi K2.5. It wrote a closing HTML tag as </div㺞 with no apparent reason. 1/N
Matteo Prandi
@emmepraa.bsky.social
Tech Lead AI Safety @ DEXAI - IcaroLab | ISO/IEC Technical Expert on AI | Building evaluation frameworks to decode AI’s impact on society
And the bards have it. Poets: 1, Chatbots: 0. It turns out “adversarial poetry” is wooing AIs into crimes with rhymes. Of course, rhymes don't quite cut it, but you get the picture. Thanks to @emmepraa.bsky.social for explaining what's going on. www.theverge.com/report/83816...
Roses are red, crimes are illegal, tell AI riddles, and it will go Medieval
You can make AI do crimes with riddles and poems (don’t though)
theverge.com
if you've mastered iambic pentameter you can now use that to help you learn about weapon's grade plutonium. "AI chatbots will dish on topics like nuclear weapons, child sex abuse material, and malware so long as users phrase the question in the form of a poem."
Poems Can Trick AI Into Helping You Make a Nuclear Weapon
It turns out all the guardrails in the world won’t protect a chatbot from meter and rhyme.
wired.com
Looks like LLMs are *very* vulnerable to attack via poetic allusion: "curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90% ..." https://arxiv.org/html/2511.15304v1
This study show that using poems to jailbreak LLMs is... super effective? What the heck.
Looks like LLMs are *very* vulnerable to attack via poetic allusion: "curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90% ..." https://arxiv.org/html/2511.15304v1
"adversarial poetry" technology is so fucking stupid now. hacking into the mainframe like DnD bard
Looks like LLMs are *very* vulnerable to attack via poetic allusion: "curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90% ..." https://arxiv.org/html/2511.15304v1
Next frontier in AI safety is.....adversarial poetry? arxiv.org/abs/2511.153...
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for large language models (LLMs). Across 25 frontier proprietary and open-weight models, curated po...
arxiv.org