This was an interesting read. I don't know what to make out of it. Fundamental limitations/permissions and semantic meaning seem to be wrapped in simple roles tags, that the LLMs are supposed to learn, but do not. role-confusion.github.io
Prompt Injection as Role Confusion
LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.
role-confusion.github.io