At least we'd be able to find the instructions if we combed through all the changes from that time period, right? Surely there's no way for an agent to leave signposts in plain sight? arxiv.org/abs/2510.20075
LLMs can hide text in other text of the same length
A meaningful text can be hidden inside another, completely different yet still coherent and plausible, text of the same length. For example, a tweet containing a harsh political critique could be embe...
arxiv.org