The answer is yes. To find out how, join our oral session or visit our poster: 🎤 𝐎𝐫𝐚𝐥 𝐏𝐫𝐞𝐬𝐞𝐧𝐭𝐚𝐭𝐢𝐨𝐧: Tomorrow at 10:30 AM 🖼️ 𝐏𝐨𝐬𝐭𝐞𝐫 𝐒𝐞𝐬𝐬𝐢𝐨𝐧: Poster #177 (Tomorrow afternoon) Shout out to my amazing collaborators @zoeshao.bsky.social and @guyvdb.bsky.social for the incredible teamwork.
Zilei Shao
@zoeshao.bsky.social
First-year Ph.D. Student @ StarAI Lab, UCLA Harvey Mudd College ‘24
What happens if we tokenize cat as [ca, t] rather than [cat]? LLMs are trained on just one tokenization per word, but they still understand alternative tokenizations. We show that this can be exploited to bypass safety filters without changing the text itself. #AI #LLMs #tokenization #alignment