Natasha Johnson

@natashamarie330.bsky.social

CS PhD student at Northeastern University studying NLP + Cultural Analytics

Digital humanities researchers often care about fine-grained similarity based on narrative elements like plot or tone, which don’t necessarily correlate with surface-level textual features. Can embedding models capture this? We study this in the context of fanfiction!

Figure showing a similarity comparison between three stories. Story A and story B have the same author, and story A and story C have the same tone. A human might care about which stories are tonally the most similar, but a language model's notion of similarity is strongly informed by surface-level features like small differences in writing style across authors.