1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
Arkadiy Saakyan
@asaakyan.bsky.social
PhD student at Columbia University working on human-AI collaboration, AI creativity and explainability. prev. intern @GoogleDeepMind, @AmazonScience asaakyan.github.io
Excited to share #ICML2026 paper from my internship @ Google DeepMind! AI models are deployed globally, but AI safety datasets are largely geographically homogenous. What is the impact of culture on AI safety ratings? Is there any impact beyond standard demographics like age, gender, and ethnicity?
We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?)
Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access!
IPUMS Spatial Student Award is a tie! @gsagostini.bsky.social for "Inferring Fine-Grained Migration Patterns Across the United States." (www.nature.com/articles/s41...)
We are co-hosting the EAAMO colloquium next Monday (12pm EST) with Professor Rachel Franklin. Come hear her talk about spatial inequality and the smart city and feel free to share with colleagues! Register below to get the Zoom link: www.eaamo.org/colloquium/r...
Excited to travel to ICLR 🇧🇷 to present our work on textual creativity metrics! See me at the 10:30am-1pm Saturday poster session in Pavilion 3 or reach out to chat 🙂
N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity.
🚨Paper on AI & Copyright Courts have credited AI companies' claims that alignment prevents reproducing copyrighted data. What if finetuning on a simple writing task breaks it. Worse: tuning on just one author (e.g., Murakami) unlocks verbatim recall of 30+ other authors' books (up to 90%) (1/n)🧵
⚛️ Introducing CREATE, a benchmark for creative associative reasoning in LLMs. Making novel, meaningful connections is key for scientific & creative works. We objectively measure how well LLMs can do this. 🧵👇
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9
Merriam-Webster’s human editors have chosen ‘slop’ as the 2025 Word of the Year.
Day 7 of #30DayMapChallenge asked us to think about accessibility. @gsagostini.bsky.social considers two metrics of access simultaneously: distance to a Subway and distance to the subway.
N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity.
Are you a researcher using computational methods to understand cities? @mfranchi.bsky.social @jennahgosciak.bsky.social and I organize an EAAMO Bridges working group on Urban Data Science and we are looking for new members! Fill the interest form on our page: urban-data-science-eaamo.github.io
Urban Data Science & Equitable Cities | EAAMO Bridges
EAAMO Bridges Urban Data Science & Equitable Cities working group: biweekly talks, paper studies, and workshops on computational urban data analysis to explore and address inequities.
urban-data-science-eaamo.github.io
📢 New paper: Applied interpretability 🤝 MT personalization! We steer LLM generations to mimic human translator styles on literary novels in 7 languages. 📚 SAE steering can beat few-shot prompting, leading to better personalization while maintaining quality. 🧵1/
Can vision-language models understand figurative meaning in multimodal inputs, like visual metaphors, sarcastic captions or memes? Come find out at our #NAACL2025 poster on Friday at 9am! New task & dataset of images and captions with figurative phenomena like metaphor, idiom, sarcasm, and humor.
Migration data lets us study responses to environmental disasters, social change patterns, policy impacts, etc. But public data is too coarse, obscuring these important phenomena! We build MIGRATE: a dataset of yearly flows between 47 billion pairs of US Census Block Groups. 1/5
People often claim they know when ChatGPT wrote something, but are they as accurate as they think? Turns out that while general population is unreliable, those who frequently use ChatGPT for writing tasks can spot even "humanized" AI-generated text with near-perfect accuracy 🎯