Gabriel Agostini

@gsagostini.bsky.social

PhD student at Cornell Tech | he/him | cities + equity + spatial everything | fan of cats and Taylor Swift | gsagostini.github.io

Excited to share #ICML2026 paper from my internship @ Google DeepMind! AI models are deployed globally, but AI safety datasets are largely geographically homogenous. What is the impact of culture on AI safety ratings? Is there any impact beyond standard demographics like age, gender, and ethnicity?

Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access!

IPUMS@ipums.bsky.social · 3mo ago

IPUMS Spatial Student Award is a tie! @gsagostini.bsky.social for "Inferring Fine-Grained Migration Patterns Across the United States." (www.nature.com/articles/s41...)

We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?)

New in Nature Health: how might we move towards a world in which race is not used in clinical algorithms? We need (1) careful comparison of race-aware and race-neutral algorithms and (2) systemic efforts to address underlying disparities.

New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246

Bild

November is over, but we still have some #30DayMapChallenge entries to share! And for our transport-themed day 26 map, MBTA data analyst Joe Hilleary takes us on a ride back in time: he shows current bus routes in Greater Boston by the earliest known year in which a direct percursor route ran a bus.

Bild

We have a new paper in Science Advances proposing a simple test for bias: Is the same person treated differently when their race is perceived differently? Specifically, we study: is the same driver likelier to be searched by police when they are perceived as Hispanic rather than white? 1/

Bild

My best workflow improvement since starting to work with spatial libraries in Python was to always include a `crs` dictionary on a variables file listing crs for lat-long projections, equidistant projections, and "maybe not satisfying any desiderata but the prettiest out there" projections.

Urban Data@urban-data.bsky.social · 8mo ago

#30DayMapChallenge day 19: projections @jennahgosciak.bsky.social created a gif that visualizes NYC neighborhoods under different map projections. While most atrocities are clear only at national or global levels, her maps show interesting local deformities!

Great map(s) by @jennahgosciak.bsky.social ---can we count that for 6 days of mapping??---that show both the permanence and the vulnerability of ecological concepts in our urban landscapes! #30DayMapChallenge

Urban Data@urban-data.bsky.social · 9mo ago

We might be a few days delayed on the #30DayMapChallenge, but our day 5 submission spans almost 250 years of history! Our "Earth" map comes from @jennahgosciak.bsky.social, who compared the original ecology of New York City to present day variables.

N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity.

N-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measuring textual creativity. However, theoretical work on creativity suggests that this approach may be inadequate, as it does not account for creativity's dual nature: novelty (how original the text is) and appropriateness (how sensical and pragmatic it is). We investigate the relationship between this notion of creativity and n-gram novelty through 7542 expert writer annotations (n=26) of novelty, pragmaticality, and sensicality via close reading of human and AI-generated text. We find that while n-gram novelty is positively associated with expert writer-judged creativity, ~91% of top-quartile expressions by n-gram novelty are not judged as creative, cautioning against relying on n-gram novelty alone. Furthermore, unlike human-written text, higher n-gram novelty in open-source LLMs correlates with lower pragmaticality. In an exploratory study with frontier close-source models, we additionally confirm that they are less likely to produce creative expressions than humans. Using our dataset, we test whether zero-shot, few-shot, and finetuned models are able to identify creative expressions (a positive aspect of writing) and non-pragmatic ones (a negative aspect). Overall, frontier LLMs exhibit performance much higher than random but leave room for improvement, especially struggling to identify non-pragmatic expressions. We further find that LLM-as-a-Judge novelty scores from the best-performing model were predictive of expert writer preferences.

New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.

Bild

Very happy Divya has been around during my PhD. I might be deep into maps and she might be deep into health (...and so much more!) but I could always count on learning something from her. She's such a kind researcher and great science communicator!

Divya Shanmugam@dmshanmugam.bsky.social · 10mo ago

I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.