Khai Loong Aw
@khaiaw.bsky.social
CS PhD student @Stanford. Research on AI, cognitive science, and neuroscience.
Excited to be at #CCN2026 in NYC! I'll be presenting our spotlight poster, Topo-Omni, with @hannesmehrer.bsky.social 🧠 📍 Poster F14 — Session F 🗓️ Thursday, Aug 6 Come by if you're around, we'd love to talk about how our multimodal model can discover functionally selective brain regions!
Will be presenting our work on building a Universal Vision-Language World Model. It uses visual abstractions as streams of thought, e.g., camera pose, depth, optical flow, point tracks, text. Our model solves vision-language tasks using world modeling and inverse dynamics, unlike standard VLMs.
Our Stanford NeuroAI Lab is presenting at the Cognitive Computational Neuroscience conference #CCN2026 🥳 Come say hi to us!! Imran, @norcalneuro.bsky.social, @dyamins.bsky.social, Khaled, @ynshah.bsky.social @seojinlee.bsky.social , @clionaod.bsky.social , Khai
Our Stanford NeuroAI Lab is presenting at the Cognitive Computational Neuroscience conference #CCN2026 🥳 Come say hi to us!! Imran, @norcalneuro.bsky.social, @dyamins.bsky.social, Khaled, @ynshah.bsky.social @seojinlee.bsky.social , @clionaod.bsky.social , Khai
Our Stanford NeuroAI Lab is presenting at the Cognitive Computational Neuroscience conference #CCN2026 🥳 Come say hi to us!! Imran, @norcalneuro.bsky.social, @dyamins.bsky.social, Khaled, @ynshah.bsky.social @seojinlee.bsky.social , @clionaod.bsky.social , Khai
nice review! amazing that there has been more than a decade of work mapping AI vision & language model representations to human brain responses (yamins et al 2014, wehbe et al 2014)
1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.
This review began with the great Peter Hagoort asking after one of my talks: "What have we actually learned about the brain from brain–DNN comparisons?" 🔥 So eye-opening to survey NeuroAI work across vision and language--many questions have been answered in one field but remain open in the other!
1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.
1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.
If you’re at #cogsci2026, please come see presentations by some of the great folks collaborating with the Language and Cognition Lab at Stanford!
Excited to share our review in @cp-neuron.bsky.social with @lauriebayet.bsky.social and @mickbonner.bsky.social! We describe how implementing principles from child development can advance the mechanistic plausibility and capacities of AI models We packed A LOT into this review, here's a quick 🧵
(1/n) Thrilled to share my first paper at Meta FAIR! "EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data" 👶 Human infants learn language from sparse, noisy multimodal input. Today's VLMs can't. We built a benchmark + challenge to close that gap. 🧵
New preprint! AI agents have shown impressive scientific automation capabilities. Can we apply them to psychology research, *including* human data collection? We introduce auto-psych, a framework that proposes cognitive models and uses them to design and run human experiments. 1/
If a student reads a typical psychology paper, picks an experiment, and tries to replicate it, they have roughly a coin-flip chance of success. That's the punchline of a decade of metascience, and it's the focus of Ch 3 of Experimentology. 🧵 experimentology.io
1/ New preprint with @dyamins.bsky.social + team! Ventral visual representations within areas evolve over the course of the response along the same hierarchical complexity axis that distinguishes the visual areas, potentially driven by local recurrence.
A hierarchical computational motif unifies neural dynamics across the ventral visual stream https://www.biorxiv.org/content/10.64898/2026.05.18.726101v1
What is a psychological theory? Here's our take on this tricky and controversial question in this week's Experimentology chapter summary. Many things called "theories" in psychology aren't actually theories — they're frameworks. 🧵 experimentology.io
Exciting analysis of objects in kids' everyday visual input (BabyView). Today's best categorization models still train on curated photos — not the long-tailed categories and non-canonical viewpoints kids actually see. A difference in kind, not just quantity, pointing to a fundamental algorithmic gap
Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process actually look like? New preprint! arxiv.org/abs/2605.14990
Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process actually look like? New preprint! arxiv.org/abs/2605.14990
Characterizing the visual representation of objects from the child's view
Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process look like? We analyzed first-person videos ...
arxiv.org
What is the function of function words ⁉️Using head-mounted eye-tracking and machine learning, we show how tiny words like THIS and THAT actively power word learning 👶🤖. Check our preprint, we’d love to hear your feedback!!!
In our latest paper on how multimodal interaction supports language acquisition, Zhang, Dong et al use head mounted eyetracking technology to demonstrate (!) the role of demonstratives (this/that) in helping children identify the correct referent in a complex visual scene osf.io/preprints/ps...
For a year and a half, @carorowland.bsky.social, @lehersingh.bsky.social, Marisa Casillas, Shanley Allen, and I have been meeting to discuss whether innateness is still a useful concept to think about in studying language acquisition. Here's our take: osf.io/preprints/ps...
Just wrote a new blogpost trying to summarize my thoughts on the question of how and whether to use AI for research in psychology and cognitive science: babieslearninglanguage.blogspot.com/2026/04/usin...
Using AI to improve (not automate away) academic research
Blog about fatherhood, langauge, developmental psychology, and cognitive science.
babieslearninglanguage.blogspot.com
There's a lot to like here! - Very smart way to use a masked autoencoder (unsupervised technique!) to build a world model from visual data. This makes other visually based world models I've seen seem clumsy in comparison. 1/4
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵
Beautiful use of the BabyView dataset to train a visual learning model!
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵
I think we finally made really significant progress on the biggest unsolved "developmental AI" problem: learning from human-scale data. Key idea: zero-shot world models that support concept extraction via approximate causal inference. amazing collab w/ @mcxfrank.bsky.social @khaiaw.bsky.social
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵
So excited about this work using our data of children’s first-person experiences to train efficient, flexible visual learning models!
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵
Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵