Khai Loong Aw

@khaiaw.bsky.social

CS PhD student @Stanford. Research on AI, cognitive science, and neuroscience.

Will be presenting our work on building a Universal Vision-Language World Model. It uses visual abstractions as streams of thought, e.g., camera pose, depth, optical flow, point tracks, text. Our model solves vision-language tasks using world modeling and inverse dynamics, unlike standard VLMs.

Bild
Khai Loong Aw@khaiaw.bsky.social · 2w ago

Our Stanford NeuroAI Lab is presenting at the Cognitive Computational Neuroscience conference #CCN2026 🥳 Come say hi to us!! Imran, @norcalneuro.bsky.social, @dyamins.bsky.social, Khaled, @ynshah.bsky.social @seojinlee.bsky.social , @clionaod.bsky.social , Khai

nice review! amazing that there has been more than a decade of work mapping AI vision & language model representations to human brain responses (yamins et al 2014, wehbe et al 2014)

Dota Tianai Dong@dotadotadota.bsky.social · 4w ago

1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.

This review began with the great Peter Hagoort asking after one of my talks: "What have we actually learned about the brain from brain–DNN comparisons?" 🔥 So eye-opening to survey NeuroAI work across vision and language--many questions have been answered in one field but remain open in the other!

Dota Tianai Dong@dotadotadota.bsky.social · 4w ago

1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.

(1/n) Thrilled to share my first paper at Meta FAIR! "EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data" 👶 Human infants learn language from sparse, noisy multimodal input. Today's VLMs can't. We built a benchmark + challenge to close that gap. 🧵

Bild

New preprint! AI agents have shown impressive scientific automation capabilities. Can we apply them to psychology research, *including* human data collection? We introduce auto-psych, a framework that proposes cognitive models and uses them to design and run human experiments. 1/

Bild

If a student reads a typical psychology paper, picks an experiment, and tries to replicate it, they have roughly a coin-flip chance of success. That's the punchline of a decade of metascience, and it's the focus of Ch 3 of Experimentology. 🧵 experimentology.io

1/ New preprint with @dyamins.bsky.social + team! Ventral visual representations within areas evolve over the course of the response along the same hierarchical complexity axis that distinguishes the visual areas, potentially driven by local recurrence.

Bild
bioRxiv Neuroscience@biorxiv-neursci.bsky.social · 3mo ago

A hierarchical computational motif unifies neural dynamics across the ventral visual stream https://www.biorxiv.org/content/10.64898/2026.05.18.726101v1

What is a psychological theory? Here's our take on this tricky and controversial question in this week's Experimentology chapter summary. Many things called "theories" in psychology aren't actually theories — they're frameworks. 🧵 experimentology.io

Exciting analysis of objects in kids' everyday visual input (BabyView). Today's best categorization models still train on curated photos — not the long-tailed categories and non-canonical viewpoints kids actually see. A difference in kind, not just quantity, pointing to a fundamental algorithmic gap

jane-yang.bsky.social@jane-yang.bsky.social · 3mo ago

Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process actually look like? New preprint! arxiv.org/abs/2605.14990

Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵

Bild

Children exhibit visual understanding from limited experience, orders of magnitude less than our best models. We introduce the Zero-shot World Model (ZWM). Trained on a single child's visual experience, BabyZWM rapidly generates competence across diverse benchmarks with no task-specific training. 🧵

Bild