Zory Zhang

@zoryzhang.bsky.social

Student of children's cognition with models and experiments. Brown PhD student with @daphnab.bsky.social‬

On the other hand, any conclusions we draw will have a very short shelf life, because the machines are in constant flux. Meanwhile, the ease with which one can do this research will sap the attention of researchers away from the harder work of understanding humans.

The viral "Definition of AGI" paper tells you to read fake references which do not exist! Proof: different articles present at the specified journal/volume/page number, and their titles exist nowhere on any searchable repository. Take this as a warning to not use LMs to generate your references!

BildBildBildBild

New Paper Alert ‼️ Current VLMs completely fail human gaze understanding 🙀 and scaling does NO help ‼️ However, humans, since an extremely age 🧒, are extremely sensitive to other people's gaze 🙄 👀 No mentors, no labs, only pre-doc students, 111 VLMs, and we did it 😎

Bild

👁️ 𝐂𝐚𝐧 𝐕𝐢𝐬𝐢𝐨𝐧 𝐋𝐚𝐧𝐠𝐮𝐚𝐠𝐞 𝐌𝐨𝐝𝐞𝐥𝐬 (𝐕𝐋𝐌𝐬) 𝐈𝐧𝐟𝐞𝐫 𝐇𝐮𝐦𝐚𝐧 𝐆𝐚𝐳𝐞 𝐃𝐢𝐫𝐞𝐜𝐭𝐢𝐨𝐧? Knowing where someone looks is key to a Theory of Mind. We test 111 VLMs and 65 humans to compare their inferences. Project page: grow-ai-like-a-child.github.io/gaze/ 🧵1/11

Sam is 100% correct on this. Indeed, human babies have essential cognitive priors such as permanence, continuity, and boundary of objects, 3D Euclidean understanding of space, etc. We spent 2 years to systematically to examine and show the lack of such in MLLMs: arxiv.org/abs/2410.10855

Bild
Sam Gershman@gershbrain.bsky.social · last yr.

I think the BabyLM Challenge is really interesting, but also feel that there is something fundamentally ill-posed about how it maps onto the challenge facing human children. It's true that babies only get a relatively limited amount of linguistic experience, but...