Zizhao Chen

@ch272h.bsky.social

chenzizhao.github.io tearing down natural stupidity while phding @cornelltech.bsky.social

🧩Natural language isn’t all you need. We’re great at evaluating text-based reasoning (MATH, AIME…) but what about long-horizon visual reasoning? Enter 𝗞𝗻𝗼𝘁𝗚𝘆𝗺: a minimalistic testbed for evaluating agents on spatial reasoning along a difficulty ladder

- Coding interview without copilot: I can’t type - IELTS writing test without Gmail autocompletion: I can’t spell I guess these evaluation formats are out of date. Or more likely, tab-AI made me dumber. I wonder how it feels like to be born in 2022 and grow up in a world with llms.

So I was volunteering today. I prompted folks randomly this question after they collected their neurips thermos: Do you think AIs today are intelligent? Answer with yes or no. Here is the break down: Yes: 57 No: 62 Total: 119 Pretty close!

Bild