Hallucinations pose a substantial challenge to the reliability of LLMs in real-world scenarios. Zhang et al. survey methods of detection, explanation, & mitigation of hallucination, & provide a taxonomy & list of benchmarks for evaluation in this paper: doi.org/10.1162/COLI... @fredashi.bsky.social
Freda Shi
@fredashi.bsky.social
Assistant Professor & Canada CIFAR AI Chair, University of Waterloo & Vector Institute | Excited about "grounding" in any form | Feeder of 2 🐈 | 🏸, 🏐, 🏂 | she/her
My student and her undergrad advisor submitted their very elegant work on X to "Transactions of X" and got desk-rejected, with the reason: "Your paper is very interesting and a very hardcore contribution to X, but we can't find appropriate reviewers, so we have to desk-reject it." Excuse me???
Wow, learned about an even more impressive statement today: We are Sorry to inform you that you Didn’t Get the Honor for Voluntary Contribution because the speakers you invited (including yours truly) are Not Famous Enough.
We are Sorry to inform you that you Didn't Get the Honor for Voluntary Contribution because you onboarded Junior People.
We are Sorry to inform you that you Didn't Get the Honor for Voluntary Contribution because you onboarded Junior People.
We are Delighted to Congratulate you on the Honor of the Opportunity of Volunteering
We are Delighted to Congratulate you on the Honor of the Opportunity of Volunteering
Recent takes from this effort: Humans are largely linear animals, and we are so used to all kinds of linear structures, including doing tasks one by one.
I've been enjoying developing a (safe & fairly reliable) personal assistant bot, partly motivated by some light experience trying OpenClaw. For safety, everything is run & saved locally. I can now manage my todo list, notebook, and calendar simply by speaking to my bot. (1/)
Open PhD/Postdoc position (start: Oct 2026). Topic: AI/LLMs and child language/communicative/cognitive development. The exact project will be shaped with the candidate. Join our team @univ-amu.fr at the intersection of computer and cognitive science (& right next to the Calanques!). Send me your CV!
I truly feel this AI-assisted development is something completely new. The key difference is the design philosophy: I also like Notion, for example, but I had to adapt myself to their design. Now the assistant bot and myself are doing "bidirectional alignment"---AI coding agents enable this.
I've been enjoying developing a (safe & fairly reliable) personal assistant bot, partly motivated by some light experience trying OpenClaw. For safety, everything is run & saved locally. I can now manage my todo list, notebook, and calendar simply by speaking to my bot. (1/)
I've been enjoying developing a (safe & fairly reliable) personal assistant bot, partly motivated by some light experience trying OpenClaw. For safety, everything is run & saved locally. I can now manage my todo list, notebook, and calendar simply by speaking to my bot. (1/)
I frankly don't think the ICML policy of having authors rank their papers will work. I have 3-5 submissions (depending on ICLR results) this time to quite different subcommunities. I know my favorite submissions could be quite controversial---people either like or dislike them a lot. (1/2)
Thrilled to announce the 1st Workshop on Computational Developmental Linguistics (CDL) at ACL 2026 🎉 A new venue at the intersection of development linguistics × modern NLP, spearheaded by @fredashi.bsky.social @marstin.bsky.social, and and outstanding team of colleagues! A thread 🧵
AI coding is now the best at implementing things that are (1) either already standardized or not very ambiguous in their natural-language description, and (2) with many details. Yes, I just vibe-coded a helper for some organizational matters, and will do more. 1/
It's #NSF #GRFP application season again so it's time to re-up my GRFP application advice post! Also, check out the cool bsky comment integration I've added to the blog! Engagement with this post will go under the blogpost on my site as comments! saxon.me/blog/2024/gr...
NSF GRFP Application Tips for NLP, AI, CS
Reflections and advice from my successful NSF GRFP proposal in NLP. Why I think my applications worked well, what I wish I did differently, and links to my actual statements and feedback from the GRFP...
saxon.me
Reviewer timeliness by area (based on a small sample of 14 papers * 4 reviewers each): Multilingual LMs >> VLMs > LLM Reasoning.
We have updated our collection of multilingual poetry corpora PoeTree. With addition of Norwegian, and few metadata fixes it now has a loud label of 1.0.0 release! 🌳 versologie.cz/poetree/vers...
PoeTree. Poetry corpora in 11 languages
PoeTree is a standardized collection of poetry corpora comprising nearly 335,000 poems in ten languages (Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Russian, Slovenian, ...
versologie.cz
Should be fun to read!
@fredashi.bsky.social and I wrote a blog for our new mechinterp paper (arxiv.org/abs/2510.13796), including many unpublished and even negative results that we found meaningful to share. An Open-Notebook Exploration of Emergent Grounding in LMs mars-tin.github.io/blogs/posts/...
I can only go through <40 slides in a one-hour talk... verified multiple times. My job talk has 39 slides (including acknowledgement), and so are my recent talks. Always impressed when folks present 100 slides in the same amount of time.
In 2016 Hinton predicted that AI would replace all radiologists in five years. Ten years later, why hasn't it happened? This post is a great explainer. www.understandingai.org/p/ai-isnt-re...
AI isn't replacing radiologists
Radiology combines digital images, clear benchmarks, and repeatable tasks. But demand for human radiologists is at an all-time high.
understandingai.org
🚀 ACL ARR is looking for a Co-CTO to join me lead our amazing tech team and drive the future of our workflow. If you’re interested or know someone who might be, let’s connect! RTs & recommendations appreciated.
🚨 ARR is looking for a volunteer Co-CTO to help improve tech infrastructure! 🛠️ Preferred: • 5+ years in NLP research • Git, CLI tools, Python, and basic HTML • 2-year role, overlapping with current Co-CTO Interested? DM @fredashi.bsky.social or email fhs@uwaterloo.ca #ARR #ACL #NLProc
Same! I‘ve literally unbidden 0 out of my recommended batch.
I have soooo many interesting-looking papers in my ICLR - AC bidding batch... I am so happy... 🤩
Is there a specific reason that NeurIPS does not show AC identity to reviewers? I'm very curious about who sent the polite reminders, and of course, even more curious about the rude ones.
#CoreCognition #LLM #multimodal #GrowAI We spent 3 years to curate 1503 classic experiments spanning 12 core concepts in human cognitive development and evaluated on 230 MLLMs with 11 different prompts for 5 times to get over 3.8 millions inference data points. A thread (1/n) - #ICML2025 ✅
Just in case this personal latex environmental setup is helpful to a broader crowd ⬇️
I use local complication w/ VSCode latex workshop plugin + git for version control. Pros: better hierarchical organization & doesn’t require internet. Cons: sometimes need to resolve conflict (usually fine if commit frequently) & require a powerful laptop to match overleaf compilation speed.
I'm late to the NAACL party, but I've just arrived in Albuquerque! I'll talk about measuring faithfulness of verbalized reasoning on *Sunday* at Repl4NLP at *9:45*.
hello NAACL friends I'm giving a keynote today at RepL4NLP at 1:30PM local time, come say hi! I'll mostly be musing about things with light research discussions
On my way to NAACL✈️! If you're also there and interested in grounding, don't miss our tutorial on "Learning Language through Grounding"! Mark your calendar: May 3rd, 14:00-17:30, Ballroom A. Another exciting collaboration with @marstin.bsky.social @kordjamshidi.bsky.social, Jiayuan, and Joyce!
📢Curious why your LLM behaves strangely after long SFT or DPO? We offer a fresh perspective—consider doing a "force analysis" on your model’s behavior. Check out our #ICLR2025 Oral paper: Learning Dynamics of LLM Finetuning! (0/12)
✈ Just landed in Singapore for #ICLR 2025! DM or email me if you'd like to chat about - Grounded language acquisition and learning - What do vision-language models "know," and what they don't - (Computational) linguistics with/for language models, especially grounded LMs (1/)