🤔💭What even is reasoning? It's time to answer the hard questions! We built the first unified taxonomy of 28 cognitive elements underlying reasoning Spoiler—LLMs commonly employ sequential reasoning, rarely self-awareness, and often fail to use correct reasoning structures🧠
Stella Li
@stellali.bsky.social
PhD student @uwnlp.bsky.social @uwcse.bsky.social | visiting researcher @MetaAI | previously @jhuclsp.bsky.social https://stellalisy.com
Because Olmo 3 is fully open, we decontaminate our evals from our pretraining and midtraining data. @stellali.bsky.social proves this with spurious rewards: RL trained on a random reward signal can't improve on the evals, unlike some previous setups
Day 1 (Tue Oct 7) 4:30-6:30pm, Poster Session 2 Poster #77: ALFA: Aligning LLMs to Ask Good Questions: A Case Study in Clinical Reasoning; led by @stellali.bsky.social & @jiminmun.bsky.social
WHY do you prefer something over another? Reward models treat preference as a black-box😶🌫️but human brains🧠decompose decisions into hidden attributes We built the first system to mirror how people really make decisions in our recent COLM paper🎨PrefPalette✨ Why it matters👉🏻🧵
Want to quickly sample high-quality images from diffusion models, but can’t afford the time or compute to distill them? Introducing S4S, or Solving for the Solver, which learns the coefficients and discretization steps for a DM solver to improve few-NFE generation. Thread 👇 1/
Asking the right questions can make or break decisions in fields like medicine, law, and beyond✴️ Our new framework ALFA—ALignment with Fine-grained Attributes—teaches LLMs to PROACTIVE seek information through better questions through **structured rewards**🏥❓ (co-led with @jiminmun.bsky.social) 👉🏻🧵
New preprint! Metaphors shape how people understand politics, but measuring them (& their real-world effects) is hard. We develop a new method to measure metaphor & use it to study dehumanizing metaphor in 400K immigration tweets Link: bit.ly/4i3PGm3 #NLP #NLProc #polisky #polcom #compsocialsci 🐦🐦
The feeling when you thought you were scooped from the title and the abstract but then unscooped after reading the actual paper is my best christmas present🤌💕🎁
Excited to present MediQ at #NeurIPS ! 📍Stop by my poster: East Exhibit Hall A-C #4805📷 🕚Thu, Dec 12 | 11am–2pm 🗓️tinyurl.com/mediq2024 Love to chat about anything--reasoning, synthetic data, multi-agent interaction, multilingual nlp! Message me if you want to chat☕️🍵🧋
31% of US adults use generative AI for healthcare 🤯But most AI systems answer questions assertively—even when they don’t have the necessary context. Introducing #MediQ a framework that enables LLMs to recognize uncertainty🤔and ask the right questions❓when info is missing: 🧵
31% of US adults use generative AI for healthcare 🤯But most AI systems answer questions assertively—even when they don’t have the necessary context. Introducing #MediQ a framework that enables LLMs to recognize uncertainty🤔and ask the right questions❓when info is missing: 🧵
study music:1, taylor swift:0 (only for September I swear😅) Spotify knows too much about me….
We created an Allen School Starter Pack to help you find and connect with @uofwa.bsky.social #UWAllen labs and researchers on 🦋! We'll add to this list as we grow our community here: go.bsky.app/RyHBLJd #AcademicBluesky #CompSci #AI #CompBio #UbiComp #Accessibility #NLP #HCI #DataViz #mHealth
Allen School Starter Pack
Join the conversation
go.bsky.app
OpenReview turned into Reddit🤯 Can we now add upvote/downvote buttons to reviews and rebuttals plz?? Would be a very rich and interesting source of preference data🤡
Looking at ICLR submissions with the lowest score - What a work of art! 🧵
I noticed a lot of starter packs skewed towards faculty/industry, so I made one of just NLP & ML students: go.bsky.app/vju2ux Students do different research, go on the job market, and recruit other students. Ping me and I'll add you!
I'm biased ofc but IMO the ~right theoretical framework for thinking and talking about instructable LLM-based agents is assistance games (a.k.a. Cooperative Inverse Reinforcement Learning (CIRL) games).
Setting aside the fact that people are literally re-inventing a bunch of terminology around agents when we have several decades of framing from RL (POMDPs, anyone?) that we can rely on to describe what we're doing, the main challenges are as follows., working backwards from the end point. [2/11]
Ready for another Computational Social Science Starter Pack? Here is number 2! More amazing folks to follow! Many students and the next gen represented! go.bsky.app/GoEyD7d
1/ Introducing ᴏᴘᴇɴꜱᴄʜᴏʟᴀʀ: a retrieval-augmented LM to help scientists synthesize knowledge 📚 @uwnlp.bsky.social & Ai2 With open models & 45M-paper datastores, it outperforms proprietary systems & match human experts. Try out our demo! openscholar.allen.ai
I noticed that all the X posts about migrating to 🦋 (with the app name, the emoji, or even “the other platform”) get extremely down ranked😅 very sad/disturbed to see the level of content control over there🤯
Had so much fun presenting our work ValueScope (arxiv.org/abs/2407.024...) at #EMNLP2024 and (mostly) enjoying Miami with friends last week🫶🏻🏖️!! Also learned about so many amazing works and now I’m inspired to start working again🫡💕