Joseph Chang

@josephc.bsky.social

https://josephcc.com

Can AI really help with literature reviews? 🧐 Meet Ai2 ScholarQA, an experimental solution that allows you to ask questions that require multiple scientific papers to answer. It gives more in-depth and contextual answers with table comparisons and expandable sections 💡 Try it now: scholarqa.allen.ai

Ai2 ScholarQA logo

🤔Giving complex tasks to AI agents is easy—getting them to do exactly what you want isn’t. How can human-AI collaboration give us more reliable & steerable agents? 🍫Introducing Cocoa, our new interaction paradigm for balancing human & AI agency in complex human-AI workflows. 🧵

A screenshot of the first page of an academic figure with a list of authors and a high-level schematic figure describing Cocoa, an interactive system for co-planning and co-execution with AI agents.

How do scholars of diff backgrounds use LLMs as research tools? How do we perceive the risks & benefits of this new practice? Are we willing to disclose to peers and reviewers? We conducted a large-scale survey of verified authors of different fields, race, gender, seniority to find out - results🧵

Simona Liao@simonaliao.bsky.social · 2y ago

Hi everyone, I am excited to share our large-scale survey study with 800+ researchers, which reveals researchers’ usage and perceptions of LLMs as research tools, and how the usage and perceptions differ based on demographics. See results in comments! 🔗 Arxiv link: arxiv.org/abs/2411.05025

I'm recruiting 1-2 PhD students to work with me at the University of Colorado Boulder! Looking for creative students with interests in #NLP and #CulturalAnalytics. Boulder is a lovely college town 30 minutes from Denver and 1 hour from Rocky Mountain National Park 😎 Apply by December 15th!

A photo of Boulder, Colorado, shot from above the university campus and looking toward the Flatirons.

Excited to share ✨ Contextualized Evaluations ✨! Benchmarks like Chatbot Arena contain underspecified queries, which can lead to arbitrary eval judgments. What happens if we provide evaluators with context (e.g who's the user, what's their intent) when judging LM outputs? 🧵↓

Bild

Lit reviews often involves comparing sets of papers using common aspects in tables or spreadsheets - @benn9.bsky.social & Yoonjoo Lee's #EMNLP paper explored ways to create such tables using LLMs, and how to evaluate them against a large set of lit review tables we extracted from arXiv.

Ben Newman@benn9.bsky.social · 2y ago

✨EMNLP Paper! ✨ Have you ever constructed a table to organize your literature review process? Can we use LMs to generate these automatically? We are excited to present ArxivDIGESTables 🍽️ a study of collecting, generating, and evaluating 🎓 scientific literature review tables 📃!

A screenshot of the first page of the paper discussed in the thread. Figure 1 contains a set of three cartoon papers with related text highlighted in three different colors. To its left, there's an arrow pointing to a cartoon table with a column corresponding to each color and a row corresponding to each paper.