Really excited about this project from our team! AgentOdyssey allows you to create open-world, long-horizon environments on demand— which can then be used to evaluate LLM agents on exploration, memory, and planning over time. x.com/zheyuanzhan...
Daniel Khashabi
@danielkhashabi.bsky.social
I play with intuitions and data. Now: @jhuclsp @jhucompsci Past: @allen_ai @uwnlp @Penn @cogcomp @Illinois_Alma @MSFTResearch
Very honored and excited to receive the NSF CAREER Award! HUGE thank you to my amazing students, collaborators, mentors, and advisors, who helped make this happen. And to my family who are the real heroes in my story! ♥️
How should multiple agents 𝗰𝗼𝗺𝗺𝘂𝗻𝗶𝗰𝗮𝘁𝗲 𝘁𝗼 𝗰𝗼𝗼𝗿𝗱𝗶𝗻𝗮𝘁𝗲 in tight spaces? This remains challenging! See @suyu_ye's solution 👇 x.com/suyu_ye/sta...
LLMs continue to struggle with long-context tasks—such as needle-in-a-haystack problems—because of “positional bias.” What can we do if we only have 𝘣𝘭𝘢𝘤𝘬-𝘣𝘰𝘹 access to the model? (i.e., we can’t modify the model weights or attention patterns, as is often the case with API models.)
We have been busy building our science co-pilot for Genomics AI Agent at @DataTecnica which is specialized in Alzheimer’s and neurodegenerative disease research. This system: * Synthesizes complex biomedical data across literature and genomics databases
Postdoc positions: ai.jhu.edu/careers/pos... Applications are due January 23, 2026. Positions are for 2 years with the possibility of an extension.
Postdoctoral Fellowship Program - Johns Hopkins Data Science and AI Institute
The Johns Hopkins Data Science and AI (DSAI) Institute welcomes applications for its postdoctoral fellowship program, seeking disciplinarily diverse scholars to advance foundational methods of data science and artificial intelligence,…
ai.jhu.edu
Overdue update — CARDBiomedBench will be featured in @LancetDigitalH! 🎉 If you're looking for a high-quality and challenging science benchmark for your AI model, this could be it! 🤗 Dataset: huggingface.co/datasets/NI... 📄 Paper: biorxiv.org/content/10.... x.com/DanielKhash...
CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research
Backgrounds Biomedical research requires sophisticated understanding and reasoning across multiple specializations. While large language models (LLMs) show promise in scientific applications, their capability to safely and accurately support complex biomedical research remains uncertain. Methods We present CARDBiomedBench , a novel question-and-answer benchmark for evaluating LLMs in biomedical research. For our pilot implementation, we focus on neurodegenerative diseases (NDDs), a domain requiring integration of genetic, molecular, and clinical knowledge. The benchmark combines expert-annotated question-answer (Q/A) pairs with semi-automated data augmentation, drawing from authoritative public resources including drug development data, genome-wide association studies (GWAS), and Summary-data based Mendelian Randomization (SMR) analyses. We evaluated seven private and open-source LLMs across ten biological categories and nine reasoning skills, using novel metrics to assess both respon
biorxiv.org
Postdoc positions: ai.jhu.edu/careers/pos... Applications are due January 23, 2026. Positions are for 2 years with the possibility of an extension.
Postdoctoral Fellowship Program - Johns Hopkins Data Science and AI Institute
The Johns Hopkins Data Science and AI (DSAI) Institute welcomes applications for its postdoctoral fellowship program, seeking disciplinarily diverse scholars to advance foundational methods of data science and artificial intelligence,…
ai.jhu.edu
Postdoc positions: ai.jhu.edu/careers/pos... Applications are due January 23, 2026. Positions are for 2 years with the possibility of an extension.
Postdoctoral Fellowship Program - Johns Hopkins Data Science and AI Institute
The Johns Hopkins Data Science and AI (DSAI) Institute welcomes applications for its postdoctoral fellowship program, seeking disciplinarily diverse scholars to advance foundational methods of data science and artificial intelligence,…
ai.jhu.edu
Postdoc positions: ai.jhu.edu/careers/pos... Applications are due January 23, 2026. Positions are for 2 years with the possibility of an extension.
Postdoctoral Fellowship Program - Johns Hopkins Data Science and AI Institute
The Johns Hopkins Data Science and AI (DSAI) Institute welcomes applications for its postdoctoral fellowship program, seeking disciplinarily diverse scholars to advance foundational methods of data science and artificial intelligence,…
ai.jhu.edu
For years since the GPT-2 paper, emergent in-context learning (ICL) from 'next-token' training has been treated as something deeply tied to 𝐡𝐮𝐦𝐚𝐧 𝐥𝐚𝐧𝐠𝐮𝐚𝐠𝐞. But … is it?
Big congrats to @jackjingyuzhang for being named an Amazon AI PhD Fellow! 🎉 Grateful for @AmazonScience @RohitPrasadAI’s support as we work together to advance AI research at JHU. x.com/jackjingyuz...
ICL and SFT are the two most studied ways to adapt LMs. We understand each in isolation — but far less about how they might 𝗰𝗼𝗺𝗽𝗹𝗲𝗺𝗲𝗻𝘁 𝗼𝗻𝗲 𝗮𝗻𝗼𝘁𝗵𝗲𝗿.
Imagine this: excited about the recent progress, you’ve built an agentic system that uses 🔧tools (API calls) to solve complex problems. What could go wrong? We studied agentic tool recovery—when your LLM selects a set of tools to execute, but one turns out to be unavailable or incorrect.
A core hurdles in AI safety eval is that benchmarks (e.g., those on jailbreak attacks) quickly become outdated shortly after they are released (e.g., saturate, contaminate, patched).
A core hurdles in AI safety eval is that benchmarks (e.g., those on jailbreak attacks) quickly become outdated shortly after they are released (e.g., saturate, contaminate, patched).
Excited to collaborate up with LMArena, NIH, and DataTecnica to launch BiomedArena! Our goal is to advance the use of LLMs in biomedical discovery and incorporate community-driven insights to help shape the future of biomedical AI. ⚔️ Check it out: biomedarena.ai
What’s really going on inside LLMs when they handle non-English queries? Niyati Bafna @niyatibafna.bsky.social 's recent work introduces the **translation barrier hypothesis**, a framework for understanding multilingual model behavior. Paper: huggingface.co/papers/2506...
Paper page - The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
huggingface.co
🔈When LLMs solve tasks with a mid-to-low resource input or target language, their output quality is poor. We know that. But can we put our finger on what breaks inside the LLM? We introduce the 💥 translation barrier hypothesis 💥 for failed multilingual generation with LLMs. arxiv.org/abs/2506.22724
🚨New LLM benchmark🚨 We're releasing BiomedSQL🔬 for tabular reasoning over large-scale biomedical databases. This includes questions based on implicit scientific conventions—like statistical thresholds, effect direction, and drug approval status. 📄 Preprint: arxiv.org/pdf/2505.20321
Long-form inputs (e.g., needle-in-haystack setups) are the crucial aspect of high-impact LLM applications. While previous studies have flagged issues like positional bias and distracting documents, they've missed a crucial element: the size of the gold/relevant context.
There have been various efforts on disentangling "task learning" vs "task recall" in LLMs. We've recently explored a fresh angle by borrowing from cryptography: with substitution ciphers, we transform a given task into an equivalent, but cryptic (no pun intended!!) forms.
What is a university without "freedom of speech"? Apparently, ChatGPT has a better grasp than @nyuniversity. x.com/nebedaay/st...
**Certified Mitigation of Worst-Case LLM Copyright Infringement** TL;DR: We propose BloomScrub a framework to certifiably remove long verbatim quotes to reduce the risk of copyright violations.
Can LLMs can be co-pilots for peer review? Answering this requires evaluating *evaluate* whether LLMs can provide critiques that are *grounded* in the context of science papers. See @JiefuOu's dataset which has a collection of paper claims and their critiques: arxiv.org/pdf/2503.21717
📣📣📣 Tianjian @tli104 and I have refreshed our course material! self-supervised.cs.jhu.edu/sp2025/ These resources may be helpful if you're: (1) looking for slides to teach about LLMs, or (2) interested in diving deeper into the field.
CSCI 601.771: Self-supervised Models
Discussing latest breakthroughs in self-supervised language models
self-supervised.cs.jhu.edu
I will be at #NAACL2025 to present our LLM creativity benchmark. Drop by if interested (Poster Session 8, Fri, May 2)! I'd love to chat about RL and its interpretability, data influence for post-training, CogSci for LLM. Feel free to reach out and let's have some coffee together ☕ !
"Benchmarking Language Model Creativity: A Case Study on Code Generation" arxiv.org/abs/2407.09007 TLDR— Proposed a framework for benchmarking LLMs' 𝒄𝒓𝒆𝒂𝒕𝒊𝒗𝒊𝒕𝒚. x.com/Yining__Lu/...
People rely on search engines/chatbots to access science. But what if you want a bird’s-eye view of science, or to identify over- and under-explored areas? We introduce 🔺Science Hierarchography🔺, the goal of organizing science papers into conceptual hierarchies. arxiv.org/abs/2504.13834
Highlighting our #ICLR2025 papers 🧵🧵🧵 (1) "GenEx: Generating an Explorable World" openreview.net/pdf?id=8NlU... TLDR— Physical exploration can be expensive, and even impossible. Our proposed policy mitigates this by enabling agents to form an imaginative model of the 3D world.