We have released #AgentCoMa, an agentic reasoning benchmark where each task requires a mix of commonsense and math to be solved 🧐 LLM agents performing real-world tasks should be able to combine these different types of reasoning, but are they fit for the job? 🤔 🧵⬇️
Joe Stacey
@joestacey.bsky.social
NLP PhD student at Imperial College London and Apple AI/ML Scholar.
Here’s my review of the US after a few days here. Did I miss anything? 🤔 The good: - Americans are the most charming, friendly and hospitable people - it’s super fun how the country is split into states that all have different laws and stuff, with different vibes state to state
Any chance Keir Starmer can reshuffle himself in as foreign secretary, and shuffle in another prime minister who actually has some vague idea about what they want to achieve? 🙏🤦♂️
Finally the heatwave has ended, and the UK is once again a bearable place to be 😍😍 If you have any UK-based collaborations, their productivity is about to increase like 10 fold
We have a fun new #NLProc paper on arXiv about improving the robustness of fine-tuned NLI models! Have a look :) arxiv.org/abs/2505.20209
Should I use an LLM to help refine my paper writing for the ARR deadline? 🤔🤔 It will improve the paper for sure, but probably also making the tone a whole lot more annoying
If you're at #NAACL2025 and want to hear about similarity effects for property inheritance in LMs, please stop by! I will be presenting this work on Wednesday at the 11-12:30 poster session on Interpretability & analysis for language models (Hall 3). aclanthology.org/2025.naacl-l...
Characterizing the Role of Similarity in the Property Inferences of Language Models
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...
aclanthology.org
How do language models organize concepts and their properties? Do they use taxonomies to infer new properties, or infer based on concept similarities? Apparently, both! 🌟 New paper with my fantastic collaborators @amuuueller.bsky.social and @kanishka.bsky.social
Excited to share our ICLR and NAACL papers! Please come and say hi, we're super friendly :)
Wow, the old ITV Agatha Christie’s Poirot is brilliant. Some tv for 1989… Gonna go binge watch the 13 seasons now 😍
I feel like the length of the ARR author rebuttals keep growing every cycle Is this a good thing for authors or reviewers that the responses can be so long? I feel like it’s a bit sub-optimal for both at the moment
Had a great time presenting my research on building more helpful QA systems @imperialcollegeldn.bsky.social! Thank you @joestacey.bsky.social for letting me invite myself 🫶 And loved visiting London+Edinburgh this week, hope to be back soon! 🙏
Was fantastic to have you here at Imperial! Thanks for your excellent talk, and looking forward to following what you do next 🙂
Had a great time presenting my research on building more helpful QA systems @imperialcollegeldn.bsky.social! Thank you @joestacey.bsky.social for letting me invite myself 🫶 And loved visiting London+Edinburgh this week, hope to be back soon! 🙏
Do LLMs need rationales for learning from mistakes? 🤔 When LLMs learn from previous incorrect answers, they typically observe corrective feedback in the form of rationales explaining each mistake. In our new preprint, we find these rationales do not help, in fact they hurt performance! 🧵
Today was the launch event of the @genaihub.bsky.social. We announced the development of Nightingale AI, a foundation world model for health. It was great to be on the panel for GenAI in Healthcare, among such amazing experts. www.genai.ac.uk
Thanks so much to everyone who has helped make this switch to BlueSky work. Honestly, making this switch was a pretty massive achievement, so thanks everyone for contributing ❤️❤️
This paper is really cool. They decompose NLI (and defeasible NLI) hypotheses into atoms, and then use these atoms to measure the logical consistency of LLMs. E.g. for an entailment NLI example, each hypothesis atom should also be entailed by the premise. Very nice idea 👏👏
I’m a week into my trip from Cairo to Riyadh, and wow what a place Egypt is! Honestly its been one of the funnest places I’ve travelled, and for sure I need to come back again Crossed into Aqaba (Jordan) yesterday, so now onto Saudi 🙂
I’m going away to do a bit of travelling, going overland from Cairo to Riyadh 😍 I love travelling in the Middle East so it should be interesting I’ve got that feeling of nervous excitement I always get before a trip 😬😁
Insanely jealous to everyone who has papers at #NAACL in Albuquerque! Albuquerque just sounds so exotic, and is such a cool place for a conference. No offence to Vienna, but Albuquerque sounds way more fun 😉
Feeling gooooood after submitting my #ARR reviews early 😍 Time to enjoy the weekend! 🕺
I was super excited to read the ModernBERT paper! Love this interest in creating a better encoder model. "ModernBERT-base is the first encoder to beat DeBERTaV3-base since its release in 2021" 🤯- arxiv.org/pdf/2412.13663 Pretty amazing how successful DeBERTa has been!
Excited to start my #ARR #NLP reviews! I'll try my best and see if I can get 100% of my reviews to be 'great' this round. If you didn't see it already, ARR publishes how many of your reviews are considered to be 'great': stats.aclrollingreview.org Join me for the challenge :)
ARR Dashboard
stats.aclrollingreview.org
At some point in life I realised I actually really love travelling by train. Kind of a strange hobby, but wow it is fun 😍 Here are my top ten train journeys so far.
Imperial are hiring computing lecturers (including for AI/ML/NLP)! Here's a little thread about why you should consider applying :)
We are hiring 6 lecturers at @imperialcollegeldn.bsky.social to work on AI, ML, graphics, vision, quantum and software engineering. This includes researchers working on LLMs, NLP, generative models and text applications. Deadline 6 Jan. @imperial-nlp.bsky.social www.imperial.ac.uk/jobs/search-...
Made it to northern Sweden (Kiruna) by train from London. Freezing cold with northern lights 😍 Just over a week ago and I was in the crazy Miami heat for #EMNLP2024
This papers' findings about testing LLMs on NLI aligns with many of personal thoughts: 1) NLI remains a difficult task for LLMs 2) Having more few-shot examples is helpful (in my view, helping LLMs better understand class boundaries) 3) Incorrect predictions are often a result of ambiguous labels
I’ve seen some pretty amazing metros before (like Moscow), but wow Stockholm is wild. Never seen anything like it!
If anyone is getting annoyed with their BlueSky feed, try 'Popular with Friends' - you can add this from the 'Feeds' tab. I'm finding it works a bit better for me, and is more like what I had on Twitter. Thanks @lasha.bsky.social for suggesting!