Joe Stacey

@joestacey.bsky.social

NLP PhD student at Imperial College London and Apple AI/ML Scholar.

We have released #AgentCoMa, an agentic reasoning benchmark where each task requires a mix of commonsense and math to be solved 🧐 LLM agents performing real-world tasks should be able to combine these different types of reasoning, but are they fit for the job? 🤔 🧵⬇️

Bild

Here’s my review of the US after a few days here. Did I miss anything? 🤔 The good: - Americans are the most charming, friendly and hospitable people - it’s super fun how the country is split into states that all have different laws and stuff, with different vibes state to state

Any chance Keir Starmer can reshuffle himself in as foreign secretary, and shuffle in another prime minister who actually has some vague idea about what they want to achieve? 🙏🤦‍♂️

Finally the heatwave has ended, and the UK is once again a bearable place to be 😍😍 If you have any UK-based collaborations, their productivity is about to increase like 10 fold

Should I use an LLM to help refine my paper writing for the ARR deadline? 🤔🤔 It will improve the paper for sure, but probably also making the tone a whole lot more annoying

I feel like the length of the ARR author rebuttals keep growing every cycle Is this a good thing for authors or reviewers that the responses can be so long? I feel like it’s a bit sub-optimal for both at the moment

Do LLMs need rationales for learning from mistakes? 🤔 When LLMs learn from previous incorrect answers, they typically observe corrective feedback in the form of rationales explaining each mistake. In our new preprint, we find these rationales do not help, in fact they hurt performance! 🧵

Bild

Thanks so much to everyone who has helped make this switch to BlueSky work. Honestly, making this switch was a pretty massive achievement, so thanks everyone for contributing ❤️❤️

This paper is really cool. They decompose NLI (and defeasible NLI) hypotheses into atoms, and then use these atoms to measure the logical consistency of LLMs. E.g. for an entailment NLI example, each hypothesis atom should also be entailed by the premise. Very nice idea 👏👏

Bild

I’m a week into my trip from Cairo to Riyadh, and wow what a place Egypt is! Honestly its been one of the funnest places I’ve travelled, and for sure I need to come back again Crossed into Aqaba (Jordan) yesterday, so now onto Saudi 🙂

I’m going away to do a bit of travelling, going overland from Cairo to Riyadh 😍 I love travelling in the Middle East so it should be interesting I’ve got that feeling of nervous excitement I always get before a trip 😬😁

Insanely jealous to everyone who has papers at #NAACL in Albuquerque! Albuquerque just sounds so exotic, and is such a cool place for a conference. No offence to Vienna, but Albuquerque sounds way more fun 😉

At some point in life I realised I actually really love travelling by train. Kind of a strange hobby, but wow it is fun 😍 Here are my top ten train journeys so far.

Imperial are hiring computing lecturers (including for AI/ML/NLP)! Here's a little thread about why you should consider applying :)

Marek Rei@marekrei.bsky.social · 2y ago

We are hiring 6 lecturers at @imperialcollegeldn.bsky.social to work on AI, ML, graphics, vision, quantum and software engineering. This includes researchers working on LLMs, NLP, generative models and text applications. Deadline 6 Jan. @imperial-nlp.bsky.social www.imperial.ac.uk/jobs/search-...

Okay genius idea to improve quality of #nlp #arr reviews. Literally give gold stars to the best reviewers, visible on open review next to your anonymously ID during review process. Here’s why it would work, and why would you should RT this fab idea:

This papers' findings about testing LLMs on NLI aligns with many of personal thoughts: 1) NLI remains a difficult task for LLMs 2) Having more few-shot examples is helpful (in my view, helping LLMs better understand class boundaries) 3) Incorrect predictions are often a result of ambiguous labels

Bild