Itay Itzhak @ COLM 🍁

@itay-itzhak.bsky.social

NLProc, deep learning, and machine learning. Ph.D. student @ Technion and The Hebrew University. https://itay1itzhak.github.io/

We vibe-tested our own paper, but the reviewers gave us the actual metrics: "From Feelings to Metrics" has been accepted to #COLM2026! 🎉 We turn "this model just feels better" into structured, user-aware evaluation! #COLM people - let’s grab a coffee and vibe-test some research in person! ☕️🌴

Bild
Itay Itzhak @ COLM 🍁@itay-itzhak.bsky.social · 3mo ago

Ever used a top-ranked LLM that just... felt wrong for you? You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation? New paper: "From Feelings to Metrics" 🧵

Ever used a top-ranked LLM that just... felt wrong for you? You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation? New paper: "From Feelings to Metrics" 🧵

Bild

Had a blast at CoLM! It really was as good as everyone says, congrats to the organizers 🎉 This week I’ll be in New York giving talks at NYU, Yale, and Cornell Tech. If you’re around and want to chat about LLM behavior, safety, interpretability, or just say hi - DM me!

BildBild

At #ACL2025 and not sure what to do next? GEM 💎² is the place to be for awesome talks on the future of LLM evaluation. Come hear @GabiStanovsky, @EliyaHabba, @LChoshen and others rethink what it means to actually evaluate LLMs beyond accuracy and vibes. Thursday @ Hall C!

Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵

Bild

New paper alert! Curious how small prompt tweaks impact LLM accuracy but don’t want to run endless inferences? We got you. Meet DOVE - a dataset built to uncover these sensitivities. Use DOVE for your analysis or contribute samples -we're growing and welcome you aboard!

Eliya Habba@eliyahabba.bsky.social · last yr.

Care about LLM evaluation? 🤖 🤔 We bring you ️️🕊️ DOVE a massive (250M!) collection of LLMs outputs  On different prompts, domains, tokens, models... Join our community effort to expand it with YOUR model predictions & become a co-author!

1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇

Bild

🚨🚨 New preprint 🚨🚨 Ever wonder whether verbalized CoTs correspond to the internal reasoning process of the model? We propose a novel parametric faithfulness approach, which erases information contained in CoT steps from the model parameters to assess CoT faithfulness. arxiv.org/abs/2502.14829

Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps

When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite mu...

arxiv.org

We usually blame hallucinations on uncertainty or missing knowledge. But what if I told you that LLMs hallucinate even when they *know* the correct answer - and they do it with *high certainty* 🤯? Check out our new paper that challenges assumptions on AI trustworthiness! 🧵👇

Adi Simhi@adisimhi.bsky.social · last yr.

🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov