Mark Ibrahim

@markibrahim.bsky.social

Researching the dark arts of deep learning at Meta's FAIR (Fundamental AI Research) Lab https://markibrahim.me/

We introduce, Common-O, a new multimodal benchmark for hallucination when reasoning across scenes. We find leading multimodal LLMs can reliably identify objects, yet hallucinate when reasoning across scenes. 🧵1/3

Bild

Open-weights for our Llip multimodal vision-language model led by @lavoiems.bsky.social are public! LLIP proposes new pre-training objective to capture the many ways to describe an image leading to strong performance across a suite of 22-zero shot benchmarks. bsky.app/profile/lavo...

Samuel Lavoie@lavoiems.bsky.social · last yr.

The code and model weights for Llip are finally out! I hope you will find this model useful! Paper: arxiv.org/abs/2405.00740 Code: github.com/facebookrese... Models: - ViT-G: huggingface.co/lavoies/llip... - ViT-B: huggingface.co/lavoies/llip...

A good language model should say “I don’t know” by reasoning about the limits of its knowledge. Our new work AbstentionBench carefully measures this overlooked skill in an open-codebase others can build on! We find frontier reasoning degrades models’ ability to know when NOT to answer. 🧵1/2

Bild