Andi

@andimara.bsky.social

Multimodal research @huggingface

New Blog📖✨: nanoVLM: The simplest way to train your own Vision-Language Model in pure PyTorch explained step-by-step! Easy to read, even easier to use. Train your first VLM today!

Bild

🚀 We just dropped SmolDocling: a 256M open-source vision LM for complete document OCR! 📄✨ Lightning fast, process a page in 0.35 sec on consumer GPU using < 500MB VRAM ⚡ SOTA in document conversion, beating every competing model we tested (including models 27x more params) 🤯 But how? 🧶⬇️

Bild

Extremely bullish on @CohereForAI's Aya Vision (8B & 32B) - new SOTA open-weight VLMs - 8B wins up to 81% of the time in its class, better than Gemini Flash - 32B beats Llama 3.2 90B! - Integrated on @hf.co from Day 0! Check out their blog! huggingface.co/blog/aya-vis...

A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Fuck it, today we're open-sourcing the codebase used to train SmolVLM from scratch on 256 H100s 🔥 Inspired by our team's effort to open-source DeepSeek's R1, we are releasing the training and evaluation code on top of the weights 🫡 Now you can train any SmolVLM—or create your own custom VLMs!

Bild

Introducing the smollest VLMs yet! 🤏 SmolVLM (256M & 500M) runs on <1GB GPU memory. Fine-tune it on your laptop and run it on your toaster. 🚀 Even the 256M model outperforms our Idefics 80B (Aug '23). How small can we go? 👀

Bild

We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute 🔥 How? By combining step-wise reward models with tree search algorithms :) We're open sourcing the full recipe and sharing a detailed blog post 👇

Bild

Let's go! We are releasing SmolVLM, a smol 2B VLM built for on-device inference that outperforms all models at similar GPU RAM usage and tokens throughputs. SmolVLM can be fine-tuned on a Google collab and be run on a laptop! Or process millions of documents with a consumer GPU!

Bild

It's pretty sad to see the negative sentiment towards Hugging Face on this platform due to a dataset put by one of the employees. I want to write a small piece. 🧵 Hugging Face empowers everyone to use AI to create value and is against monopolization of AI it's a hosting platform above all.

Let's go! We are releasing SmolVLM, a smol 2B VLM built for on-device inference that outperforms all models at similar GPU RAM usage and tokens throughputs. SmolVLM can be fine-tuned on a Google collab and be run on a laptop! Or process millions of documents with a consumer GPU!

Bild