Sergio Paniego

@sergiopaniego.bsky.social

AI PhD. Technology enables us to be more human. 🏳️‍🌈

🧠 Following Hugging Face's blog on scaling test-time compute with open models—letting models "think longer," inspired by OpenAI & DeepMind—I created a recipe to extend inference time for Instruct LLMs, tackling harder tasks like complex math problems. Links below 👇

Scaling test-time compute with open models diagram

I’m a big fan of smol models—compact, efficient, and perfect for inference/training on limited resources. Even better when they’re multimodal! 🤏✨ I explored fine-tuning SmolVLM, a multimodal smol model using TRL with SFT and DPO, creating 2 hands-on projects! 🔗Links below👇

Bild

💡I've been exploring how to go smol with multimodal RAG. I've created a project using SmolVLM and ColSmolVLM to create a multimodal RAG that can run on Colab's free tier. Featuring: 🤏👀 SmolVLM (VLM) 🤏📚ColQwen2 (Doc Retrieval) ⚙️ Runs in Colab's free-tier GPU Link below

Bild

💡 New Multimodal RAG Recipe with Re-Ranking 💡 I explored how to enhance a multimodal RAG pipeline by integrating a re-ranker! Featuring: ✨ Qwen2-VL-7B (VLM) 📚 ColQwen2 (Doc Retrieval) 🔍 MonoQwen2 (Re-ranking) 🔥 Optimized for consumer GPUs with quantized VLMs. Link below:

Bild

✨ Gave a talk on autonomous driving today to undergrad students! We covered everything from definitions to real-world examples, plus cutting-edge concepts like Generative World Models and Vision-Language Models (VLMs). Exciting future ahead! 🚗💡

Bild

I've been exploring the latest Llama 3.2 releases and working on a couple of projects you may find interesting: 1️⃣ Understanding tool calling with Llama 3.2 🔧 2️⃣ Using Text Generation Inference (TGI) with Llama models 🦙 (links in the next post)

Bild

TRL is a cornerstone of LLM post training and imo it's the default to learn. There are great alternatives like Unsloth, Axolotl, and AutoTrain. But if you want a daily drive that does experimentation to production, it's TRL. 🧵 these community notebooks guide you through TRL's core: