merve

@merve.bsky.social

proud mediterrenean 🧿 open-sourceress at hugging face 🤗 multimodality, zero-shot vision, vision language models, transformers

I'm so hooked on @hf.co Inference Providers (specifically Qwen2.5-VL-72B) for multimodal agentic workflows with smolagents 🥹 get started ⤵️ > filter models provided by different providers > test them through widget or Python/JS/cURL

Hello friends 👋🏼 If visit Turkey this summer, know that millions of Turkish people are doing a boycott, once a week not buying anything and rest of the week only buying necessities if you have plans, here's a post that summarizes where you should buy stuff from www.instagram.com/share/BADrkS...

Login • Instagram

Welcome back to Instagram. Sign in to check out what your friends, family & interests have been capturing & sharing around the world.

instagram.com

DO NOT SLEEP ON THIS MODEL Kimi-VL-A3B-Thinking is the first ever capable open-source reasoning VLM with MIT license ❤️ > it has only 2.8B activated params 👏 > it's agentic 🔥 works on GUIs > surpasses gpt-4o I've put it to test (see below ⤵️) huggingface.co/spaces/moons...

Bild

InternVL3 is out 💥 > 7 ckpts with various sizes (1B to 78B) > Built on InternViT encoder and Qwen2.5VL decoder, improves on Qwen2.5VL > Can do reasoning, document tasks, extending to tool use and agentic capabilities 🤖 > easily use with Hugging Face transformers 🤗 huggingface.co/collections/...

Bild

All the multimodal document retrieval models (ColPali, DSE et al) are now under visual document retrieval at @hf.co 📝🤗 take your favorite VDR model out for multimodal RAG 🤝

Bild

Introducing the smollest VLMs yet! 🤏 SmolVLM (256M & 500M) runs on <1GB GPU memory. Fine-tune it on your laptop and run it on your toaster. 🚀 Even the 256M model outperforms our Idefics 80B (Aug '23). How small can we go? 👀

Bild

ByteDance just dropped SA2VA: a new family of vision LMs combining Qwen2VL/InternVL and SAM2 with MIT license 💗 The models are capable of tasks involving vision-language understanding and visual referrals (referring segmentation) both for images and videos ⏯️

Bild

you can now stay up-to-date with big AI research labs' updates on @hf.co easily over org activity page 🥹 I have been looking forward to this feature as I felt most back to back releases are overwhelming and I tend to miss out 🤠

Aya by Cohere For AI can now see! 👀 C4AI community has built Maya 8B, a new open-source multilingual VLM built on SigLIP and Aya 8B 🌱 works on 8 languages! 🗣️ The authors extend Llava dataset using Aya's translation capabilities with 558k examples! works very well ⬇️ huggingface.co/spaces/kkr51...

screenshot of model conversation turns

VLMs go MoE ✨ DeepSeek AI dropped three new commercially permissive vision LMs based on SigLIP encoder and their DeepSeek-MoE decoder 🐳 the models come in 1.0B, 2.8B and 4.5B active params 🥹 models seem to catch up with state-of-the-art with less active parameters! huggingface.co/collections/...

Bild