XiaomiMiMo just released 2 SoTA models. One might be the new BEST open model yet🔥 huggingface.co/collections/... Both are: - Sparse MoE + 1M context + MIT licensed - Native omni: text/image/video/audio
Adina Yakup
@adinayakup.bsky.social
AI Research @Hugging Face 🤗 Contributing to the Chinese ML community.
HyperFlow⚡New MiniMax-H3 variant from Video Rebirth huggingface.co/videorebirth... - Open weight 8 step LoRA - 3× faster with data free self distillation - Video + stereo audio intact
DeepSeek v4.1 Flash is just another level 🤯 huggingface.co/deepseek-ai/... - Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B - Native vision merged into one endpoint - KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first model
deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Ling-3.0-flash-VL just dropped from Ant Group huggingface.co/inclusionAI/... - Native image + video: understand > reason > act > verify - 124B/5.5B active - 1M context - MIT licensed
inclusionAI/Ling-3.0-flash-VL · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
3 new open agentic models from Nex-AGI just dropped on @hf.co 🔥 Nex-N2.5: mini / Pro / Max - mini: 35B, tool-calling on 2×H100 - Pro: 397B hybrid-attention MoE, single 8×H100 node ( weights coming soon ) - Max: 1.6T, MoE - All Apache 2.0 huggingface.co/collections/...
Nex-N2.5 - a nex-agi Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
When everyone is releasing big models, OpenBMB keeps shipping small but strong one. Here is the latest one: MiniCPM5-2B 🔥 huggingface.co/openbmb/Mini... - Dense 2B - Full data pipeline open as well: web, code, agent, RL - 131 context - Apache 2.0
Ling-3.0-flash-Fin 💰 A finance tuned build of Ling-3.0-flash from Ant Group huggingface.co/collections/... - 124B/ 5.1B active - 256K context - MIT license - Finance native agent workflows - Already being quantized for llama.cpp
LingBot-Video 🎬 MoE video model built for embodied AI from Ant group huggingface.co/collections/... - 30B/3B - Apache 2.0 - Trained on web videos + 70K hours of embodied data - Tops RBench: ahead of Cosmos3/Veo 3/Seedance 1.5 pro
Agents-A1 🤖🔬 New agentic model from Shanghai AI Lab, InternScience team huggingface.co/collections/... - 35B MoE (built on Qwen3.5-35B-A3B) - Apache 2.0 - 256K context - Trained for long-horizon agent work - Includes quantized variants
Agents-A1 - a InternScience Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
LingBot Vision 👀🤖 A self-supervised vision backbone family for dense spatial perception from Ant Group huggingface.co/collections/...
Tencent just released HY3 🔥 huggingface.co/collections/... - 295B / 21B MoE , 256K context - Apache 2.0 - FP8 version included 👀 - Switchable reasoning: no think / low / high - Hallucinates half as often as before - Stable across agent frameworks
BAAI just released the Orca paper 🔥 ( weights coming soon ) huggingface.co/papers/2606.... A Multimodal Latent World Model: it learns the world itself first, and text/images/actions are just different ways to read it out 💡
Unlimited-OCR 🔥New OCR from Baidu huggingface.co/baidu/Unlimi... It can parse hundreds of pages in a single pass while maintaining stable speed. The key is R-SWA (Reference Sliding Window Attention), which keeps KV cache constant during decoding. 🏆 93% on OmniDocBench 📈 +6% over DeepSeek-OCR
Really cool to see the GLM 5.2 blog on Hugging Face 🔥 huggingface.co/blog/zai-org...
GLM-5.2: Built for Long-Horizon Tasks
A Blog post by Z.ai on Hugging Face
huggingface.co
GLM 5.2 is here 🔥 huggingface.co/collections/... ✨ 753B - 1M context ✨ MIT license ✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M ✨ AIME 2026: 99.2 (beats GPT-5.5, Claude Opus 4.8) ✨ vLLM / SGLang / Transformers supported
MiniMax-M3 just dropped 🔥 huggingface.co/MiniMaxAI/Mi... ✨ 428B / 23B active ✨ 1M context ✨ MiniMax Sparse Attention (MSA) And it’s not just weights! - paper: huggingface.co/papers/2606.... - kernel: huggingface.co/kernels/Mini... - Transformers support Love how this was released❤️
PP-OCRv6 just released by Baidu huggingface.co/collections/... ✨ tiny 1.5M / small 7.7M / medium 34.5M ✨ 48+ languages ✨ Supports handwritten/printed/industrial/screen and card text ✨ Edge friendly deployment
Macaron-V1-Preview-749B 👀 a Mixture-of-LoRA personal agent model from MindLab ✨ 744B base + 5 specialist LoRAs ✨ Generative UI as a core skill ✨ Personal agent focused ✨ 202K context ✨ MIT license huggingface.co/collections/...
Macaron-V1 - a mindlab-research Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
dots.tts 🔊 New TTS from Xiaohongshu (RedNote) huggingface.co/collections/... ✨ 2B - Apache 2.0 ✨ Fully continuous architecture (no codec tokens) ✨ 48kHz synthesis ✨ Zero-shot voice cloning
Step-3.7-Flash 🔥 New VL model from StepFun_ai huggingface.co/collections/... ✨ 198B / 11B active - MoE ✨ 256K context ✨ 3 reasoning level ✨ Up to 400 tokens/sec 🤯
Step-3.7-Flash - a stepfun-ai Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Qwen just dropped a new Text to Image benchmark + a judge model huggingface.co/collections/... ✨ 56 fine-grained evaluation facets ✨ Measures creativity beyond prompt alignment ✨ Covers storytelling/typography/design & physical logic ✨ Human aligned judge model (ρ = 0.92)
Qwen-Image-Bench - a Qwen Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
MiniCPM5-1B is an impressive release in the 1B class! huggingface.co/collections/... ✨ 1B - Apache 2.0 ✨ Hybrid reasoning with Think / No-Think modes ✨ 128K context ✨ Runs on CPU/Apple Silicon/GPU ✨ Strong eval result in the same size class
BitCPM4-CANN 🔥Native 1.58-bit LLM training system on Ascend NPUs huggingface.co/collections/... ✨ 0.5B/1B/3B/8B - Apache 2.0 ✨ 6× less memory at inference ✨ Only 4.5% training throughput overhead
BitCPM4-CANN - a openbmb Collection
Full-pipeline ternary quantized model trained on CANN.
huggingface.co
LongCat-Video-Avatar 1.5🐱 an audio driven avatar video generation framework from Meituan huggingface.co/meituan-long... ✨ Multi-character + multi-audio support ✨ Drive video from audio alone or audio + image + text ✨ 8-step inference ✨ Whisper-Large powered lip sync ✨ MIT license
Hy-MT2 🔥 New translation model family from Tencent Hunyuan ✨ 1.8B / 7B / 30B-A3B MoE ✨ Supports 33 languages ✨ 1.8B > 440MB with 1.25-bit quantization ✨ Runs on device with faster inference ✨ 1.8B outperforms some commercial APIs
HiDream-O1-Image is getting a lot of attention🔥 A few things that make it different: ✨ Interesting architecture: no VAE, no disjoint encoders, just raw pixels and text in one shared token space ✨ 8B + MIT license ✨ Native 2048×2048 ✨ Built in reasoning agent
ByteDance dropped Lance👀 huggingface.co/bytedance-re... This 3B model can generate images + edit images + generate videos + edit videos, and understand both images/videos. It's trained from scratch on only 128 A100s, and beats several 7B+ models on GenEval and VBench!
✨Big update from Baidu The PaddleOCR now supports Transformers as an inference backend 🔥 Really cool to see it becoming easier to use within the @hf.co ecosystem! huggingface.co/blog/PaddleP...
PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend
A Blog post by PaddlePaddle on Hugging Face
huggingface.co
Intern S2 preview 🔥 A scientific multimodal model from Shanghai AI Lab huggingface.co/internlm/Int...
internlm/Intern-S2-Preview · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Ant group just dropped Ring-2.6-1T 🔥 1T reasoning model, built for real world agent workflows. ✨ MIT license ✨ 128K >> 256K context (YaRN) ✨ Async RL + IcePop training architecture ✨ Dual reasoning : "high" for fast agent loops, "xhigh" for deep reasoning = Better cost/performance tradeoff 👀