Inkling (Thinking Machines): open-weights MoE (975B/41B active) with multimodal reasoning, 1M context, controllable thinking effort. NVFP4 for NVIDIA Blackwell. huggingface.co/thinkingmach... thinkingmachines.ai/news/introdu...
AIME
@aime-hq.bsky.social
AIME provides GPU cloud compute and develops AI-machines for deep learning and model inference (Multi-GPU workstations & HPC servers). We are in Berlin, Germany.
The SOOFI project released Soofi-S-Base, a new European open-source foundation model. - Hybrid Mamba-2/MoE architecture - 30B total / 3.5B active params - Sovereign base for fine-tuning Details: huggingface.co/Soofi-Projec...
Soofi-Project/Soofi-S-Base · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Zhipu AI releases GLM-5.2 (MIT license) with a 1M token context. The open-source model is optimized for coding marathons and nearly matches Opus 4.8 in benchmarks. Curious detail: It cheated during training by downloading solutions from GitHub. huggingface.co/zai-org/GLM-...
zai-org/GLM-5.2 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Nex-AGI introduces Nex-N2: an open-source agent model for complex workflows. Core: Agentic Thinking for adaptive reasoning & tool use. Available in Pro & Mini versions on Hugging Face, delivering top-tier performance. github.com/nex-agi/Nex-N2
GitHub - nex-agi/Nex-N2
Contribute to nex-agi/Nex-N2 development by creating an account on GitHub.
github.com
NVIDIA's Nemotron-3-Ultra-550B-A55B-NVFP4 is now on Hugging Face: 550B total params (55B active), 1M token context, NVFP4 quantization, 11 languages. OpenMDW-1.1 license permits commercial use. Optimized for agents, RAG & long-context inference. huggingface.co/nvidia/NVIDI...
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Moonshot AI open-sources Kimi-Audio-7B: a unified foundation model for audio understanding, generation, and conversation. github.com/MoonshotAI/K...
GitHub - MoonshotAI/Kimi-Audio: Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation - MoonshotAI/Kimi-Audio
github.com
EAGLE-3 boosts speculative decoding with direct token prediction + multi-layer feature fusion RedHatAI now offers a ready-to-use EAGLE-3 speculator for Gemma-4-26B-A4B-it on HuggingFace huggingface.co/RedHatAI/gemma-4-26B-A4B-it-speculator.eagle3
RedHatAI/gemma-4-26B-A4B-it-speculator.eagle3 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
TurboQuant compresses LLM KV-caches via 3-bit key & 2/4-bit value quantization, cutting memory use by up to 4.4×. Enables longer contexts & higher throughput under GPU constraints. Open-source (GPL-3.0). github.com/0xSero/turboquant #LLM #Inference #Quantization #vLLM #OpenSource
GitHub - 0xSero/turboquant: TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration - 0xSero/turboquant
github.com
Xiaomi releases MiMo-V2.5 (310B params) and MiMo-V2.5-Pro (1.02T params) as open-source models. Both support 1M token context, hybrid attention, and native multimodal capabilities. Available via API and open weights for agentic AI and code generation tasks. huggingface.co/XiaomiMiMo/M...
XiaomiMiMo/MiMo-V2.5 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Robbyant released LingBot-Map A new feed-forward 3D foundation model for streaming-based reconstruction – now open source. ✨ Geometric Context Transformer ⚡ ~20 FPS streaming inference 🏆 SOTA results on multiple benchmarks 💻 github.com/Robbyant/lin...
GitHub - Robbyant/lingbot-map: A feed-forward 3D foundation model for reconstructing scenes from streaming data
A feed-forward 3D foundation model for reconstructing scenes from streaming data - Robbyant/lingbot-map
github.com
Cohere Transcribe is here – a new open-weights ASR model (Apache 2.0) that's up to 4× faster than Whisper Large v3 with a best-in-class 5.42 WER. 🚀 🔹 2B params | 14 languages | Conformer architecture 🔹 Easy 🤗 Transformers integration 🔗 huggingface.co/CohereLabs/c...
CohereLabs/cohere-transcribe-03-2026 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
🎙️ Mistral AI's Voxtral TTS is live: 4B params, 9 languages, ~70ms latency, zero-shot voice cloning. Open weights on HF. huggingface.co/mistralai/Vo...
mistralai/Voxtral-Mini-4B-Realtime-2602 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Hume AI just dropped TADA (Text-Acoustic Dual Alignment) – an open-source framework for expressive speech generation. ✨ 1:1 text-audio token alignment ✨ Precise prosody & timing control ✨ Multilingual (DE/EN/FR/ES/JA/AR) ✨ Built on Llama 3.2 (1B/3B) 🔗 github.com/HumeAI/tada
GitHub - HumeAI/tada: Open Source Speech Language Model
Open Source Speech Language Model. Contribute to HumeAI/tada development by creating an account on GitHub.
github.com
SakanaAI's new Text-to-LoRA (T2L) uses a hypernetwork to generate task-specific LoRAs from simple text descriptions—no expensive fine-tuning required. ✅ Compresses 100s of adapters ✅ Generalizes to unseen tasks ✅ ICML 2025 Paper & Code: github.com/SakanaAI/tex...
GitHub - SakanaAI/text-to-lora: Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input
Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input - SakanaAI/text-to-lora
github.com
Voicebox: open-source, locally run TTS studio—no cloud, no subscriptions. ✅ Powered by Qwen3-TTS for expressive voice cloning ✅ Multi-track editor + inline audio editing ✅ Tauri/Rust app: 10× smaller than Electron ✅ MIT license, full privacy github.com/jamiepine/vo...
GitHub - jamiepine/voicebox: The open-source voice synthesis studio powered by Qwen3-TTS.
The open-source voice synthesis studio powered by Qwen3-TTS. - jamiepine/voicebox
github.com
AKI.IO is now live: Curated open-source and open-weight models such as #MiniMax M2.5, #Apertus 70B, #Qwen Image Edit and many more as an API – hosted entirely in European data centers w/o hyperscalers. Happy to get your feedback! The playground is open, API key via free registration at aki.io
Home - AKI.IO
Token-based access to leading open-source AI models on EU infrastructure. Evaluate, build and scale your AI product without self-hosting or vendor lock-in.
aki.io
Qwen3.5 is out: Alibaba's open-weight series built for agentic AI with native multimodality. ✅ Flagship: 397B total / 17B active params (MoE) ✅ 1M-token context → 2h audio/video in one pass ✅ 60% cheaper, 8× more efficient than predecessor ✅ MIT license, full open weights github.com/QwenLM/Qwen3.5
GitHub - QwenLM/Qwen3.5: Qwen3.5 is the large language model series developed by Qwen team, Alibaba Cloud.
Qwen3.5 is the large language model series developed by Qwen team, Alibaba Cloud. - QwenLM/Qwen3.5
github.com
Microsoft's VibeVoice-ASR transcribes 60-minute audio in a single pass—no chunking needed. ✅ 9B params, 64K-token context ✅ ASR + speaker diarization + timestamps in one inference ✅ MIT license, fully open source A leap for meeting/podcast transcription 👇 huggingface.co/microsoft/Vi...
microsoft/VibeVoice-ASR · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
MiniMax M2.5 is out: a frontier model optimized via massive RL for agentic workflows. Forge RL Framework enables near-linear scaling across 10k+ real agent scenarios. Weights not yet public—MiniMax historically releases them later. www.minimax.io/news/minimax...
MiniMax M2.5: 更快更强更智能,为真实世界生产力而生
minimax.io
GLM-5 is Z.ai's new MoE flagship (744B total/40B active) built for agentic engineering. ✅ #1 open-source on Vending Bench 2 ✅ Closes gap with Claude Opus on CC-Bench-V2 ✅ DeepSeek Sparse Attention for efficient 200K context ✅ Apache 2.0 license, commercial use allowed github.com/zai-org/GLM-5
GitHub - zai-org/GLM-5: GLM-5: From Vibe Coding to Agentic Engineering
GLM-5: From Vibe Coding to Agentic Engineering. Contribute to zai-org/GLM-5 development by creating an account on GitHub.
github.com
ACE-Step v1.5 is out: an open-source music generation model that runs locally with <4 GB VRAM. 8 diffusion steps → full songs in ~2s (A100) 4-min tracks with lyrics, 50+ languages MIT license, full training code included A leap for accessible, commercial-grade audio AI 👇 github.com/ace-step/ACE...
GitHub - ace-step/ACE-Step-1.5: The most powerful local music generation model that outperforms most commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
The most powerful local music generation model that outperforms most commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices. - ace-step/ACE-Step-1.5
github.com
Qwen3-Coder-Next is out: an open-weight MoE model (80B total / 3B active params) built for agentic coding workflows. ✅ 256K context length ✅ Tool-use & multi-step reasoning optimized ✅ Apache 2.0 license for local/dev use Great step for open coding agents 👇 huggingface.co/Qwen/Qwen3-C...
Qwen/Qwen3-Coder-Next · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Qwen3-ASR drops: open-source speech recognition that transcribes speech, music, and singing across 52 languages — with accuracy rivaling GPT-4o and Gemini. 1.7B & 0.6B variants. Unified streaming/offline inference. Apache 2.0. github.com/QwenLM/Qwen3...
GitHub - QwenLM/Qwen3-ASR: Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detectio...
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp predicti...
github.com
DeepSeek OCR 2 is a 3B VLM that reads documents like humans do. "Visual Causal Flow" dynamically reorders tokens by semantic meaning, not left-to-right, unlocking 91.09% accuracy on complex layouts. Fully open source (Apache 2.0). github.com/deepseek-ai/...
GitHub - deepseek-ai/DeepSeek-OCR-2: Visual Causal Flow
Visual Causal Flow. Contribute to deepseek-ai/DeepSeek-OCR-2 development by creating an account on GitHub.
github.com
Kimi released K2.5 — a native multimodal model trained on 15T visual-text tokens that generates full interactive UIs from prompts and orchestrates 100-agent swarms for complex tasks. 4.5× faster execution, 59% productivity boost. Open weights available now. www.kimi.com/blog/kimi-k2...
Kimi K2.5: Visual Agentic Intelligence | Technical Report
Kimi K2.5 defines Visual Agentic Intelligence. Trained on 15T tokens, it introduces SOTA visual coding and autonomous agent swarm. Read the full tech report.
kimi.com
Alibaba released Qwen3-TTS, a new text-to-speech model with discrete multi-codebook LM architecture under Apache license. Features 97ms synthesis latency, 3-second voice cloning, and 10-language support including German. Available on Hugging Face and ModelScope. github.com/QwenLM/Qwen3...
GitHub - QwenLM/Qwen3-TTS: Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice...
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice...
github.com
NVIDIA just dropped PersonaPlex - a speech-to-speech model that lets you control AI personas through text prompts AND voice conditioning! github.com/NVIDIA/perso...
GitHub - NVIDIA/personaplex: PersonaPlex code.
PersonaPlex code. Contribute to NVIDIA/personaplex development by creating an account on GitHub.
github.com
Z.AI just released GLM-4.7-Flash - a 30B-A3B MoE model that dominates the 30B parameter class! Perfect balance of power & efficiency for enterprise deployment. Supports vLLM, SGLang & native tool integration. huggingface.co/zai-org/GLM-...
zai-org/GLM-4.7-Flash · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
FLUX.2 [klein] 9B Base released by @BlackForestLabs 🔥 9B-parameter undistilled foundation model ⚡ End-to-end inference in <1 second 💻 Runs on RTX 4090+ (~29GB VRAM) 🎨 Perfect for fine-tuning & research Non-commercial license only huggingface.co/black-forest...
black-forest-labs/FLUX.2-klein-base-9B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Z.AI released GLM-Image, an innovative image generation model that establishes new benchmarks in specific application areas through its hybrid architecture. github.com/zai-org/GLM-...
GitHub - zai-org/GLM-Image: GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation. - zai-org/GLM-Image
github.com