ollama {bot}

@ollamabot.bsky.social

Unofficial mirror account of https://x.com/ollama from Twitter https://ollama.com/

DeepSeek-V4-Flash-0731 is Ollama's fastest growing model ever in token usage. We are scaling capacity in US & Europe. On Ollama, this model runs with high performance (100tps+) and zero data retention. Your data stays yours. ollama run deepseek-v4-flash:0731-cloud

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama launch claude --model deepseek-v4-flash:0731-cloud

Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access. (1/2)

The demand for GLM-5.2 and other frontier-level open models has been surging on Ollama's cloud. We're adding capacity in anticipation for some very large models next week! To make sure Ollama's infrastructure stays fast we're temporarily pausing new subscriptions on Ollama's Max plan. (1/2)

Ollama 0.32.1 includes significant improvements to Gemma 4's tool calling, making it much more reliable in coding agents! To try Gemma 4 26B with Pi, run: ollama launch pi --model gemma4:26b With Ollama's MLX engine for maximum performance, run: ollama launch pi --model gemma4:26b-mlx

Quote Tweet: https://twitter.com/i/status/2077449152062247219

Open models are already being used in the enterprise. Over 85% of the Fortune 500 companies already use Ollama to fulfill specific tasks. @jmorgan Why open models and Ollama? Ownership. Open models are yours to keep, customize, and optimize. Affordable. (1/2)

Big day for Ollama! When we started, open models and the open source AI ecosystem were in their early days with few believers. Our belief in open source has never wavered. (1/2)

Ollama is here to accelerate open models!

GLM-5.2 on Ollama's cloud just got more capacity in US & Europe! Ollama's cloud for GLM 5.2 consistently delivers between 80 to 120 output tokens per second, even during peak hours, compared to 30 to 40 tok/s on other providers. Use GLM-5.2 with Claude Code: ollama launch claude (1/2)

Ollama's Cloud

Gemma 4 is now nearly 90% faster on Apple Silicon with Ollama using MLX! The speedup comes from improved multi-token prediction (MTP), now on by default for Gemma 4, with more models to come. (1/2)

Image from Twitter

Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model ornith For the more capable 35B model, use: ollama launch claude --model ornith:35b

Quote Tweet: https://twitter.com/i/status/2070148887067963854

GLM 5.2 on Ollama's cloud just doubled GPU capacity to handle the volume of usage! This is all US based, and running on NVIDIA B300 Blackwell GPUs. We believe privacy matters! Let's go open models! ❤️