WebBrain

@webbrain-one.bsky.social

The Open-Source AI Browser Agent https://webbrain.one https://github.com/webbrain-one/webbrain (pls star!)

DeepSeek V4 Flash is no longer text-only 👀 In our internal benchmarks, it delivered significantly better price-performance than others in its class. We’ve added vision for screen-level understanding, capabilities needed for WebBrain Available on HF: huggingface.co/webbrain-one... #deepseek #ai

webbrain-one/DeepSeek-V4-Flash-0731-Vision-NVFP4 · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

We tested 13 OpenRouter models across 1,300 browser-planning calls. DeepSeek V4 Flash led consensus, Gemini 3.6 Flash delivered 100/100 valid calls, and Kimi K3 agreed with Sonnet 5 on 86% of tool choices. Full benchmark: www.webbrain.one/blog/america...

Two AI gaps are narrowing: our thirteen-model planner benchmark

Thirteen OpenRouter models, 1,300 comparable calls, no reference-model judge: consensus quality, latency, modality, architecture, parameter count, and real test cost.

webbrain.one

Benchmarked Poolside's new Laguna S 2.1 (118B-A8B) as a browser-agent planner. 71% Sonnet alignment. Near Hy3/MiniMax M3 but below them, and below Gemma 4 31B QAT / Qwen 3.6 27B (~77%). Best US open-weight in its class, half the size of M3 👇 www.webbrain.one/blog/poolsid... #offlinellm #poolside

Laguna S 2.1 reaches 71% in WebBrain with a high-reasoning request

Poolside Laguna S 2.1 scored 65% with default reasoning and 71% with a high-effort request. It is fast and cheap, but the second run also increased no-tool outputs.

webbrain.one

kids, don't do this at home. we handed our AI browser agent the keys and it: – swiped on Bumble on our behalf – read the NYT past the paywall – auto-replied to strangers at night – ripped a YouTube video – beat a CAPTCHA to sign itself up webbrain.one

Ornith 35B claims to beat Gemma4 31B and Qwen3.6 35B on many fronts. We ran it through WebBrain's frozen browser-agent planner benchmark. Result: Ornith is solid — edges Qwen 3.6 on alignment — but Gemma4 remains the better option. webbrain.one/blog/ornith-... #Ornith #Gemma4 #Qwen

Ornith-1.0-35B enters WebBrain's frozen planner benchmark

Ornith's coding-agent claims are impressive. In WebBrain's browser-agent planner benchmark, it lands near the hosted/local planner tier but does not overtake Gemma 4 31B.

webbrain.one