Kyle Corbitt

@corbt.com

If you're fine-tuning LLMs, Gemma 3 is the new 👑 and it's not close. Gemma 3 trounces Qwen/Llama models at every size! - Gemma 3 4B beats 7B/8B competition - Gemma 3 27B matches 70B competition Vision benchmarks soon!

Bild

Training models with RL subjectively feels much more like gardening than engineering. You do your best to set the right conditions, provide the right inputs... and then wait and watch what grows. Very rewarding/magical feeling when it works!

Big news: we've figured out how to train models 80-90% cheaper than before. Cheaper than renting your own GPUs. Cheaper than any other service. And 0 quality regression. Super proud of the team on this one. New pricing is now live!

Bild

This holiday season I am legitimately grateful that my kids are all <8 and not 16+. I have no idea what career prep advice I'd give someone in this moment. We're in for a ride.

Helpful intuition that folks new to LLMs may not know: if you have a lot of data, small models are often just as good as much, much larger ones for tasks like classification and information extraction. Here I compare a 1B vs 8B on a hard classification task, and I bet you can't tell which is which!

Bild

OpenAI's Reinforcement Fine-Tuning (RFT) is far more data efficient than SFT—can generalize from 10-20 labeled examples. Huge deal bc as compute costs drop to 0, the pain of gathering high-quality training data is the biggest barrier to deploying AI. RFT needs much less of it!

SUPER PUMPED to announce that Gemini fine-tuning is available to all OpenPipe users! Gemini Flash provides the lowest cost fine-tuning of any model in its quality class. Comparable to gpt-4o-mini, but 4x cheaper inference and FREE fine-tuning!

Bild

One of the new features I'm most excited about at OpenPipe is "criteria distillation". This allows you to distill an expensive LLM-as-judge criteria into a super fast, cheap, low-latency reward model that approximates the LLM-as-judge's outputs. DM for access!

Amazon's Nova models have excellent price/perf ratio. We'd love to support them, but to deploy fine-tuned versions you need to purchase "provisioned throughput", which costs $100/hr/model. 😬 Putting out the bat signal—if you know someone at AWS Bedrock, pls put me in contact!

BildBild

Ok I am terrible at sharing product updates here, but we now support Llama 3.2 1B and 3B (the best small LLMs) as well as Qwen 2.5 72B and 32B Coder (the best open general and code-specific models) on OpenPipe!

Bild

Kinda feels like the product engineer and IC PM roles are quickly converging. A really good AI-enabled SWE can produce the same output as a former team of 5 SWE+1 PM. Are AI-native companies still hiring IC PMs who don't code?

What is the current SOTA on language autoencoders? Can you run lossy compression on a 20K-word Wikipedia article to give you an archive that's just a few KB in size, but decompresses into text semantically indistinguishable from the original?

This may become an official Qwen-stan account. ✅ Open source SOTA on code ✅ Open source SOTA in general for 14B+ ✅ Almost SOTA <14B ✅ Works great for LM, RM and classification tasks ✅ SOTA open source multimodal

Qwen 2.5 Coder 32B is a 🐐 ✅ Benchmarks at or above GPT-4 and Claude 3.5 ✅ Subjectively feels fantastic for code (been trying it) ✅ Fine-tunable on your own data on OpenPipe!

Bild

Last week Huggingface released "SmolLM v2," several <2B models designed for edge deployment. Interested in how they perform when fine-tuned? You're in luck! We've compared their performance with other edge models. (Spoiler: Qwen remains the champion 👑)

Bild