Jonathan Frankle

@jfrankle.com

Chief AI Scientist at Databricks. Founding team at MosaicML. MIT/Princeton alum. Lottery ticket enthusiast. Working on data intelligence.

This is how it's done. A strong and principled response by WilmerHale to the illegal Executive Order attack - a form of attempted government intimidation declared unconstitutional by a federal judge. This is how to guard the rule of law.

Bild

The hardest part about finetuning is that people don't have labeled data. Today, @databricks.bsky.social introduced TAO, a new finetuning method that only needs inputs, no labels necessary. Best of all, it actually beats supervised finetuning on labeled data. www.databricks.com/blog/tao-usi...

TAO: Using test-time compute to train efficient LLMs without labeled data

LIFT fine-tunes LLMs without labels using reinforcement learning, boosting performance on enterprise tasks.

databricks.com

Excited to share our work with friends from MIT/Google on Learned Asynchronous Decoding! LLM responses often contain chunks of tokens that are semantically independent. What if we can train LLMs to identify such chunks and decode them in parallel, thereby speeding up inference? 1/N

We're probably a little too obsessed with zero-shot retrieval. If you have documents (you do), then you can generate synthetic data, and finetune your embedding. Blog post lead by @jacobianneuro.bsky.social shows how well this works in practice. www.databricks.com/blog/improvi...

Improving Retrieval and RAG with Embedding Model Finetuning

Fine-tune embedding models on Databricks to enhance retrieval and RAG accuracy with synthetic data—no manual labeling required.

databricks.com

In case it is not clear from my reposts, the Trump administration is engaged in an illegal AND unconstitutional to seize power over the federal government away from Congress and the courts. "Pausing" payment on the government's bills is just one part of it, but it is among the worst.

All the more convinced that the markets don't understand AI. Both the irrational hype and the irrational pessimism. DeepSeek is incredibly bullish for GPU sales...

🧵 Super proud to finally share this work I led last quarter - the @databricks.bsky.social Domain Intelligence Benchmark Suite (DIBS)! TL;DR: Academic benchmarks ≠ real performance and domain intelligence > general capabilities for enterprise tasks. 1/3

Bild

Reflections on NeurIPS: There's always a big theme people seem to be preoccupied with. This year, it was the continuation of scaling/progress. Will it continue? What will the next generation of models hold? I even got to sass Dylan Patel (not on bsky) over it. Here are my personal thoughts 🧵

We just released the Databricks synthetic evals SDK! 🎉 We’ve found that synthesizing evals is a great way to hill climb your AI system before you’re able to get labels from domain experts. We’ve also recently release a new diff UI that lets you both qualitatively and quantitatively view results!

Bild

Interested in massively reducing the toil of code maintenance? Permanence AI is at #NeurIPS2024 this week! Tyler Holloway and Ethan Elenberg will be presenting work at the ML For Systems workshop, and our Founder/CEO Joseph Hackman is attending the conference as well. Let's catch up! #AI #ML

Mat is not on 🦋—posting on his behalf! It's time to revisit common assumptions in IR! Embeddings have improved drastically, but mainstream IR evals have stagnated since MSMARCO + BEIR. We ask: on private or tricky IR tasks, are rerankers better? Surely, reranking many docs is best?

A plot showing that reranking improves recall as we increase the number of reranked docs, but with increasing docs we diminishing returns and eventually a performance dip.

Incredibly excited for the team at @datologyai.bsky.social for huge progress on data efficiency for CLIP model training. I'm a big fan of lines that go up and to the left, and this is a really substantial improvement.

Matthew Leavitt@leavittron.bsky.social · 2y ago

🧵We’ve spent the last few months at @datologyai.bsky.social building a state-of-the-art data curation pipeline and I’m SO excited to share our first results: we curated image-text pretraining data and massively improved CLIP model quality, training speed, and inference efficiency 🔥🔥🔥