I noticed that I'm not using bsky much anymore. Not sure why, vibes. Anyways, someone noticing that DeepSeek refuses to answer *anything* about Xi Jinping, even the question whether he exists at all, triggered me writing a short snippet on safety fine-tuning: lb.eyer.be/s/safety-sft...
Lucas Beyer (bl16)
@giffmana.ai
Researcher (OpenAI. Ex: DeepMind, Brain, RWTH Aachen), Gamer, Hacker, Belgian. Anon feedback: https://admonymous.co/giffmana 📍 Zürich, Suisse 🔗 http://lucasb.eyer.be
First candidate for banger of the year appeared, only 2 days in:
OpenAI skips o2, previews o3 scores, and they're truly crazy. Huge progress on the few benchmarks we think are truly hard today. Including ARC AGI. Rip to people who say any of "progress is done," "scale is done," or "llms cant reason" 2024 was awesome. I love my job.
A post by @cloneofsimo on Twitter made me write up some lore about residuals, ResNets, and Transformers. And I couldn't resist sliding in the usual cautionary tale about small/mid-scale != large-scale. Blogpost: lb.eyer.be/s/residuals....
Sup cats
Good morning Vancouver! Things are different here: this guy is alone, chonky, and not scared at all, I was more scared of him towards the end lol. Also look at that … industrialization
Good morning! On my way to NeurIPS, slightly sad to leave this beautiful place and my family for the week, but also excited to meet many new and old friends at NeurIPS!
Good morning! On my way to NeurIPS, slightly sad to leave this beautiful place and my family for the week, but also excited to meet many new and old friends at NeurIPS!
One of the best tutorials for understanding Transformers! 📽️ Watch here: www.youtube.com/watch?v=bMXq... Big thanks to @giffmana.ai for this excellent content! 🙌
[M2L 2024] Transformers - Lucas Beyer
YouTube video by Mediterranean Machine Learning (M2L) summer school
youtube.com
Attending #NeurIPS2024? If you're interested in multimodal systems, building inclusive & culturally aware models, and how fractals relate to LLMs, we've 3 posters for you. I look forward to presenting them on behalf of our GDM team @ Zurich & collaborators. Details below (1/4)
The fourth nice thing we* have for you this week: PaliGemma 2. It’s also a perfect transition: this v2 was carried a lot more by @andreaspsteiner.bsky.social André and Michael than by us. Crazy new sota tasks! Interesting res vs LLM size study! Better OCR! Less hallucination!
🚀🚀PaliGemma 2 is our updated and improved PaliGemma release using the Gemma 2 models and providing new pre-trained checkpoints for the full cross product of {224px,448px,896px} resolutions and {3B,10B,28B} model sizes. 1/7
OpenAI is coming to Switzerland, and the founding team of the new Zurich office is nothing short of stellar🇨🇭⭐️ What a fantastic win for the European AI landscape 🇪🇺 Congrats to @giffmana.ai @kolesnikov.ch @xzhai.bsky.social for that move, and to OpenAI for making these hires!
So, now that our move to OpenAI became public, @kolesnikov.ch @xzhai.bsky.social and I are drowning in notifications. I read everything, but may not reply. Excited about this new journey! 🚀 Quick FAQ thread...
@francois.fleuret.org hey, can you buy me a few of tomorrow's Le Temps, if the news about us is printed in it? I'll pay you in beers at neurips.
So, now that our move to OpenAI became public, @kolesnikov.ch @xzhai.bsky.social and I are drowning in notifications. I read everything, but may not reply. Excited about this new journey! 🚀 Quick FAQ thread...
Ok, it is yesterdays news already, but good night sleep is important. After 7 amazing years at Google Brain/DM, I am joining OpenAI. Together with @xzhai.bsky.social and @giffmana.ai, we will establish OpenAI Zurich office. Proud of our past work and looking forward to the future.
Ok, it is yesterdays news already, but good night sleep is important. After 7 amazing years at Google Brain/DM, I am joining OpenAI. Together with @xzhai.bsky.social and @giffmana.ai, we will establish OpenAI Zurich office. Proud of our past work and looking forward to the future.
JetFormer: An Autoregressive Generative Model of Raw Images and Text @mtschannen.bsky.social @asusanopinto.bsky.social @kolesnikov.ch tl;dr: VQGAN quality w/o VQGAN, but optimizing pixel tokens with NLL. Also some normalizing flow magic, which I have to read up arxiv.org/abs/2411.19722
So this was the second cool thing that we* got this week. * this time the « we » is really just me, but hey, academic we 🙃
Our big_vision codebase is really good! And it's *the* reference for ViT, SigLIP, PaliGemma, JetFormer, ... including fine-tuning them. However, it's criminally undocumented. I tried using it outside Google to fine-tune PaliGemma and SigLIP on GPUs, and wrote a tutorial: lb.eyer.be/a/bv_tuto.html
Our big_vision codebase is really good! And it's *the* reference for ViT, SigLIP, PaliGemma, JetFormer, ... including fine-tuning them. However, it's criminally undocumented. I tried using it outside Google to fine-tune PaliGemma and SigLIP on GPUs, and wrote a tutorial: lb.eyer.be/a/bv_tuto.html
The first of the cool things we* got this week! Typically, you'd train a VQ-VAE/GAN tokenizer first, and then use its tokens for your LLM/DiT/... But we all know eventually end-to-end wins over pipelines. With flow models, you can actually learn pixel-LLM-pixel end-to-end!
Have you ever wondered how to train an autoregressive generative transformer on text and raw pixels, without a pretrained visual tokenizer (e.g. VQ-VAE)? We have been pondering this during summer and developed a new model: JetFormer 🌊🤖 arxiv.org/abs/2411.19722 A thread 👇 1/
Some recent discussions made me write up a short read on how I think about doing computer vision research when there's clear potential for abuse. Alternative title: why I decided to stop working on tracking. Curious about other's thoughts on this. lb.eyer.be/s/cv-ethics....
segment cow not on the beach, but as you may know I like cows in "out of distribution" places and poses. huggingface.co/spaces/big-v...
I think that it's about implementation because the idea as @giffmana.bsky.social has noted was well-known and unsurprisingly claimed by the Father of AI. But the method, which is the implementation of the technology, convinced people to use it
Sorry, but I was strangely attracted... @5trange4ttractor.bsky.social =D
Here's a fun real-life prompt where 30min of Googling didn't really help me. The small/fast/mini chatbots all failed miserably The large ones work great (except Gemini). Also: 1. I really like C3.6 giving only answer and asking if explain 2. Wild that Chat understood the calls and added comments!
Reminder that robots.txt is a gentleman’s agreement, not a legal document. The good actors all respect it, but there’s many bad actors out there on the internet.
Bluesky is an open and public social network, much like websites on the Internet itself. Websites can specify whether they consent to outside companies crawling their data with a robots.txt file, and we’re investigating a similar practice here.
No need to be a deep pocketed company, or HF. Scraping one million posts is not much. I did more than 14M just for fun on my single cheap desktop during PhD. Please don’t be intimidated too easily.
I'm glad @hf.co is doing this. It brings down the barriers to allow more people to benefit from AI, rather than keeping it exclusively in the realm of deep pocketed giant companies. AI can help open the gates, to allow regular people to do things they couldn't do before. (Which can be threatening!)
A sunny winter morning near the lake in Zürich is probably one of the best mood enhancers available world-wide. I don't think I'll ever get tired of this.
In the end I got emotional and couldn't bring myself to kill little `riesling`, who's remained loyal this whole decade. So I embarked on a journey to update Ubuntu 14.04 to 24.04. What an adventure! `failed to execute process '/usr/bin/ln'` almost killed me. But I made it, a decade of updates!
soooo, this marks the 10th anniversary of using digitalocean for my personal server. It's been running smoothly on 512MB RAM all that time. Now I'm going to splurge on a 1GB RAM, just so it feels like an upgrade! (Same price.)
So how long do we give it before the bots and spam arrive here, and how well do we think @bsky.app is prepared? 1. L I N K I N B I O. 2. Crypto spam 3. If you’re not doing these 10 GenAI things, you’re falling behind 4. Wow. This changes everything. My guesses: 2 months, and not well.