Jon Barron

@jonbarron.bsky.social

Principal research scientist at Google DeepMind. Synthesized views are my own. 📍SF Bay Area 🔗 http://jonbarron.info This feed is a mostly-incomplete mirror of https://x.com/jon_barron, I recommend you just follow me there.

I see many folks are pledging not to use AI in their writing. I pledge the opposite: I will use the latest LLMs, and for that matter any other available tool, to best improve my research or the way I communicate it. That way, if my name is on it, you can be sure it reflects my own best judgment.

Is basic image understanding solved in today’s SOTA VLMs? Not quite. We present VisualOverload, a VQA benchmark testing simple vision skills (like counting & OCR) in dense scenes. Even the best model (o3) only scores 19.8% on our hardest split.

Bild

Here’s what I’ve been working on for the past year. This is SkyTour, a 3D exterior tour utilizing Gaussian Splat. The UX is in the modeling of the “flight path.” I led the prototyping team that built the first POC. I was the sole designer and researcher on the project, one of the 1st inventors.

A thread of thoughts on radiance fields, from my keynote at 3DV: Radiance fields have had 3 distinct generations. First was NeRF: just posenc and a tiny MLP. This was slow to train but worked really well, and it was unusually compressed --- The NeRF was smaller than the images.

Bild

I made this handy cheat sheet for the jargon that 6DOF math maps to for cameras and vehicles. Worth learning if you, like me, are worried about embarrassing yourself in front of a cinematographer or naval admiral.

Bild

Next week is the one year anniversary of this paper showing that videos generated from Sora are nearly 3D-consistent. I'm surprised we never saw any follow-up papers in this line evaluating other videos models this way, it would be helpful to track these metrics over time. arxiv.org/abs/2402.17403

Bild

I just pushed a new paper to arXiv. I realized that a lot of my previous work on robust losses and nerf-y things was dancing around something simpler: a slight tweak to the classic Box-Cox power transform that makes it much more useful and stable. It's this f(x, λ) here:

My pitch for an "LLM-native" alternative to citation count/h-index etc: 1) Train an LLM on a new paper and record the average loss during training. 2) Evaluate the retrained LLM on benchmarks. Your Google Scholar records avg_train_loss * benchmark_delta, and you go write your next paper.

fun test for image and video generation systems: add "the camera is upside down" to the prompt (especially for shots of people) and then vertically mirror the output. Even the best models struggle, with upside-down teeth and blinks, and gravity tugging everything up. Here's Veo 2.

I fed the "spinning dancer" illusion (a silhouette of a spinning figure that can be seen as rotating clockwise or counter-clockwise, left) into Runway Gen-3 (right). It resolved the ambiguity by having the dancer face the camera and oscillate, which is kinda clever.

`A 1960s NASA scientist with a white button down shirt and black heavy rimmed glasses, with a giant thick alien umbilical cord coming out of the back of his body. The cord is holding him up in space, and he is levitating around his office. Wide angle, full body shot.` #Veo2

"A fun children's educational program where Mr. See-thru teaches kids about how the gastrointestinal system works using his semitransparent abdomen." #Veo2 I'm surprised by how well Veo 2 understands human anatomy, and amused by the things that it doesn't yet understand.