I see many folks are pledging not to use AI in their writing. I pledge the opposite: I will use the latest LLMs, and for that matter any other available tool, to best improve my research or the way I communicate it. That way, if my name is on it, you can be sure it reflects my own best judgment.
Jon Barron
@jonbarron.bsky.social
Principal research scientist at Google DeepMind. Synthesized views are my own. 📍SF Bay Area 🔗 http://jonbarron.info This feed is a mostly-incomplete mirror of https://x.com/jon_barron, I recommend you just follow me there.
Radiance Meshes for Volumetric Reconstruction Alexander Mai, Trevor Hedstrom, @grgkopanas.bsky.social, Janne Kontkanen, Falko Kuester, @jonbarron.bsky.social tl;dr: Delaunay tetrahedralization->constant density and linear color radian->radiance mesh->radiance field arxiv.org/abs/2512.04076
Is basic image understanding solved in today’s SOTA VLMs? Not quite. We present VisualOverload, a VQA benchmark testing simple vision skills (like counting & OCR) in dense scenes. Even the best model (o3) only scores 19.8% on our hardest split.
Here’s what I’ve been working on for the past year. This is SkyTour, a 3D exterior tour utilizing Gaussian Splat. The UX is in the modeling of the “flight path.” I led the prototyping team that built the first POC. I was the sole designer and researcher on the project, one of the 1st inventors.
🚀🚀🚀Announcing our $13M funding round to build the next generation of AI: 𝐒𝐩𝐚𝐭𝐢𝐚𝐥 𝐅𝐨𝐮𝐧𝐝𝐚𝐭𝐢𝐨𝐧 𝐌𝐨𝐝𝐞𝐥𝐬 that can generate entire 3D environments anchored in space & time. 🚀🚀🚀 Interested? Join our world-class team: 🌍 spaitial.ai youtu.be/FiGX82RUz8U
SpAItial AI: Building Spatial Foundation Models
YouTube video by SpAItial AI
youtu.be
📺 Now available: Watch the recording of Aaron Hertzmann's talk, "Can Computers Create Art?" www.youtube.com/watch?v=40CB... @uoftartsci.bsky.social
“Can Computers Create Art?” with Aaron Hertzmann — C.C. “Kelly” Gotlieb Distinguished Lecture Series
YouTube video by Arts & Science - University of Toronto
youtube.com
Thank you to all who joined us yesterday for “Can Computers Create Art?” presented by @aaronhertzmann.bsky.social. If you missed the lecture, stay tuned for the full recording.
Here's a recording of my 3DV keynote from a couple weeks ago. If you're already familiar with my research, I recommend skipping to ~22 minutes in where I get to the fun stuff (whether or not 3D has been bitter-lesson'ed by video generation models) www.youtube.com/watch?v=hFlF...
Radiance Fields and the Future of Generative Media
YouTube video by Jon Barron
youtube.com
A thread of thoughts on radiance fields, from my keynote at 3DV: Radiance fields have had 3 distinct generations. First was NeRF: just posenc and a tiny MLP. This was slow to train but worked really well, and it was unusually compressed --- The NeRF was smaller than the images.
Here's Bolt3D: fast feed-forward 3D generation from one or many input images. Diffusion means that generated scenes contain lots of interesting structure in unobserved regions. ~6 seconds to generate, renders in real time. Project page: szymanowiczs.github.io/bolt3d Arxiv: arxiv.org/abs/2503.14445
I made this handy cheat sheet for the jargon that 6DOF math maps to for cameras and vehicles. Worth learning if you, like me, are worried about embarrassing yourself in front of a cinematographer or naval admiral.
Can someone point me to a video of renderings of some 3DGS-like algorithm *as it is being optimized*? I want to see all those little Gaussians wobbling around.
Next week is the one year anniversary of this paper showing that videos generated from Sora are nearly 3D-consistent. I'm surprised we never saw any follow-up papers in this line evaluating other videos models this way, it would be helpful to track these metrics over time. arxiv.org/abs/2402.17403
Veo 2 now has a public price point: $0.50 per second. Very important number to keep in mind when considering the future of generative and non-generative media. Taken from cloud.google.com/vertex-ai/ge...
I just pushed a new paper to arXiv. I realized that a lot of my previous work on robust losses and nerf-y things was dancing around something simpler: a slight tweak to the classic Box-Cox power transform that makes it much more useful and stable. It's this f(x, λ) here:
My pitch for an "LLM-native" alternative to citation count/h-index etc: 1) Train an LLM on a new paper and record the average loss during training. 2) Evaluate the retrained LLM on benchmarks. Your Google Scholar records avg_train_loss * benchmark_delta, and you go write your next paper.
Great interview with @jascha.sohldickstein.com about diffusion models! This is the first in a series: similar interviews with Yang Song and yours truly will follow soon. (One of these is not like the others -- both of them basically invented the field, and I occasionally write a blog post 🥲)
History of Diffusion - Jascha Sohl-Dickstein
YouTube video by Bain Capital Ventures
youtube.com
With the CVPR 2025 rebuttal deadline over, it’s the perfect time to submit a demo application for CVPR 2025. Demos can be any application of computer vision in the real world and are a great way to show off your work to other #computervision enthusiasts. Submit here docs.google.com/forms/d/e/1F...
CVPR 2025 Demo Submission
Submission page for the Demo Track at CVPR 2025. Read the call for demos on the CVPR Website.
docs.google.com
fun test for image and video generation systems: add "the camera is upside down" to the prompt (especially for shots of people) and then vertically mirror the output. Even the best models struggle, with upside-down teeth and blinks, and gravity tugging everything up. Here's Veo 2.
I fed the "spinning dancer" illusion (a silhouette of a spinning figure that can be seen as rotating clockwise or counter-clockwise, left) into Runway Gen-3 (right). It resolved the ambiguity by having the dancer face the camera and oscillate, which is kinda clever.
`A 1960s NASA scientist with a white button down shirt and black heavy rimmed glasses, with a giant thick alien umbilical cord coming out of the back of his body. The cord is holding him up in space, and he is levitating around his office. Wide angle, full body shot.` #Veo2
"A fun children's educational program where Mr. See-thru teaches kids about how the gastrointestinal system works using his semitransparent abdomen." #Veo2 I'm surprised by how well Veo 2 understands human anatomy, and amused by the things that it doesn't yet understand.
I have been having fun turning the Mario Brothers into a 1940s industrial film using Veo 2.
In case it's helpful to others writing papers or doing comparisons, I'm hosting raw mp4s for the #Veo2 results I've been posting, with no Bluesky-induced compression: drive.google.com/drive/folder.... Here's `A kraken emerging from the ocean at the beach on coney island`, which I forgot to post.
`An advertisement for a shoe made of broccoli romanesco, rotating on a turntable` #Veo2
Happy Hanukkah! "A man using an upside-down Hanukkah menorah as a jetpack, blasting off" #Veo2
"A colorized 1950s TV commercial. Left: a dad struggling to pick up a heavy box. Right: that dad wearing a full-body mechanical exoskeleton, lifting the box with ease." #Veo2
"A giraffe riding four skateboards, one for each foot" #Veo2
Happy Hanukkah! "A row of industrial robot arms grabbing latkes from a conveyor belt and dunking them into a row of bowls of sour cream" #Veo2