Will Smith

@willsmithvision.bsky.social

Professor in Computer Vision at the University of York, vision/graphics/ML research, Boro @mfc.co.uk fan and climber 📍York, UK 🔗 https://www-users.york.ac.uk/~waps101/

I am delighted (that pun will make sense in a second) that @alistairfoggin.bsky.social's first paper, CroCoDiLight, has been accepted to ICLR. The idea came from a group discussion on the CroCo paper from @naverlabseurope.bsky.social and realising it might implicitly already understand relighting.

Alistair Foggin@alistairfoggin.bsky.social · 5mo ago

Very excited to announce my first paper, CroCoDiLight! Co-authored with @willsmithvision.bsky.social and accepted at #ICLR2026. Image relighting, shadow removal, and albedo estimation, all in our single base model with swappable task components. See you in Rio @iclr-conf.bsky.social! [1/7]

Teaser image for CroCoDiLight paper. Three panels showing the same scene of a street in three different lighting conditions: sunny with shadows, even lighting without shadows, and night time with lights switched on. A badge for ICLR 2026 is in the bottom right, and CroCoDiLight is in big letters across the top.

The LLM obsession with em-dashes has created a weird sort of paradox. I see so much AI generated content that I now see how I should have been using em-dashes all along. But if I start using them, everyone will assume what I've written is AI generated.

My students @fhudson.bsky.social and @jadgardner.bsky.social are presenting TAPVid-360 at NeurIPS this week. We introduce an interesting new problem, a benchmark dataset and a baseline adaptation of an existing TAP model for our task. More importantly, they've also created a genre-defining poster...

Finlay Hudson@fhudson.bsky.social · 8mo ago

TAPVid-360 will be at NeurIPS25 this week in sunny San Diego!! This work involves models having to understand beyond a camera’s field of view without the need for expensive 3D data. 👇 1/5

Has anyone ever tried a very non-standard tone for a rebuttal? I'm thinking something like "Hey reviewers! Sit back, relax and let me convince you that you actually want to accept this paper..." or "You wouldn't let a little thing like that stop you accepting the paper would you? WOULD YOU?!!!"

I just pushed a new paper to arXiv. I realized that a lot of my previous work on robust losses and nerf-y things was dancing around something simpler: a slight tweak to the classic Box-Cox power transform that makes it much more useful and stable. It's this f(x, λ) here:

#CVPR2025 Area Chair update: depending on which time zone the review deadline is specified in, we are past or close to the review deadline. Of the 60 reviews needed for my batch, I currently have 52 and they have been coming in quite fast this morning. In general, review standard looks good.

This simple pytorch trick will cut in half your GPU memory use / double your batch size (for real). Instead of adding losses and then computing backward, it's better to compute the backward on each loss (which frees the computational graph). Results will be exactly identical

Bild

Entropy is one of those formulas that many of us learn, swallow whole, and even use regularly without really understanding. (E.g., where does that “log” come from? Are there other possible formulas?) Yet there's an intuitive & almost inevitable way to arrive at this expression.

Introducing MegaSaM! Accurate, fast, & robust structure + camera estimation from casual monocular videos of dynamic scenes! MegaSaM outputs camera parameters and consistent video depth, scaling to long videos with unconstrained camera paths and complex scene dynamics!

How to drive your research forward? “I tested the idea we discussed last time. Here are some results. It does not work. (… awkward silence)” Such conversations happen so many times when meetings with students. How do we move forward? You need …

I am a first time Area Chair for #CVPR2025 so, in the interests of transparency, I'll post some updates here on the various stages of the process. There are 708 (!) ACs (not that long ago, CVPR could have coped with 708 *reviewers*!) We've been allocated 18.27 papers on average (I have 20).

Introducing Generative Omnimatte: A method for decomposing a video into complete layers, including objects and their associated effects (e.g., shadows, reflections). It enables a wide range of cool applications, such as video stylization, compositions, moment retiming, and object removal.

We've released our paper "Generating 3D-Consistent Videos from Unposed Internet Photos"! Video models like Luma generate pretty videos, but sometimes struggle with 3D consistency. We can do better by scaling them with 3D-aware objectives. 1/N page: genechou.com/kfcw