SGFormer: Structure-Guided Transformer for Robust Local Feature Matching Runyu Zhu tl;dr: network attention->overlapping regions; shallow local features from early network layers->deeper layer->salient structure arxiv.org/abs/2608.03423
Zhenjun Zhao
@ericzzj.bsky.social
ericzzj1989.github.io Postdoc@UniZar | PhD@CUHK | 3D vision, SLAM, Image matching (http://github.com/ericzzj1989/Awesome-Global-Solvers-for-3D-Vision)
SLAMFormer-∞: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing Zhijian Fang, Weicheng Zheng, Yijun Yuan, Weibang Wang, Zhuoguang Chen, Chang Sun, Junhao Huang, Kenan Li, Minghui Qin, Hang Zhao tl;dr: long-range SLAM-Former arxiv.org/abs/2608.03429
QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction Yinglong Li, Donghui Shen, Xiaoyu Zhang, Zhichao Ye, Hongyu Wu, Aimin Hao, Guofeng Zhang, Haomin Liu tl;dr: geometry-appearance dual-query decoder arxiv.org/abs/2608.01186
UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization Inha Lee, Dongjae Jeong, Junhee Lee, Kyungdon Joo tl;dr: keyframe poses + submap-local poses->Sim(3) factor graph with view-to-view & view-to-submap & submap-to-submap edges arxiv.org/abs/2608.01706
StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting Changhao Song, Yuxuan Wang, Qibiao Li, Youcheng Cai, Ligang Liu tl;dr: geometry-grounded historical tokens->depth estimation and Gaussian-token regression arxiv.org/abs/2608.01659
Rolling Shutter Camera Self-Calibration Yongcong Zhang, Navid Rabbani, Bangyan Liao, Chengbo Wang, Yizhen Lao, Adrien Bartoli tl;dr: B-spline-based trajectory estimation + correction fields->self-calibration->intrinsics and readout time ratio arxiv.org/abs/2608.01509
UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction Haixu Song, Xiaoke Yang, Shengjun Zhang, Jiwen Lu, Yueqi Duan tl;dr: view-conditioned information->prior->Gaussians arxiv.org/abs/2608.02145
Stipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking Kilian Northoff, Mateo de Mayo, @dcremers.bsky.social tl;dr: Basalt + 3DGS arxiv.org/abs/2608.00931
NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction Jiaheng Li, Binsheng Zhang, Xinhai Chang, Wenzheng Chen tl;dr: in title arxiv.org/abs/2607.24495
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion Zador Pataki, @pesarlin.bsky.social, @marcpollefeys.bsky.social arxiv.org/abs/2607.27194
Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis Maria Peribañez, @jcivera.bsky.social, Rudolph Triebel, Riccardo Giubilato tl;dr: MVSAnywhere-based depth & normal alignment loss + LiDAR-guided Chamfer-based loss arxiv.org/abs/2607.22147
SM4RT: Learning Structured Motion Geometry for 4D Reconstruction Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, Yuqi Wu, Jie Zhou, Jiwen Lu tl;dr: scene motion->a set of N motion bases; base->a temporal sequence of 6D twists in se(3) arxiv.org/abs/2607.22534
DB-VIO: Dual-Branch Visual Inertial Odometry with Enhanced Visual-Inertial Representation Ziyu Wan, Lin Zhao tl;dr: decouple rotational & translational motion modeling; Metric3D->visual features; gyroscope measurements->attitude prior and->IMU features arxiv.org/abs/2607.22123
Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering Oriol Jiménez-Ayguadé, Antonio Agudo tl;dr: augment triangle with K learnable control points per edge to represent shapes as a soup of non-uniform primitives arxiv.org/abs/2607.22446
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao, Xiaolin Zhou, Fangzhou Hong, Zhizhong Su, Dingwen Zhang, Ziwei Liu tl;dr: in title arxiv.org/abs/2607.19228
GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition Panagiotis Mermigkas, Argyris Manetas, Petros Maragos tl;dr: ORB-SLAM2 + Scaffold-GS + LiteFlowNet3 arxiv.org/abs/2607.21416
Exploration Matters for Escaping the Blur Trap in 3D Gaussian Splatting Chengbo Wang, Guozheng Ma, Jinhong Wu, Tie Ji, Yizhen Lao tl;dr: optimization limitation of 3DGS->viewing-ray gradient orthogonality and blending-induced gradient attenuation arxiv.org/abs/2607.17965
Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement Lingyu Kong, Ruicheng Li, Ruicheng Wang, Sicheng Xu, Chengtang Yao, Jianfeng Xiang, Jiaolong Yang tl;dr: aggregate features based on 3D spatial locality arxiv.org/abs/2607.17967
CSS-BA: Gate-Guided Column Space Search for Bundle Adjustment Ayano Kaneda, Takafumi Taketomi, Shugo Yamaguchi, Shigeo Morishima tl;dr: gate-guided camera-block coordinate subspaces->Lanczos/Ritz basis->LM arxiv.org/abs/2607.15652
HETA++: Global Structure-from-Motion with Hybrid Explicit Translation Averaging Peilin Tao, Hainan Cui, Mengqi Rong, Shuhan Shen tl;dr: HETA, but not BCD arxiv.org/abs/2607.15912
SeeSE3: Emergence of 3D Space in Vision Features Caroline Chen, Sayna Ebrahimi, Fedor Kitashov, Ming-Hsuan Yang, Leonidas Guibas, Viorica Pătrăucean, Maks Ovsjanikov tl;dr: vision features capture the structure of Euclidean space? arxiv.org/abs/2607.14228
Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency @andreasmeuleman.bsky.social, @linusfranke.bsky.social, Boris Zhestiankin, Camille Montemagni, George Drettakis tl;dr: 3DGS + immediate feedback arxiv.org/abs/2607.14481
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos Ziren Gong, Xiaohan Li, Fabio Tosi, Ninghui Xu, Stefano Mattoccia, Jianfei Cai, @mattpoggi.bsky.social tl;dr: multi-agent VGGT arxiv.org/abs/2607.15211
SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment Saad Ejaz, Miguel Fernandez-Cortizas, @jcivera.bsky.social, Holger Voos, Jose Luis Sanchez-Lopez tl;dr: scaling + geometry-aware matching arxiv.org/abs/2607.15058
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation tl;dr: in title arxiv.org/abs/2607.14203
Video Generation Models are General-Purpose Vision Learners Letian Wang, Chuhan Zhang, @rkabra.bsky.social, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu arxiv.org/abs/2607.09024
Glob3R: Global Structure-from-Motion with 3D Foundation Models Junyuan Deng, Heng Li, Kejie Qiu, Lingteng Qiu, Rui Peng, Weichao Shen, Weihao Yuan, Siyu Zhu, Zilong Dong, Ping Tan tl;dr: Pi3X->keyframe selection+dense warps->sparse tracks->GLOMAP arxiv.org/abs/2607.09225
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility Filippo Ziliotto, Luciano Serafini, @lambertoballan.bsky.social, Tommaso Campari tl;dr: early layers->3D-aware scene representation; late layers->co-visibility; VGGT+MoE head arxiv.org/abs/2607.09503
DGSfM: Depth-Guided Scale-Aware Global Structure-from-Motion Sithu Aung, Viktor Kocur, Yaqing Ding, @sattlertorsten.bsky.social, Zuzana Kukelova arxiv.org/abs/2607.09507
Incremental Online Scene Reconstruction by 3D Gaussian Triangulation Yanjin Zhu, Shaofan Liu, Jianke Zhu tl;dr: optimized 3D Gaussians->surfel primitives->direct triangulation arxiv.org/abs/2607.10690