Yang Yang, Qinyu Zhao, Mouxiang Chen, Xiaohui Li, Lixin Gu, Wenhai Wang, Hongjie Zhang, Wenwei Zhang: ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs https://arxiv.org/abs/2608.04010 https://arxiv.org/pdf/2608.04010 https://arxiv.org/html/2608.04010
arXiv cs.CV Computer Vision and Pattern Recognition
@cscv-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/cs.CV/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Wanli Ma, Jiangwen Lu, Qinmu Peng, Xinge You: Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation https://arxiv.org/abs/2608.03991 https://arxiv.org/pdf/2608.03991 https://arxiv.org/html/2608.03991
Fang, Zeng, Huang, Zhao, Huang, Ren, Lu, Ren, Su, Wang, Yin, Chen, Chen, Chen, Yin, Hu, Lin, Ouyang, Cao, Zhao: Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent https://arxiv.org/abs/2608.03979 https://arxiv.org/pdf/2608.03979 https://arxiv.org/html/2608.03979
Yicheng Xiao, et al.: JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion https://arxiv.org/abs/2608.03974 https://arxiv.org/pdf/2608.03974 https://arxiv.org/html/2608.03974
Li, Yan, Bai, Chen, Sun, Wang, Wu, Yuan, Lin, Liu, Niu, Yuan: UniWorld-Design: From Pixel Generation to Layer-Native Design https://arxiv.org/abs/2608.03971 https://arxiv.org/pdf/2608.03971 https://arxiv.org/html/2608.03971
Noor Hussein, Anil K. Jain, Karthik Nandakumar: Progressive Learning of a Diffusion-based Inpainting Model for Separating Overlapped Fingerprints https://arxiv.org/abs/2608.03937 https://arxiv.org/pdf/2608.03937 https://arxiv.org/html/2608.03937
Lu Gan, Hanyu Yan, Chaofeng Chen, Junqi Hu, Dan Zeng: GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration https://arxiv.org/abs/2608.03923 https://arxiv.org/pdf/2608.03923 https://arxiv.org/html/2608.03923
Peng Xia, Junbiao Pang, Zheng Huang: Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization https://arxiv.org/abs/2608.03919 https://arxiv.org/pdf/2608.03919 https://arxiv.org/html/2608.03919
Li, Chen, Li, Zheng, Zou, Zhang, Liu, Chen: When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding https://arxiv.org/abs/2608.03918 https://arxiv.org/pdf/2608.03918 https://arxiv.org/html/2608.03918
Xiang Chen: StreamDAM: Presence-Aware Memory for Real-Time Streaming Video Object Segmentation https://arxiv.org/abs/2608.03912 https://arxiv.org/pdf/2608.03912 https://arxiv.org/html/2608.03912
Zhang, Li, Hu, Yang, Zou, Zhang, Dong: UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution https://arxiv.org/abs/2608.03911 https://arxiv.org/pdf/2608.03911 https://arxiv.org/html/2608.03911
Wenbin Pan, Wanhao Liu, Liwei Luo, Panshuo Li, Yong Xu, Renquan Lu: NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection https://arxiv.org/abs/2608.03895 https://arxiv.org/pdf/2608.03895 https://arxiv.org/html/2608.03895
Mercy Prasanna Ranjit, et al.: CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement https://arxiv.org/abs/2608.03890 https://arxiv.org/pdf/2608.03890 https://arxiv.org/html/2608.03890
Liu, Wang, Liu, Wu, Wang, Wang, Chen, Ji: MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization https://arxiv.org/abs/2608.03885 https://arxiv.org/pdf/2608.03885 https://arxiv.org/html/2608.03885
Sadab Shiper, Tawsif Tashwar Dipto, Mir Md Inzamam, Eshat Tanzeem: BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models https://arxiv.org/abs/2608.03884 https://arxiv.org/pdf/2608.03884 https://arxiv.org/html/2608.03884
Yvan Richard: CPrefix: A Combinatorial Tensor Framework for Structured Discrete Color Mappings https://arxiv.org/abs/2608.03863 https://arxiv.org/pdf/2608.03863 https://arxiv.org/html/2608.03863
Tianbao Zhang, Zeyu Liu, Shuyu Wu, Fanxing Li, Zhaoxin Fan, Wenjun Wu, Danping Zou: LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation https://arxiv.org/abs/2608.03851 https://arxiv.org/pdf/2608.03851 https://arxiv.org/html/2608.03851
Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu: Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding https://arxiv.org/abs/2608.03826 https://arxiv.org/pdf/2608.03826 https://arxiv.org/html/2608.03826
Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli: FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis https://arxiv.org/abs/2608.03822 https://arxiv.org/pdf/2608.03822 https://arxiv.org/html/2608.03822
Ezzati, Rezaee, Kariminia, Yousefi, Mamaqani, Samimi, Rohban: UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space https://arxiv.org/abs/2608.03817 https://arxiv.org/pdf/2608.03817 https://arxiv.org/html/2608.03817
Su, Shi, Liu, Yu, Min, Zhang, Wang, Wang, Liu, Zhang, Wu, Huo, Ding: OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models https://arxiv.org/abs/2608.03812 https://arxiv.org/pdf/2608.03812 https://arxiv.org/html/2608.03812
Yuxiang Duan, Huining Li, Ao Li, Shuai Feng, Lanju Kong, Ning Liu, Jian Zhang, Xingdong Sheng, Yuntao Du: AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding https://arxiv.org/abs/2608.03779 https://arxiv.org/pdf/2608.03779 https://arxiv.org/html/2608.03779
Qingxi Du, Junbo Wang, Yuke Li, Yining Zhu: TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding https://arxiv.org/abs/2608.03763 https://arxiv.org/pdf/2608.03763 https://arxiv.org/html/2608.03763
Maccarone, Di Stefano, Longari, Frigerio, Rizzato, Prudentino, Agarwal, Ciceri, Peruzzo, Melzi: Towards Reliable and Reproducible Fetal Brain Biometry: A Deep Learning Approach Using MRI https://arxiv.org/abs/2608.03724 https://arxiv.org/pdf/2608.03724 https://arxiv.org/html/2608.03724
Maximilian Dillitzer, Tin Stribor Sohn, Jason J. Corso, Michael Auerbach: Attention is Case-Sensitive https://arxiv.org/abs/2608.03711 https://arxiv.org/pdf/2608.03711 https://arxiv.org/html/2608.03711
Ruirui Zhang, Zhengkai Zhao, Pan Gao: MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding https://arxiv.org/abs/2608.03708 https://arxiv.org/pdf/2608.03708 https://arxiv.org/html/2608.03708
Yanning Hou, Jingyuan Zhang, Xiaoyun Wang, Qixiang Ma, Sihang Zhou, Ke Xu: Keep the Needle, Prune the Haystack: Defect-Preserving Token Pruning for Efficient Zero-Shot Anomaly Detection https://arxiv.org/abs/2608.03681 https://arxiv.org/pdf/2608.03681 https://arxiv.org/html/2608.03681
Elena Izzo, Riccardo Toniolo, Lamberto Ballan: XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation https://arxiv.org/abs/2608.03666 https://arxiv.org/pdf/2608.03666 https://arxiv.org/html/2608.03666
Jiaming Liang, QiHui Han, Haolin Chen, Chengxin Ye, Jiawen Liu, Jiazhou Chen, Xiaoqi Sheng, Hongmin Cai: Morphology-Aware Implicit Super-Resolution Network for Pathological Images https://arxiv.org/abs/2608.03664 https://arxiv.org/pdf/2608.03664 https://arxiv.org/html/2608.03664
Hao Dou, Ruiwen Tian: When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware https://arxiv.org/abs/2608.03649 https://arxiv.org/pdf/2608.03649 https://arxiv.org/html/2608.03649