| 1 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | Junchao Huang, Guian Fang, Shengju Qian +15 | cs.CV | 2026-09-02 |
| 2 | Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation | Yutong Liu, Nan Huang, Xu Cao +1 | cs.CV | 2026-09-02 |
| 3 | PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation | Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2 | cs.CV | 2026-09-02 |
| #4 | MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion | Aidan Bradshaw, Marco Giordano, David Rode +8 | cs.CV | 2026-09-02 |
| #5 | RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation | Xiaolei Lang, Ze Kang, Zehao Huang +1 | cs.CV | 2026-09-02 |
| #6 | Efficient All-in-One Weather Restoration using Spectral Harmonization | Paula Garrido-Mellado, Daniel Feijoo, Yuning Cui +2 | cs.CV | 2026-09-02 |
| #7 | Benchmarking RAW and RGB Restoration in Image Signal Processors | Zihao Lu, Radu Timofte, Marcos V. Conde | cs.CV | 2026-09-02 |
| #8 | GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design | Adrienne Deganutti, Purvanshi Mehta, Simon Hadfield +1 | cs.CV | 2026-09-02 |
| #9 | AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels | Javier Tirado-Garín, Alan Savio Paul, Shuai Chen +5 | cs.CV | 2026-09-02 |
| #10 | ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding | Jitai Hao, Ke Yang, Qiang Huang +1 | cs.CV | 2026-09-02 |
| #11 | Video-Based Palm-Vein Authentication under Challenging Conditions | Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh +3 | cs.CV | 2026-09-02 |
| #12 | Multi-Tool Image Editing Attribution in Facial Forgery | Sheng Liu, Qiang Sheng, Danding Wang +3 | cs.CV | 2026-09-02 |
| #13 | Balancing Frequencies and Pixels in Flow Matching | Lucas Degeorge, Paul Couairon, Arijit Ghosh +3 | cs.CV | 2026-09-02 |
| #14 | InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling | Tianheng Lu, Guangyu Wang, Ruqi Huang +1 | cs.CV | 2026-09-02 |
| #15 | RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models | Canjie Liu, Jiawen Kang, Jinbo Wen +1 | cs.CV | 2026-09-02 |
| #16 | MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception | Guido Caccianiga, Sergey Prokudin, Yutong Chen +9 | cs.CV | 2026-09-02 |
| #17 | A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames | Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh +1 | cs.CV | 2026-09-02 |
| #18 | Generating Medical Image Counterfactuals using Causal Explanations | David A. Kelly, Tom Yaacov, Nathan Blake +2 | cs.CV | 2026-09-02 |
| #19 | GaLe: memory-efficient Global Approximate and Local Exact features | Alberto Ancilotto, Elisabetta Farella | cs.CV | 2026-09-02 |
| #20 | Genesis: A Generative Engine for Hierarchical Satellite Image Synthesis | Subash Khanal, Yangzhi Cui, Daniel Cher +4 | cs.CV | 2026-09-02 |
| #21 | LoFi RADIO: A Distilled In-Domain Backbone Applied for Artifact-Severity Grading of Ultra-Low-Field Neonatal Brain MR | Jonathan B. Martin, Yashwant Kurmi, Charlotte R. Sappo | eess.IV | 2026-09-02 |
| #22 | Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations | Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran | cs.CV | 2026-09-02 |
| #23 | Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling | Ziquan Liu, Zhewei Zhu, Xuyang Shi | cs.CV | 2026-09-02 |
| #24 | Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face) | El Ouanas Belabbaci, Bhavesh Wani, Philipp Terhörst | cs.CV | 2026-09-02 |
| #25 | From Detection to Localization: A Unified Forensics Framework for Fully Synthetic and Tampered Images | Annalisa Gallina, Marco Fiorucci, Marco Brigo +2 | cs.CV | 2026-09-02 |
| #26 | AffectDelta: Beyond Emotion Labels for Image Editing | Xingzu Zhan, Lin Gu, Ruogu Fang | cs.CV | 2026-09-02 |
| #27 | Generalizable Brain Tumor Segmentation with Self-Training and Tumor-Aware Deformations | Henrique Zan Grande, Jeovane Honorio Alves, Rayson Laroca +1 | cs.CV | 2026-09-02 |
| #28 | Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment | Zihao Wang, Xi Xiang, Yuwen Sun +5 | cs.CV | 2026-09-02 |
| #29 | MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval? | Uicheol Jung, Juyoung Hong, Geuntaek Lim +1 | cs.CV | 2026-09-02 |
| #30 | Stereo 4D Radar for 3D Object Detection: Integrating Geometric Alignment and Absolute Velocity Estimation | Seung-Hyun Song, Dong-Hee Paek, Woong-Chan Byun +1 | cs.CV | 2026-09-02 |
| #31 | RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains | Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden +4 | cs.CV | 2026-09-02 |
| #32 | Spatially Aware World Action Model via Geometric Latent Diffusion | Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid | cs.CV | 2026-09-02 |
| #33 | Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework | Yan Zhong, Gefei Chen, Qiufang Ma +4 | cs.CV | 2026-09-02 |
| #34 | Doppio: A Dataset for Contactless Weight Estimation of Falling Particles | Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5 | cs.CV | 2026-09-02 |
| #35 | Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness | Santiago Grandas, Juan Sebastian Cely-Acosta, Mohit Mendiratta +2 | cs.CV | 2026-09-02 |
| #36 | Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition | Naoto Nishida, Yoshio Ishiguro | cs.CV | 2026-09-02 |
| #37 | SR-Edit: Region-Aware Image Editing via Self-Refinement | Andong Wang, Zehua Chen, Yuxuan Jiang +1 | cs.CV | 2026-09-02 |
| #38 | Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models | Chuer Chen, Zichen Wang, Yi He +2 | cs.CV | 2026-09-02 |
| #39 | ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering | Adrien Mialland, Marc Plantevit, Julien Gallois +1 | cs.IR | 2026-09-02 |
| #40 | UnCapsTSR: An Unsupervised Transformer-based Image Super-Resolution Approach for Capsule Endoscopy Images | Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal +3 | cs.CV | 2026-09-02 |
| #41 | Learning to Track from Privileged Target Appearances | Xin Chen, Jiao Xu, Dong Wang +2 | cs.CV | 2026-09-02 |
| #42 | VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation | Hoonhee Cho, Jae-Young Kang, Giwon Lee +3 | cs.CV | 2026-09-02 |
| #43 | WiFlow: Estimating Optical Flow using WiFi Channel State Information | Thomas Weigel, Simon Kiefhaber, Fabian Portner +2 | cs.CV | 2026-09-02 |
| #44 | Adapting a Foundation Model for Lunar Surface Height Estimation | Patrick Bauer, Marius Schwinning, Melanie Siegel +2 | cs.CV | 2026-09-02 |
| #45 | Uncertainty-Guided Adverse Weather Restoration via Gated Transformer Network | Zheke Jin, Yuning Cui, Tianle Jin +2 | cs.CV | 2026-09-02 |
| #46 | The Diagnosis a Reporter Leaves Unspoken: Surfacing Frozen Tumor Features for Brain-Tumor MRI Reporting | Khawaja Murad ul Hassan, Ruqiyya Adil, Adil Qayyum +4 | cs.CV | 2026-09-02 |
| #47 | CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction | Menghao Li, Linjie Mu, Yin Wang +4 | cs.CV | 2026-09-02 |
| #48 | Seeing Beyond the Lesion: Disease Recognition from Reactive CNS Tissue | Jan Schnorrenberg, Jan Ernsting, Enrico Küllenberg +3 | eess.IV | 2026-09-02 |
| #49 | ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution | Byoungwoo Kim, Munchurl Kim | cs.CV | 2026-09-02 |
| #50 | Information Density Imbalance in Visual Object Detection | Ziwei Zhao, Yanxi Lu, Yuwei Hu +8 | cs.CV | 2026-09-02 |
| #51 | The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation | Yichen Liu, Quanwei Zhang, Haozhe Wang +7 | cs.MM | 2026-09-02 |
| #52 | TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models | Leqian Ding, Junning Qiu, Manwen Yang +2 | cs.CV | 2026-09-02 |
| #53 | LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory | Kun-Yang Yu, Yingzhe Li, Hongyu Xu +8 | cs.CV | 2026-09-02 |
| #54 | GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors | Qiang Xiang, Shuang Sun, Binglei Li +4 | cs.CV | 2026-09-02 |
| #55 | Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation | Luo Li, Chongchong Huang, Jun Jia +4 | cs.CV | 2026-09-02 |
| #56 | Towards Zero-Shot Transfer Across Embodiments For Driving VLAs | Caio Azevedo, Stefano Sabatini, Sascha Hornauer +1 | cs.CV | 2026-09-02 |
| #57 | ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans | Amirhosein Azarpour | cs.CV | 2026-09-02 |
| #58 | YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification | Quansheng Hu, Qin Sun, Qiansen Dai +4 | cs.CV | 2026-09-02 |
| #59 | Domain shift-robust object detection with GenAI image editing | Isabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel +4 | cs.CV | 2026-09-02 |
| #60 | VoRTeC: Taming Foundation Flow for One-step Real time Video Compression | Yichong Xia, Qinhong Wu, Qinhong Wu +3 | cs.CV | 2026-09-02 |
| #61 | If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection | Yinghao Sun, Shuguang Li, Jinliang Shao +1 | cs.CV | 2026-09-02 |
| #62 | Diffusion-Encoding Gaussian Field for Joint k-q dMRI Reconstruction | Zhibo Chen, Yajuan Huang, Yu Guan +3 | cs.CV | 2026-09-02 |
| #63 | RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification | Jierui Li, Zhiyuan Qi, Hao Zhu +7 | cs.CV | 2026-09-02 |
| #64 | Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics | Yijie Lin, Ching-Chun Chang, Isao Echizen +2 | cs.CR | 2026-09-02 |
| #65 | MAOL: Morphology-Aware Ordinal Learning for Fine-Grained Industrial Defect Severity Grading | Zhaoyang Wang, Haiyong Chen, Binyi Su +4 | cs.CV | 2026-09-02 |
| #66 | T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation | Yan Wang, Xinyi Hou, Weiguo Lin +2 | cs.CV | 2026-09-02 |
| #67 | Handwriting Trajectory Recovery via Autoregressive Ordered Stroke Instance Prediction | En-Guang Wang, Yan-Ming Zhang, Fei Yin +1 | cs.CV | 2026-09-02 |
| #68 | SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation | Qin Lu, Zheyang Jing, Yujie Yang +3 | cs.CV | 2026-09-02 |
| #69 | InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models | Chengyin Hu, Dingyi Lu, Jiaju Han +5 | cs.CV | 2026-09-02 |
| #70 | Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging | Khawaja Murad ul Hassan, Mehran Ebrahimi | cs.CV | 2026-09-02 |
| #71 | Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis | Siddhant Shete, Hilmi Dogu Kücüker, Udo Frese +1 | cs.RO | 2026-09-02 |
| #72 | FuDU: A Fuzzy Dual-dimensional Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection | Zhaoyang Wang, Haiyong Chen, Binyi Su +1 | cs.CV | 2026-09-02 |
| #73 | Asymmetric Paired-Annotation Learning for Multi-Structure ULF Pediatric Brain MRI Segmentation | Ha-Hieu Pham, Dang P. M. Cao, Minh Hoang Pham +4 | cs.CV | 2026-09-02 |
| #74 | LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images | Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov | cs.CV | 2026-09-02 |
| #75 | TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval | Uicheol Jung, Juyoung Hong, Hojung Kwon +1 | cs.CV | 2026-09-02 |
| #76 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #77 | CC-4DGS: Computational Deformation and Point-Cloud Compression for Storage-Efficient Dynamic Gaussian Splatting | Kyungdae Park, Chae Eun Rhee | cs.CV | 2026-09-02 |
| #78 | Progressive Pseudo-Label Optimization for Point-Supervised Change Detection | Hailong Ning, Hao Wang, Yimeng Wang +3 | cs.CV | 2026-09-02 |
| #79 | World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models | Chuhan Zhang, Seiji Ito, Kenta Hoshino +2 | cs.CV | 2026-09-02 |
| #80 | Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology | Mingxin Liu, Chengfei Cai, Anwen Lu +5 | cs.CV | 2026-09-02 |
| #81 | Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap | Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2 | cs.CV | 2026-09-02 |
| #82 | A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization | Xianghong Fang, Wenlong Mou, Yuan Yuan +2 | cs.LG | 2026-09-02 |
| #83 | Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts | Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal +2 | cs.LG | 2026-09-02 |
| #84 | Evidence-Guided Detection, Localization and Explanation for Text-Centric Image Forensics | Peifeng Liu, Bin Li, Qingsong Zhang +3 | cs.CV | 2026-09-02 |
| #85 | Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development | Yilong Guo, Hanqi Chen, Zixiao Ye +3 | cs.CV | 2026-09-02 |
| #86 | TC-Next: Zero-Shot Multimodal Cyclone Forecasting | Zhe Wang, Sijie Chen, Yiming Luo +2 | cs.LG | 2026-09-02 |
| #87 | KSG-Net: Key-Sparse and Global-Context Learning for Maritime 3D Ship Detection | Zhouyuan Huai, Meiqi Wan, Yan Yang +4 | cs.CV | 2026-09-02 |
| #88 | DPA: Decoupling Product-Agnostic Anomaly Representations for Zero-shot Anomaly Generation | Hang Yao, Yansheng Fu, Ming Liu +4 | cs.CV | 2026-09-02 |
| #89 | LaST-SR: Laplace-Inspired Steady-Transient Complex-Frequency Decomposition for Single Image Super-Resolution | Linhao Li, Zhaojie Pan, Langkun Chen | cs.CV | 2026-09-02 |
| #90 | DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents | Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park +5 | cs.AI | 2026-09-02 |
| #91 | Test-Time Logit Prompting for Source-Free Missing Modality Adaptation | Taixi Chen, Nancy Guo | cs.CV | 2026-09-02 |
| #92 | SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition | Tingyan Wen, Chenqian Yan, Xurui Peng +4 | cs.CV | 2026-09-02 |
| #93 | Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification | Xuanbing Wen, Boxu Chen, Le Yang +4 | cs.CV | 2026-09-02 |
| #94 | Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning | Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri | cs.LG | 2026-09-02 |
| #95 | Perceptually Regularized Diffusion Model for Image Super-Resolution | Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang +4 | eess.IV | 2026-09-02 |
| #96 | GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric Matching | Lu Han, Xiting Sun, Hao Wang +4 | cs.CV | 2026-09-02 |
| #97 | InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation | Ziquan Liu, Zhewei Zhu, Xuyang Shi | cs.CV | 2026-09-02 |
| #98 | InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation | Vanshika Vats, Ashwani Rathee, James Davis | cs.CV | 2026-09-02 |
| #99 | Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework | Shuyao Xiao, Shengling Wang, Haoyu Niu +4 | cs.CV | 2026-09-02 |
| #100 | Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation | Akio Hayakawa, Yusuke Mukuta, Tatsuya Harada | cs.CV | 2026-09-02 |
| #101 | Morphology signal in whole slide image foundation models can automatically triage slides | Ayushi Sinha, Shashank Yadav, Benjamin Holmes +9 | cs.CV | 2026-09-02 |
| #102 | Aggregating Neighbor Embedding Projection and Rank-Based Manifold Learning for Image Retrieval | Vinicius Atsushi Sato Kawai, Gustavo Rosseto Leticio, Lucas Pascotti Valem +1 | cs.CV | 2026-09-02 |
| #103 | Data-Efficient Networks for Multi-Contrast MRI Reconstruction based on a Generalized Content/Style Prior | Chinmay Rao, Efe Ilıcak, Matthias J. P. van Osch +5 | eess.IV | 2026-09-02 |
| #104 | Learning with Volterra Neural Networks: A System Theoretic Perspective | Haoyu Yun, Hamid Krim, Yufang Bao | cs.CV | 2026-09-01 |
| #105 | Automated Maize Ear Phenotyping Using 3D Reconstructions | Ritwesh A. Kumar, Som Tripathi, Peja Matthews +5 | cs.CV | 2026-09-01 |
| #106 | TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views | Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman +9 | cs.CV | 2026-09-01 |
| #107 | Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods | Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1 | cs.CV | 2026-09-01 |
| #108 | SignMatch: Matching Dictionary Signs to Continuous Sign Language Video | Ryan Wong, Youngjoon Jang, Liliane Momeni +2 | cs.CV | 2026-09-01 |
| #109 | RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation | Zixiang Tong, Lehu Bu, Jin Yang | cs.CV | 2026-09-01 |
| #110 | Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation | R. James Cotton, Divya Joshi, Colleen Peyton | cs.CV | 2026-09-01 |
| #111 | SliceBridge: context-consistent repair of corrupted slice intervals in T1-weighted MRI | Jiheng Li, Michael E. Kim, Trent Schwartz +6 | cs.CV | 2026-09-01 |
| #112 | Kirin: Animal Motion Generation from In-the-Wild Video | Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg +2 | cs.CV | 2026-09-01 |
| #113 | Video2Reaction: Training Foundation Video Models to Predict Audience Reaction | Sidong Zhang, Trang Nguyen, Shiv Shankar +4 | cs.CV | 2026-09-01 |
| #114 | Integrated Laser Scanning and Image-Based Topology Optimization Techniques for Detection and Quantification of Visible and Subsurface Structural Defects | Mehrdad Shafiei Dizaji, Devin Harris | cs.CV | 2026-09-01 |
| #115 | Consistency as Regularization for Unsupervised Shadow Removal | Anh-Kiet Duong, Petra Gomez-Krämer, Jean-Michel Carozza | cs.CV | 2026-09-01 |
| #116 | Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks | C. J. Moore, Alex Hurt, Jordan Malof | cs.CV | 2026-09-01 |
| #117 | Designing Versatile Samples for Learned Trajectory Scoring | Yaguang Li, Jiaru Zhang, Chuheng Wei +2 | cs.RO | 2026-09-01 |
| #118 | DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation | Atif Belal, Lilian Hollard, Marco Pedersoli +1 | cs.CV | 2026-09-01 |
| #119 | Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade | Daniela Ruiz, Manuel Castellote, Zhongqi Miao +5 | cs.SD | 2026-09-01 |
| #120 | CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer | Yisen Wang, Zhirong Wu, Limin Wang | cs.CV | 2026-09-01 |
| #121 | Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation | Ehsan Faghih, Fatemeh Ashrafi, Marguerite Moore +1 | cs.CV | 2026-09-01 |
| #122 | Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings | Song Jin, Zhongtao Jiang, Chenglei Shen +5 | cs.CV | 2026-09-01 |
| #123 | MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models | Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5 | cs.CL | 2026-09-01 |
| #124 | AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $α$-Corrected Binary Cross Entropy and Factorized Latent Supervision | Jianzhong You, Yuan Gao, Chris McIntosh | cs.CV | 2026-09-01 |
| #125 | Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics | Sejuti Basu, Ashima Sood, Vijay Kumar +1 | cs.CV | 2026-09-01 |
| #126 | SCULPT: Training Edge Vision Models for Post-Training Quantization Readiness | Bharadwaj Kavuri, Sourav Babu-PK, Varadhraj Ellapan +2 | cs.CV | 2026-09-01 |
| #127 | Evidential Deep Learning for Multi-Modal Anti-UAV Detection | Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag | cs.CV | 2026-09-01 |
| #128 | ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes | Mingda Lin, Weijie Wang, Zeyu Zhang +7 | cs.CV | 2026-09-01 |
| #129 | UAV Thermal Imagery for Inert Ordnance Screening: Multi Campaign Dataset Development,Object Detection, and Practical Recommendations | Chad Melton, PhD., Annabelle Kelton | cs.CV | 2026-09-01 |
| #130 | From Visual Cues to Spoken Narration: Rethinking Audio Description | Akshita Gupta, Aditya Arora, Federico Tombari +2 | cs.CV | 2026-09-01 |
| #131 | Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System | Penghao Wu, Haiwen Diao, Weichen Fan +3 | cs.CV | 2026-09-01 |
| #132 | UI-VISA: U-Net Initialized Vascular Image Segmentation Architecture | Asees Kaur, Suzanne S. Sindi, Erica M. Rutter | cs.CV | 2026-09-01 |
| #133 | A Benchmark for Vehicle Attribute Classification in Cross-Domain Surveillance Scenarios | Sergio M. Silva, Otavio T. Remer, Gabriel E. Lima +3 | cs.CV | 2026-09-01 |
| #134 | SpatialGuard: Harness-Guided Verifiable Spatial Reasoning for Text-to-Image Generation | Ziyun Qian, Zizhi Chen, Yizhou Liu +3 | cs.CV | 2026-09-01 |
| #135 | H3-World: Turning Language Understanding into World Control | Danze Chen, Zeqing Wang, Ziyue Lin +2 | cs.CV | 2026-09-01 |
| #136 | BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net | Marven Sherif, Amgad Elmasry, Youssef Ghazal +1 | cs.CV | 2026-09-01 |
| #137 | What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models | Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3 | cs.CV | 2026-09-01 |
| #138 | Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison | Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit | cs.CV | 2026-09-01 |
| #139 | DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting | Qian Wang, Yu Wang, Weiqi Li +4 | cs.CV | 2026-09-01 |
| #140 | TempCloze: Can Video-LLMs Identify the Missing Middle? | Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4 | cs.CV | 2026-09-01 |
| #141 | A Sensor-Adaptive Incremental Learning Framework for Artifact Detection in Satellite Precipitation Data | Andres F. Monsalve, Hernan A. Moreno, Christian D. Kummerow | physics.ao-ph | 2026-09-01 |
| #142 | Benchmarking Spatial, Spectral, and Self-Supervised Cues for Face Forgery Detection under Realistic Degradation | Lucas Cunha, Lucas Sotomaior, Lucas Gasperin +3 | cs.CV | 2026-09-01 |
| #143 | CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling | Xin Shen, Chengyou Jia, Keshuo Xing +6 | cs.CV | 2026-09-01 |
| #144 | FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making | Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza | cs.CV | 2026-09-01 |
| #145 | RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching | Charles Corbière, Léo Machado, Aubin Charley +3 | cs.CV | 2026-09-01 |
| #146 | Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept Erasure | Qinghui Gong, Xunlei Chen, Yu-Xuan Zhang +2 | cs.CV | 2026-09-01 |
| #147 | Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications | S. Sifaoui, E. Angelini, S. Toupin +2 | cs.CV | 2026-09-01 |
| #148 | Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading | Fatemeh Javadian, Zhu Chen, Zahra Aminparast +1 | cs.CV | 2026-09-01 |
| #149 | MegaStyle++: Scaling Image Style Space through Hierarchical Style Definition | Junyao Gao, Sibo Liu, Jiaxing Li +4 | cs.CV | 2026-09-01 |
| #150 | EdiTikZ: Scientific Figure Editing from Revision Trajectories | Christian Greisinger, Zhixue Zhao, Steffen Eger | cs.AI | 2026-09-01 |
| #151 | Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations | Qingde Li, Qingqi Hong, Zihan Li +1 | cs.AI | 2026-09-01 |
| #152 | Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery | Matheus F. Kovaleski, Cristiano Premebida, João Ruivo Paulo | cs.CV | 2026-09-01 |
| #153 | Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3 | Matheus F. Kovaleski, Luís Garrote, Cristiano Premebida +2 | cs.CV | 2026-09-01 |
| #154 | InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations | Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3 | cs.CL | 2026-09-01 |
| #155 | IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals | Md. Atabuzzaman, Christian Alexander, Chris Thomas | cs.CV | 2026-09-01 |
| #156 | Diffusion Based Unpaired Data Learning for Inverse Problems | Chenglong Bao, Yiming Dang, Chenguang Duan +2 | cs.CV | 2026-09-01 |
| #157 | Accurate Reconstruction of Gas Turbine Blade Geometry Using 3D/2D Rigid Registration and CT View Optimization | Hristo Valtchanov, Nicolas Piché, Vladimir Brailovski +3 | cs.CV | 2026-09-01 |
| #158 | ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence | Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao +1 | cs.CV | 2026-09-01 |
| #159 | Reliability Challenges in Diffusion Vision-Language Models | Md. Atabuzzaman, Chris Thomas | cs.CV | 2026-09-01 |
| #160 | MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval | Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro +2 | cs.IR | 2026-09-01 |
| #161 | GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation | Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10 | eess.IV | 2026-09-01 |
| #162 | CMRVision: A Foundation Model for Cardiac MR Image Analysis | Athira J. Jacob, Puneet Sharma, Daniel Rueckert | cs.CV | 2026-09-01 |
| #163 | MeshSplatBench: A Unified Benchmark for Triangle-Based Neural Rendering | Kaixuan Zhang, Minxian Li, Mingwu Ren +1 | cs.GR | 2026-09-01 |
| #164 | Agentic Multimodal Models for Environmental Hyperspectral Unmixing | Michał Cholewa, Luca Ciampi, Nicola Messina +2 | cs.CV | 2026-09-01 |
| #165 | HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives | Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera +1 | cs.CV | 2026-09-01 |
| #166 | TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models | Chao Zhou, Yiling Chen, Qi Chu +3 | cs.CV | 2026-09-01 |
| #167 | Seeing the World and the Self from Egocentric Video | Kai Guan, Minchao Jiang, Ruichen WangLi +2 | cs.CV | 2026-09-01 |
| #168 | FORGE: Forward-Only Test-Time Adaptation for Integer-Only Vision Models on Microcontrollers | Muhammad Rehan, Haider Ali, Muhammad Ali Munir +1 | cs.CV | 2026-09-01 |
| #169 | MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video Generation | Zhijian Qiao, Xinjiang Wang, Jiajie Chen +5 | cs.CV | 2026-09-01 |
| #170 | One Prompt Is Enough: Watermark Laundering Through Foundation Image Models | Jidong Yang, Qi Li, Wei Zong +5 | cs.CV | 2026-09-01 |
| #171 | S$^2$Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models | Yuanyuan Jia, Shunpu Tang, Qianqian Yang | cs.CV | 2026-09-01 |
| #172 | Compressing AI Traffic: Standardized Neural Network Coding of Visual-Token Representations in Split Vision-Language Inference | Reza Heidari, Hamed R. Tavakoli, Juho Kannala | cs.CV | 2026-09-01 |
| #173 | Monocular Depth Estimation from a Single Image: Progress and Opportunities | Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun +4 | cs.CV | 2026-09-01 |
| #174 | Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement | Chujie Qin, Zilong Zhang, Zewei Chang +5 | cs.CV | 2026-09-01 |
| #175 | On the Design Fundamentals of Pixel Text Representation Learning | Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4 | cs.CV | 2026-09-01 |
| #176 | StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization | Hongtao Kang, Die Luo, Li Chen +4 | cs.CV | 2026-09-01 |
| #177 | Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising | Xianlei Han, Shaoyu Wang, Jiancheng Fang +2 | cs.CV | 2026-09-01 |
| #178 | Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set | Michael Zang, Haiyu Wu, Mrinal Sharma +1 | cs.CV | 2026-09-01 |
| #179 | Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning | Jiyoung Park, InJae Oh, Jung Uk Kim | cs.CV | 2026-09-01 |
| #180 | P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement | Ruoyu Guo, Haonan Zhong, Maurice Pagnucco +1 | cs.CV | 2026-09-01 |
| #181 | When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP | Shota Sato, Hajime Kiyama, Tosho Hirasawa +1 | cs.CL | 2026-09-01 |
| #182 | IT-TextFusion: Iterative Text-Image Interaction with Text-Guided Residual Refinement for Degradation-Aware Image Fusion | Siyang Liu, Peiyi Zhou, Tianle Jin +3 | cs.CV | 2026-09-01 |
| #183 | Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration | Daehwan Kim, Haejun Chung, Ikbeom Jang | cs.LG | 2026-09-01 |
| #184 | Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment | Mohamad Jouni, Aurélien Godet, Mauro Dalla Mura | eess.IV | 2026-09-01 |
| #185 | Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models | Jiayu Ding, Zhuodong Liu, Lei Zhang +6 | cs.CV | 2026-09-01 |
| #186 | SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations | Yiming Luo, Rongqiang Zhao, Jie Liu | cs.LG | 2026-09-01 |
| #187 | ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives | Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos +1 | cs.CV | 2026-09-01 |
| #188 | MultiGait: A Multi-Sensor Multi-Perspective Multi-Session Biometric Inference Benchmark and its Dataset | Julian Todt, Felix Morsbach, Philip Dissert +1 | cs.CR | 2026-09-01 |
| #189 | Fi-ImageNet-1k: An OOD Benchmark From the Inside of the ImageNet-1k Validation Set | Ruslan Rozumnyi, Matěj Suchánek, Tomáš Vojíř +2 | cs.CV | 2026-09-01 |
| #190 | PyDoseRT Proton: A GPU Pencil-Beam Engine with a Convolutional Residual-Correction Network for Fast Proton Dose Calculation | Lukas Zimmermann, Hermann Fuchs, Attila Simkó +1 | physics.med-ph | 2026-09-01 |
| #191 | Low-Quality Face Recognition using Center Aligned Representations and Local Margin Constraints | Vedat Can Dilaver, Benjamin S. Riggan | cs.CV | 2026-09-01 |
| #192 | SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models | Shiyu Li, Zi-Yuan Hu, Shijia Huang +3 | cs.CV | 2026-09-01 |
| #193 | Does This Moment Justify the Recommendation? Counterfactual Behavior-Grounded Evidence Retrieval for Personalized Video Recommendation | Xin Liu | cs.CV | 2026-09-01 |
| #194 | CQF-HMR: Continuous Quaternion Flows for Probabilistic 3D Human Mesh Recovery from a Single Image | Cuong Le, Bao-Long Tran, Pavlo Melnyk +3 | cs.CV | 2026-09-01 |
| #195 | EvoGS: Modeling Deformation Evolution for Dynamic Gaussian Splatting | Wei Dong, Shahram Shirani, Jun Chen +1 | cs.CV | 2026-09-01 |
| #196 | Semi-Supervised Virtual Staining via Morphology Preservation and Histopathological Realism Constraints | Baoshun Wang, Weiping Lin, Linwu Wang +3 | cs.CV | 2026-09-01 |
| #197 | Prior-Guided Implicit Neural Representations for Single-Subject Diffusion MRI Super-Resolution | Abdulkader Ghandoura, Marsil Zakour, William Consagra +1 | eess.IV | 2026-09-01 |
| #198 | ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation | Jeonghyeok Do, Seungchul Lee, Munchurl Kim | cs.CV | 2026-09-01 |
| #199 | Conditional Flow Matching for Cross-Field MRI Harmonisation | Baris Imre, Aram Salehi, Levente Baljer +3 | cs.CV | 2026-09-01 |
| #200 | Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA | Hai-Dang Nguyen, Huy-Hieu Pham | cs.CV | 2026-09-01 |