Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Vision and Imaging
841 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China52% · 289 papers
- United States33% · 185 papers
- United Kingdom8% · 45 papers
- South Korea7.8% · 44 papers
- Hong Kong SAR China7.5% · 42 papers
- Germany6.4% · 36 papers
- Japan4.3% · 24 papers
- France3.9% · 22 papers
Across 561 papers on this subject with at least one lab located. 52 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Decoupling Spherical Reasoning from Dense Prediction for 360 Depth Estimation
Zhijie Shen, Chunyu Lin, Shuai Zheng, Feng Li, Runmin Cong, Huihui Bai, Yao Zhao · 1 October 2026
The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feature modeling challenging. We revisit panoramic depth estimation by decoupling contextual modeling in native spherical space from dense ERP prediction…
- StereoGaussians: Feed-Forward 3D Gaussian Splatting from Stereo Images
Boyuan Tian, Huangying Zhan, Zhan Li, Shin-Fang Chng, Hanwen Yang, Zirui Wang, Yi Xu · 1 October 2026
Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appear…
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang · 30 September 2026
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a fra…
- GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking
Zhuoqian Feng, Weixing Chen, Ziliang Chen, Yang Liu, Liang Lin · 29 September 2026
Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the sc…
- Depth Any Seen: Which Surfaces and How Far?
Xiaohao Xu, Xiaonan Huang · 29 September 2026
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain ab…
- Learning with Volterra Neural Networks: A System Theoretic Perspective
Haoyu Yun, Hamid Krim, Yufang Bao · 28 September 2026
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural operator for compact higher-order filter…
- Self-Supervised Perceptually Interpretable Monocular Depth Estimation
Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis · 28 September 2026
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, …
- OC-GS: Gaussian Splatting for Irregular Turntable Capture
Jae Joong Lee, Bedrich Benes · 28 September 2026
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimi…
- Geometric Inconsistency Localization in Multi-View Image Sets
Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael · 28 September 2026
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potenti…
- FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors
Dan Halperin, Mirko M\"ahlisch · 25 September 2026
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-co…
- LiFR v2: Completion-Augmented Event Propagation for High-Rate Dense Prediction
Tao Wan, Xiaoshan Wu, Yifei Yu, Bo Wang, Xiaoyang Lyu, Muxin Liu, Aoxuan Pan, Zhongrui Wang, Xiaojuan Qi · 24 September 2026
High-rate dense perception in dynamic environments is limited by the low update rate of RGB cameras, as rapid scene changes can occur between frames. Event cameras offer temporally dense but spatially sparse measurements, complementary to spatially dense RGB observations. Direct fusion cannot fully …
- Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB
Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz · 24 September 2026
As mobile robots become increasingly integrated into everyday environments, privacy risks arising from onboard cameras have become a growing concern. Ultra-low-resolution (ULR) RGB can mitigate visual privacy exposure at the source, but ULR appearance alone substantially limits semantic and spatial …
- Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness
Liang Zeng, Maarten Vergauwen · 24 September 2026
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objecti…
- SURE-Map: Self-Correcting Streaming Geometric Foundation Models
Mingkai Liu, Hao Zhao, Xingxing Zuo · 22 September 2026
Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issue: each prediction is made from limited context, which is vulnerable to dynamic objects and weak textures. Small local errors accumulate into severe …
- Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation
Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua · 22 September 2026
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera par…
- When Wider Views Fail: Stress-Testing Feed-Forward 3D Reconstruction
Daisy Li, Kyle Gao, Quanyun Wu, Hanna Chomko, John S. Zelek, Jonathan Li · 22 September 2026
Feed-forward 3D reconstruction models enable efficient geometry estimation from sparse images, but their pretrained nature can make them vulnerable to distribution shifts beyond their training data. Identifying these failure modes is important for understanding when such models can be reliably deplo…
- CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention
Xuezhi Xiang, Jiayao Liu, Heqi Xiang, Yuqi Hu, Yiming Chen, Shanjun Zhang · 22 September 2026
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and globa…
- AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission
Andr\'e Amorim, Pedro F. Proen\c{c}a · 22 September 2026
Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. However analog VTX suffers from complex spatially structured image degradation which differ fundamentally from digital image corruption (e.g. AWGN) used in standard training augmentation. This w…
- VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai · 22 September 2026
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechan…
- RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
Hongbo Mao (Harbin Institute of Technology), Junjun Jiang (Harbin Institute of Technology), Youyu Chen (Harbin Institute of Technology), Jiaxin Zhang (Harbin Institute of Technology), Zhemeng Dong (Harbin Institute of Technology), Xianming Liu (Harbin Institute of Technology) · 22 September 2026
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context…
- PanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline
Heechan Yoon, Dongki Jung, Phuc Nguyen, Ming Lin, Dinesh Manocha · 22 September 2026
We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg3R jointly predicts 3D geometry and multi-view semantic segmentation in one single forward pass. Built upon a pretrained reconstruction backbone that…
- 3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation
Yueen Ma, Zenglin Xu, Irwin King · 22 September 2026
Spatial intelligence, encompassing 3D perception and reasoning, is the essential next frontier of AI. Scaling current 3D vision-language models (VLMs) that rely on dense Transformers for spatial tasks incurs prohibitive computational costs. In this paper, we introduce 3D-MoE, a 3D VLM leveraging an …
- Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction
Sunghyun Baek, Hanna Bae, Minchan Kwon, Junmo Kim · 21 September 2026
Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video streams in an online manner for real-world deployment. However, existing methods overlook two key signals when handling long image streams: the importanc…
- 4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors
Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan · 21 September 2026
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed p…
- Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction
Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang · 21 September 2026
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effe…
Other topics in Computer vision and pattern recognition
The topics the OpenAlex classification attaches to the same theme, most active first.
- Multimodal Machine Learning Applications8,069 papers / 12 months+191%
- Generative Adversarial Networks and Image Synthesis4,992 papers / 12 months+39%
- Advanced Neural Network Applications2,354 papers / 12 months+48%
- Human Pose and Action Recognition836 papers / 12 months+457%
- Face recognition and analysis482 papers / 12 months+88%
- Image Enhancement Techniques463 papers / 12 months+20%
