Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Vision and Imaging
841 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China52 % · 289 Artikel
- Vereinigte Staaten33 % · 185 Artikel
- Vereinigtes Königreich8 % · 45 Artikel
- Südkorea7,8 % · 44 Artikel
- Sonderverwaltungsregion Hongkong7,5 % · 42 Artikel
- Deutschland6,4 % · 36 Artikel
- Japan4,3 % · 24 Artikel
- Frankreich3,9 % · 22 Artikel
Über 561 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 52 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Decoupling Spherical Reasoning from Dense Prediction for 360 Depth Estimation
Zhijie Shen, Chunyu Lin, Shuai Zheng, Feng Li, Runmin Cong, Huihui Bai, Yao Zhao · 1. Oktober 2026
The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feature modeling challenging. We revisit panoramic depth estimation by decoupling contextual modeling in native spherical space from dense ERP prediction…
- StereoGaussians: Feed-Forward 3D Gaussian Splatting from Stereo Images
Boyuan Tian, Huangying Zhan, Zhan Li, Shin-Fang Chng, Hanwen Yang, Zirui Wang, Yi Xu · 1. Oktober 2026
Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appear…
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang · 30. September 2026
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a fra…
- GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking
Zhuoqian Feng, Weixing Chen, Ziliang Chen, Yang Liu, Liang Lin · 29. September 2026
Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the sc…
- Depth Any Seen: Which Surfaces and How Far?
Xiaohao Xu, Xiaonan Huang · 29. September 2026
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain ab…
- Learning with Volterra Neural Networks: A System Theoretic Perspective
Haoyu Yun, Hamid Krim, Yufang Bao · 28. September 2026
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural operator for compact higher-order filter…
- Self-Supervised Perceptually Interpretable Monocular Depth Estimation
Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis · 28. September 2026
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, …
- OC-GS: Gaussian Splatting for Irregular Turntable Capture
Jae Joong Lee, Bedrich Benes · 28. September 2026
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimi…
- Geometric Inconsistency Localization in Multi-View Image Sets
Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael · 28. September 2026
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potenti…
- FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors
Dan Halperin, Mirko M\"ahlisch · 25. September 2026
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-co…
- LiFR v2: Completion-Augmented Event Propagation for High-Rate Dense Prediction
Tao Wan, Xiaoshan Wu, Yifei Yu, Bo Wang, Xiaoyang Lyu, Muxin Liu, Aoxuan Pan, Zhongrui Wang, Xiaojuan Qi · 24. September 2026
High-rate dense perception in dynamic environments is limited by the low update rate of RGB cameras, as rapid scene changes can occur between frames. Event cameras offer temporally dense but spatially sparse measurements, complementary to spatially dense RGB observations. Direct fusion cannot fully …
- Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB
Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz · 24. September 2026
As mobile robots become increasingly integrated into everyday environments, privacy risks arising from onboard cameras have become a growing concern. Ultra-low-resolution (ULR) RGB can mitigate visual privacy exposure at the source, but ULR appearance alone substantially limits semantic and spatial …
- Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness
Liang Zeng, Maarten Vergauwen · 24. September 2026
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objecti…
- SURE-Map: Self-Correcting Streaming Geometric Foundation Models
Mingkai Liu, Hao Zhao, Xingxing Zuo · 22. September 2026
Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issue: each prediction is made from limited context, which is vulnerable to dynamic objects and weak textures. Small local errors accumulate into severe …
- Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation
Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua · 22. September 2026
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera par…
- When Wider Views Fail: Stress-Testing Feed-Forward 3D Reconstruction
Daisy Li, Kyle Gao, Quanyun Wu, Hanna Chomko, John S. Zelek, Jonathan Li · 22. September 2026
Feed-forward 3D reconstruction models enable efficient geometry estimation from sparse images, but their pretrained nature can make them vulnerable to distribution shifts beyond their training data. Identifying these failure modes is important for understanding when such models can be reliably deplo…
- CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention
Xuezhi Xiang, Jiayao Liu, Heqi Xiang, Yuqi Hu, Yiming Chen, Shanjun Zhang · 22. September 2026
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and globa…
- AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission
Andr\'e Amorim, Pedro F. Proen\c{c}a · 22. September 2026
Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. However analog VTX suffers from complex spatially structured image degradation which differ fundamentally from digital image corruption (e.g. AWGN) used in standard training augmentation. This w…
- VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai · 22. September 2026
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechan…
- RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
Hongbo Mao (Harbin Institute of Technology), Junjun Jiang (Harbin Institute of Technology), Youyu Chen (Harbin Institute of Technology), Jiaxin Zhang (Harbin Institute of Technology), Zhemeng Dong (Harbin Institute of Technology), Xianming Liu (Harbin Institute of Technology) · 22. September 2026
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context…
- PanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline
Heechan Yoon, Dongki Jung, Phuc Nguyen, Ming Lin, Dinesh Manocha · 22. September 2026
We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg3R jointly predicts 3D geometry and multi-view semantic segmentation in one single forward pass. Built upon a pretrained reconstruction backbone that…
- 3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation
Yueen Ma, Zenglin Xu, Irwin King · 22. September 2026
Spatial intelligence, encompassing 3D perception and reasoning, is the essential next frontier of AI. Scaling current 3D vision-language models (VLMs) that rely on dense Transformers for spatial tasks incurs prohibitive computational costs. In this paper, we introduce 3D-MoE, a 3D VLM leveraging an …
- Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction
Sunghyun Baek, Hanna Bae, Minchan Kwon, Junmo Kim · 21. September 2026
Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video streams in an online manner for real-world deployment. However, existing methods overlook two key signals when handling long image streams: the importanc…
- 4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors
Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan · 21. September 2026
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed p…
- Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction
Jingke Zhou, Chenhang Ma, Zhizhou Zhong, Mingkai Liu, Zhuang Zhou, Yicheng ji, Binghua Su, Bo Cai, Xianliang Huang · 21. September 2026
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effe…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
- Image Enhancement Techniques463 Papiere / 12 Monate+20 %
