Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Vision and Imaging
841 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- Chine52 % · 289 articles
- États-Unis33 % · 185 articles
- Royaume-Uni8 % · 45 articles
- Corée du Sud7,8 % · 44 articles
- R.A.S. chinoise de Hong Kong7,5 % · 42 articles
- Allemagne6,4 % · 36 articles
- Japon4,3 % · 24 articles
- France3,9 % · 22 articles
Sur 561 articles de ce sujet dont au moins un laboratoire est situé. 52 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- MoSE3: Learning World-Space SE(3) at Every Pixel
Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang · 5 octobre 2026
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first fe…
- Feedforward Novel View Synthesis for Heterogeneous Cameras
Meng Wei, Cheng Zhang, Boying Li, Yihang Chen, Jianmin Zheng, Hamid Rezatofighi, Jianfei Cai · 5 octobre 2026
Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and pa…
- Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation
Daikun Liu, Teng Wang, Changyin Sun · 5 octobre 2026
Event cameras hold excellent dynamic properties, showing great potential for monocular depth estimation (MDE). However, existing methods mainly improve performance by optimizing contextual features, but still struggle with the ill-posed and nonlinear nature of direct full-depth regression. In this p…
- SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models
Xi Ye, Yuzhu Wang, Xiaoyang Liu, Jiayi Wang, Yangyang Xu, Ruyu Wang, Wenlin Chen, Duo Su, Jun Zhu · 5 octobre 2026
Flow-matching-based multi-view world models generate realistic videos, but are commonly restricted to fixed camera rigs. Extending them to continuously varying camera poses requires paired pose--video observations with dense pose coverage, which are costly to acquire. We introduce \emph{SymRegFlow},…
- OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection
Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang · 5 octobre 2026
Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirect…
- Decoupling Spherical Reasoning from Dense Prediction for 360 Depth Estimation
Zhijie Shen, Chunyu Lin, Shuai Zheng, Feng Li, Runmin Cong, Huihui Bai, Yao Zhao · 1 octobre 2026
The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feature modeling challenging. We revisit panoramic depth estimation by decoupling contextual modeling in native spherical space from dense ERP prediction…
- StereoGaussians: Feed-Forward 3D Gaussian Splatting from Stereo Images
Boyuan Tian, Huangying Zhan, Zhan Li, Shin-Fang Chng, Hanwen Yang, Zirui Wang, Yi Xu · 1 octobre 2026
Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appear…
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
Yajiao Xiong, Youyu Luan, Xiaoyu Zhou, Yongtao Wang · 30 septembre 2026
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a fra…
- GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking
Zhuoqian Feng, Weixing Chen, Ziliang Chen, Yang Liu, Liang Lin · 29 septembre 2026
Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the sc…
- Depth Any Seen: Which Surfaces and How Far?
Xiaohao Xu, Xiaonan Huang · 29 septembre 2026
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain ab…
- Learning with Volterra Neural Networks: A System Theoretic Perspective
Haoyu Yun, Hamid Krim, Yufang Bao · 28 septembre 2026
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural operator for compact higher-order filter…
- Self-Supervised Perceptually Interpretable Monocular Depth Estimation
Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis · 28 septembre 2026
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, …
- OC-GS: Gaussian Splatting for Irregular Turntable Capture
Jae Joong Lee, Bedrich Benes · 28 septembre 2026
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimi…
- Geometric Inconsistency Localization in Multi-View Image Sets
Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael · 28 septembre 2026
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potenti…
- FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors
Dan Halperin, Mirko M\"ahlisch · 25 septembre 2026
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-co…
- LiFR v2: Completion-Augmented Event Propagation for High-Rate Dense Prediction
Tao Wan, Xiaoshan Wu, Yifei Yu, Bo Wang, Xiaoyang Lyu, Muxin Liu, Aoxuan Pan, Zhongrui Wang, Xiaojuan Qi · 24 septembre 2026
High-rate dense perception in dynamic environments is limited by the low update rate of RGB cameras, as rapid scene changes can occur between frames. Event cameras offer temporally dense but spatially sparse measurements, complementary to spatially dense RGB observations. Direct fusion cannot fully …
- Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB
Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz · 24 septembre 2026
As mobile robots become increasingly integrated into everyday environments, privacy risks arising from onboard cameras have become a growing concern. Ultra-low-resolution (ULR) RGB can mitigate visual privacy exposure at the source, but ULR appearance alone substantially limits semantic and spatial …
- Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness
Liang Zeng, Maarten Vergauwen · 24 septembre 2026
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objecti…
- SURE-Map: Self-Correcting Streaming Geometric Foundation Models
Mingkai Liu, Hao Zhao, Xingxing Zuo · 22 septembre 2026
Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issue: each prediction is made from limited context, which is vulnerable to dynamic objects and weak textures. Small local errors accumulate into severe …
- Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation
Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua · 22 septembre 2026
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera par…
- When Wider Views Fail: Stress-Testing Feed-Forward 3D Reconstruction
Daisy Li, Kyle Gao, Quanyun Wu, Hanna Chomko, John S. Zelek, Jonathan Li · 22 septembre 2026
Feed-forward 3D reconstruction models enable efficient geometry estimation from sparse images, but their pretrained nature can make them vulnerable to distribution shifts beyond their training data. Identifying these failure modes is important for understanding when such models can be reliably deplo…
- CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention
Xuezhi Xiang, Jiayao Liu, Heqi Xiang, Yuqi Hu, Yiming Chen, Shanjun Zhang · 22 septembre 2026
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and globa…
- AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission
Andr\'e Amorim, Pedro F. Proen\c{c}a · 22 septembre 2026
Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. However analog VTX suffers from complex spatially structured image degradation which differ fundamentally from digital image corruption (e.g. AWGN) used in standard training augmentation. This w…
- VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
Abteen Arab, Guile Wu, Chengjie Huang, Dongfeng Bai · 22 septembre 2026
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechan…
- RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
Hongbo Mao (Harbin Institute of Technology), Junjun Jiang (Harbin Institute of Technology), Youyu Chen (Harbin Institute of Technology), Jiaxin Zhang (Harbin Institute of Technology), Zhemeng Dong (Harbin Institute of Technology), Xianming Liu (Harbin Institute of Technology) · 22 septembre 2026
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context…
Autres sujets du thème Vision par ordinateur et reconnaissance de formes
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Multimodal Machine Learning Applications8 069 papiers / 12 mois+191 %
- Generative Adversarial Networks and Image Synthesis4 992 papiers / 12 mois+39 %
- Advanced Neural Network Applications2 354 papiers / 12 mois+48 %
- Human Pose and Action Recognition836 papiers / 12 mois+457 %
- Face recognition and analysis482 papiers / 12 mois+88 %
- Image Enhancement Techniques463 papiers / 12 mois+20 %
