Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Image and Video Retrieval Techniques
286 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China48 % · 70 artículos
- Estados Unidos31 % · 45 artículos
- RAE de Hong Kong (China)6,9 % · 10 artículos
- Japón4,8 % · 7 artículos
- Alemania4,8 % · 7 artículos
- Corea del Sur4,1 % · 6 artículos
- Reino Unido4,1 % · 6 artículos
- Francia3,4 % · 5 artículos
Sobre 145 artículos de este tema con al menos un laboratorio localizado. 34 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Introduction to Computer Vision
Stan Birchfield · 1 de octubre de 2026
This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organized as 44 short chapters across three parts, the book builds each topic from first principles: image arithmetic and morphology; convolution, pyramids…
- Front-to-Back: Benchmarking Vision-Language Models for Asymmetric Cross-View Vehicle Re-Identification
Moseli Mots'oehli, Thulani Babeli · 1 de octubre de 2026
Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehicle handovers from 20 recording sequences in South Africa. Each exam…
- PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers
Yun Dai, Jiarui Wen, Huiping Zhuang, Cen Chen, Ziqian Zeng · 1 de octubre de 2026
Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval …
- Emergent Multi-View Geometry Through Self-Distillation
David Nordstr\"om, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl · 1 de octubre de 2026
Over a century ago, Henri Poincar\'e argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propo…
- Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um · 30 de septiembre de 2026
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-…
- NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
Jiawei Zhang, Shuhao Liu, Rong Huang, Yuancheng Li, Zhihui Li, Xiaojun Chang, Changlin Li · 30 de septiembre de 2026
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on neste…
- FM-ReID: Selective Competitive Token Routing for Object Re-Identification
Zhiqi Li, Xiaowei Zhou, Zeyuan Sun, Feng Gao, Junyu Dong · 30 de septiembre de 2026
Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, conto…
- One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
Aditya Sharma, Divya Saxena · 30 de septiembre de 2026
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment,…
- Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval
Cassandra Yang, Yufan Tang · 30 de septiembre de 2026
Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles stretch with rate, speech changes with …
- The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation
Kuniaki Saito, Yoshitaka Ushiku · 29 de septiembre de 2026
Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requi…
- ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
Kieran Sagar Parikh, Jose Luis Garcia del Castillo y Lopez · 29 de septiembre de 2026
Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representat…
- Refinement Symmetry in Multimodal Transformers
Yuhao Du, Shunian Chen · 29 de septiembre de 2026
Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention…
- QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez · 28 de septiembre de 2026
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection t…
- ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding
Yusung Choi · 25 de septiembre de 2026
The strong performance that modern semantic correspondence methods achieve at standard thresholds plateaus sharply at fine-grained thresholds. We argue that this plateau stems not from the representational capacity of backbone features, but from a grid-tied readout. Patch-based vision transformers t…
- LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations
Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur · 24 de septiembre de 2026
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated …
- Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces
Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem · 24 de septiembre de 2026
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature spac…
- Geometry-Conditioned Visual Place Recognition in Natural Environments
Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam · 24 de septiembre de 2026
Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, sparse distinctive landmarks, and substantial appearance and viewpoint variation across traversals. While visual observations of the same place can change considerably, their underlying spatial s…
- SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem · 23 de septiembre de 2026
Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance …
- SPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction
David Szczecina, Yuanpei Xiang, Jitao Hu, David Clausi, Yuhao Chen, Jason Deglint, Paul Fieguth · 23 de septiembre de 2026
Self-supervised learning (SSL) has become an effective approach for learning visual representations without manual annotations. Among SSL approaches, contrastive learning has been widely used for visual representation learning. However, existing contrastive SSL methods have focused primarily on imag…
- Parameterized Dense-Sparse Fusion for Hybrid Retrieval: Tuning a Rank-Score Mix on BEIR SciFact with Qdrant
Satyanarayan Pati, Srikanth Patil · 23 de septiembre de 2026
We study a parameterized hybrid ranker that fuses a dense embedding list and a sparse lexical list. The method has a small, explicit parameter vector: a dense prior $\alpha \in [0,1]$, a score-versus-rank mix $\lambda \in [0,1]$, an RRF smoothing parameter $\kappa > 0$, optional list-geometry coeffi…
- POI-Loc: A Fine-Grained POI Localization Benchmark and an Asymmetric Global-to-Local Matching Method
Lu Han, Xiting Sun, Hao Wang, Zhiqiang Cao, Ruihuan Du, Ziquan Zeng, Chunlong Lv · 22 de septiembre de 2026
Point-of-interest (POI) localization matches user-provided storefront close-ups to the same shops in wide, geo-tagged vehicle-mounted street views. POIs may change while the surrounding scene stays similar, so scene-level recognition alone cannot establish POI identity. Differences in target scale a…
- ReSTI: A Source-Grounded Audit and Repair of STI-Bench
Pengzhan Sun, Ramanathan Rajaraman, Shiu-Hong Kao, Junbin Xiao, Angela Yao · 22 de septiembre de 2026
Spatial--temporal benchmarks are valid only when their questions, source annotations, and answer options identify the same physical quantity. We audit STI-Bench against the official ScanNet, Waymo, and Omni6DPose sources and find systematic coordinate-system and timestamp errors, under-specified tar…
- Active Visual Sampling with a Connectome-Constrained Fly Model for One-Shot Hatch Recognition in Architectural Drawings
Dmitry Kuklev · 22 de septiembre de 2026
Architectural drawings encode material classes through repeated hatch patterns. We test whether a connectome-constrained fly visual network, pretrained for motion, can be repurposed without task-specific weight updates as a descriptor for one-shot hatch matching. Each 64 x 64 patch is translated ove…
- HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space
Zelin Peng, Zhengqin Xu, Changsong Wen, Yu Huang, Yaoming Wang, Xiaokang Yang, Wei Shen · 22 de septiembre de 2026
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly …
- GeoBalance: Geometry-Aware Monitoring and Reconstruction with Asymmetric Optimization for Balanced Multimodal Learning
Zechang Xiong, Da Li, Rong Yin, Kexin Tang, Biao Yang, Pengyuan Li, Wenkang Kong, Yulan Hu, Shengyu Zhu, Hao Peng · 22 de septiembre de 2026
Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of others. Existing balancing methods mainly adjust losses, gradients, or modality contributions, largely treating modality imbalance as an optimization p…
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
