Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Video Surveillance and Tracking Methods
364 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- China52 % · 119 artículos
- Estados Unidos21 % · 48 artículos
- Alemania5,7 % · 13 artículos
- Corea del Sur4,8 % · 11 artículos
- Francia4,4 % · 10 artículos
- Singapur4 % · 9 artículos
- Reino Unido3,5 % · 8 artículos
- RAE de Hong Kong (China)3,5 % · 8 artículos
Sobre 227 artículos de este tema con al menos un laboratorio localizado. 48 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Hard-Region Supervision: #1 on the Waymo Open Dataset 2D Video Panoptic Segmentation Leaderboard
Jinghan Yang · 1 de octubre de 2026
We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of every frame and, for countable objects, an identity that holds across 100 frames and across five overlapping cameras. We build on DVIS++, a cascade o…
- TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos
Patt Phurtivilai, Zhiyang Dou, Yifan Wu, Kinfung Chu, Yuan Liu, Lei Yang, Wenping Wang, Taku Komura · 1 de octubre de 2026
Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework fo…
- Template-Search Domain Adaptation via Multi-Stage Feature Alignment for Cross-Modal Object Tracking
Fereshteh Aghaee Meibodi, Amir Mehdi Soufi Enayati, Shadi Alijani, Homayoun Najjaran · 1 de octubre de 2026
Visual object tracking typically assumes that the initial template and subsequent search frames share the same sensing modality. In practice, sensor availability or operation may change over time, creating a substantial representation gap between template and search frames. Unlike conventional multi…
- SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation
Zhengtong Zhu, Jiaqing Fan, Hanwen Qian, Fanzhang Li · 29 de septiembre de 2026
Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learni…
- Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge
Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli · 25 de septiembre de 2026
Video object detection on edge devices runs computationally expensive detectors over long frame streams, causing high energy consumption and sustained GPU utilization. Although consecutive frames are highly redundant, naive frame skipping is content-blind: it skips during critical moments such as ob…
- TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley · 25 de septiembre de 2026
Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tr…
- LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT
Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni · 24 de septiembre de 2026
Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM2 pipelines, failures typically arise at three stages of the object lifecycle: (i) erroneous or dupl…
- Identity-Consistent Analysis of Long-Shot Windsurfing Video: A Domain-Specific Offline Tracking System
Bertil Braun · 22 de septiembre de 2026
Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not a generic MOT trace but a separate, stable rider-relative video for each surfer; one false identity merge can invalidate an otherwise useful result. …
- DOA-SORT: Directional Occlusion-Aware Multi-Object Tracking with Distributional Observations
Hao Wang · 22 de septiembre de 2026
Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existing motion-dominant trackers commonly represent occlusion as a scalar penalty. This treatment misses the directional observation bias caused by occlus…
- Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT
Linh Van Ma, Juhua Hu, Wei Cheng, Unse Fatima, Moongu Jeon · 22 de septiembre de 2026
This paper presents a comprehensive experimental evaluation and detailed analysis of state-of-the-art multi-object tracking algorithms, with an emphasis on quantifying the individual contributions of detection and association components to overall tracking performance. Unlike existing surveys that p…
- Learning to Track from Privileged Target Appearances
Xin Chen, Jiao Xu, Dong Wang, Huchuan Lu, Kede Ma · 18 de septiembre de 2026
Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better reflect the current appearance but are croppe…
- On-the-Fly Homographies Calibration for Multi-Camera Tracking
David Voihanski, Mor Sinai, Ben Zion Bobrovsky · 17 de septiembre de 2026
Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Privacy regulations frequently prohibit recording video for offline calibration; limited bandwidth precludes synchronizing high-resol…
- Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion
Mais Mohammed, Sharifa Mohammed, Hanan Awadh, Haneen Bamaas, Raghad Bawazeer, Elham Alghamdi · 16 de septiembre de 2026
Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awaren…
- GRACE: Geometry- and Ray-Aware Camera-Efficient Multi-View Pedestrian Tracking
Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno, Naoki Kato · 16 de septiembre de 2026
Reducing the number of cameras reduces the deployment cost but removes views that correct BEV responses stretched away from true pedestrian positions by projection and short score drops that can split tracks} in Bird's-Eye View (BEV) tracking. We introduce GRACE, a camera-efficient multi-view tracke…
- MAETrack: Unleashing the Potential of Pretrained Geometric Priors for 3D Single Object Tracking
Sifan Zhou, Qiwei Wang, Linyue Tan, Ziyu Liu, Ziyu Zhao, Xiaobo Lu · 16 de septiembre de 2026
Large-scale pre-training has transformed representation learning in 2D vision, yet its transferability to 3D single object tracking (SOT) remains insufficiently understood. Directly fine-tuning self-supervised 3D encoders, such as masked autoencoders (MAE), often leads to sub-optimal adaptation beca…
- SAVTrack: Selective Vote Aggregation for Reliability-Aware Point Cloud Tracking
Sifan Zhou, Linyue Tan, Qiwei Wang, Ziyu Zhao, Xiaobo Lu · 16 de septiembre de 2026
3D single object tracking (SOT) in LiDAR point clouds is essential for autonomous systems, but remains challenging under sparse and incomplete observations. In such cases, different target points provide highly uneven constraints on the object center, causing some point-to-center votes to be substan…
- A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification
Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne · 15 de septiembre de 2026
Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult due to viewpoint and illumination changes, occlusion, background clutter, and low resolution imagery. We propose a generative AI integrated multimodal ReID framework designed explicitly for…
- Revisiting Multi-Object Tracking Baselines: Hyperparameter Optimization with Multi-Fidelity Greedy Coordinate Search
Momir Ad\v{z}emovi\'c · 14 de septiembre de 2026
Multi-object tracking (MOT) is dominated by the tracking-by-detection paradigm, whose methods typically rely on a small set of hyperparameters that are conventionally chosen by hand. Tuning them requires repeated expert-guided experimentation, while the procedures used to select reported values are …
- Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification
Menglin Wang, Xiaojin Gong · 11 de septiembre de 2026
Estimating reliable cross-modality association is crucial to unsupervised visible-infrared person re-ID. While optimal transport is shown to be a practical solution for cross-modality association, it suffers from the rigidness of hard label assignment without considering the impact of cluster noise.…
- TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking
Zhaofeng Hu, Sifan Zhou, Jiahao Nie, Ziyu Zhao, Weizi Li, Ci-jyun Liang · 9 de septiembre de 2026
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames in sparse point clouds. Existing methods, rooted in the Siamese tracking paradigm from 2D vision, rely on costly dual-input designs and excessive motion…
- Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents
Daniel Davila, Ravikumar Balakrishnan, Mike Cochran · 7 de septiembre de 2026
We present a system that uses a Vision-Language Model (VLM) as a diagnostic agent for adapting a detect-to-track pipeline to a new target domain without access to target-domain labels. Rather than optimizing against annotated metrics, the VLM directly inspects rendered tracking outputs, identifies v…
- PuTR-CouT: Counting-by-Tracking in Camera-Trap Image Sequences
Fagner Cunha, Juan G. Colonna, Eulanda M. dos Santos · 7 de septiembre de 2026
Species identification in camera trap images has been widely studied, but key ecological modeling tasks such as species abundance or density estimation also require counting individual animals. However, the lack of counting labels in most datasets and low frame rates (typically ~1 frame per second) …
- Counting Beyond Instances: A Benchmark for Group-Individual Object Counting
Rui Wang, Junyi Huang, Jiahui Li, Qiao Yu, Yixue Hao, Long Hu, Baoru Huang · 7 de septiembre de 2026
Visual counting is commonly formulated at the instance level, aiming to estimate how many objects of a queried category appear in an image. However, real-world counting often involves higher-level semantic units formed by multiple instances, such as a bunch of grapes, a stack of plates, or a pair of…
- Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision
Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali, Arno Solin · 4 de septiembre de 2026
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view o…
- ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation
Javier del Pino (SperidLabs), Salvador Rodr\'iguez (SperidLabs), Alejandro Garabito (SperidLabs), Javier \'Alvarez (SperidLabs), Chema Garabito (SperidLabs) · 4 de septiembre de 2026
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to …
Otros asuntos del tema Visión por computador y reconocimiento de formas
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Multimodal Machine Learning Applications8069 artículos / 12 meses+191 %
- Generative Adversarial Networks and Image Synthesis4992 artículos / 12 meses+39 %
- Advanced Neural Network Applications2354 artículos / 12 meses+48 %
- Advanced Vision and Imaging841 artículos / 12 meses+78 %
- Human Pose and Action Recognition836 artículos / 12 meses+457 %
- Face recognition and analysis482 artículos / 12 meses+88 %
