Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Human Pose and Action Recognition
836 indexierte Paper
Die Untersuchung der Erkennung von menschlichen Posen und Aktionen zielt darauf ab, Körperbewegungen und deren Sequenzen zu analysieren, um Darstellungen zu extrahieren, die von Modellen der künstlichen Intelligenz genutzt werden können. Aktuelle Arbeiten erforschen verschiedene Ansätze, wie world models, die zukünftige Verhaltensweisen simulieren können, transformers, die an räumliche und zeitliche Daten angepasst sind, oder Methoden zur Segmentierung und Bewertung von Aktionen aus Videos. Diese Forschungen stützen sich auf Architekturen, die Computer Vision, Graphenverarbeitung und multimodales Lernen kombinieren, um Szenarien von der Einzelanalyse bis zur Interaktion zwischen mehreren agents zu bearbeiten.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China43 % · 243 Artikel
- Vereinigte Staaten31 % · 174 Artikel
- Vereinigtes Königreich8,1 % · 46 Artikel
- Deutschland7,9 % · 45 Artikel
- Südkorea6,7 % · 38 Artikel
- Japan6,2 % · 35 Artikel
- Singapur4,6 % · 26 Artikel
- Kanada4,1 % · 23 Artikel
Über 567 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 57 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video
Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI) · 2. Oktober 2026
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PAC…
- STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
Animesh Varma · 2. Oktober 2026
Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. W…
- NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video
Zhaoxu Meng, Yiming Sun, Mingyuan Gao, Jiachang Zhang, Zhuhan Dai, Yipeng Du, Zheng Lian, Jian-Qiao Zhu · 2. Oktober 2026
We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If …
- Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan · 1. Oktober 2026
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action…
- Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination
Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo · 1. Oktober 2026
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods …
- Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation
Hongjia Zhai, Xiyu Zhang, Haoran Zhang, Zhichao Ye, Haomin Liu, Guofeng Zhang, Ian Reid, Xingxing Zuo · 1. Oktober 2026
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person obser…
- HEIR: Learning Human-Entity Interactions with Functional Roles
Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang, Minheng Wu, Zhihang Chen, Haiwen Sun, Fei Teng, Zhiyuan Gao, Yufeng Zhang, Yuanhao Luo, Jingqi Zhang, Yufan Chen, Junwei Zheng, Ruiping Liu, Jiale Wei, Kailun Yang, Kunyu Peng · 30. September 2026
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI…
- Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng · 30. September 2026
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose…
- Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio · 30. September 2026
Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (u…
- De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift
Mengyuan Liu, Yuhang Wen, Yi Zhang, Songtao Wu, Hong Liu, Junsong Yuan, Beichen Ding · 29. September 2026
Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which expects individuals are independent and …
- Auditing Quality Filters for Long-Tail Human Data Curation
Rishav Agarwal, Nirshal Chandra Sekar, Anirudh Vemula · 29. September 2026
Robots on construction sites must detect workers who are kneeling or bending, which we call low poses. These workers can be lost from training datasets during automatic labeling. We study a pipeline that detects people, estimates their body joints using NLF, and groups similar poses. Low poses accou…
- Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang · 28. September 2026
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled…
- Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures
Abhiram Maddukuri, Georgios Pavlakos · 25. September 2026
Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich resource for skill learning and assessment, procedural activity understanding, and embodied AI. However, the dataset ships with only sparse 3D human pose annotations, and reconstructing dense…
- Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models
Abu Bakar, Abdullah Aftab, Amir Hamza · 25. September 2026
Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset fr…
- DeltaWAM: Delta World Action Models for Bimanual Manipulation
Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang · 25. September 2026
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-condi…
- WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
Jerrin Bright, John Zelek · 25. September 2026
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can suppo…
- PAOLI: Pose-free Articulated Object Learning from Sparse-view Images
Jianning Deng, Kartic Subr, Hakan Bilen · 22. September 2026
We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view observations and ground-truth camera poses shared across articulation states, whereas our approach operates with as few as four unposed views per articulati…
- AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos
Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad "Mahi'' Shafiullah · 22. September 2026
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry…
- STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation
Mena Kamel, Natalie Won, Amrut Sarangi, Sven Jager, Albert Pla Planas · 22. September 2026
Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based archite…
- Estimating Accurate Hand Pose in Camera Space with Vision Transformer
Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia · 22. September 2026
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. H…
- TReViS: Temporal Repetition Structure Aware Video Synthesis for Self-supervised Repetitive Action Counting
Fanqi Yu, Shengming Ma, Stefano Fiorini, Vito Paolo Pastore, Xuan Qi, Vittorio Murino, Cigdem Beyan · 22. September 2026
Fully supervised repetitive action counting (RAC) has achieved strong performance, but requires dense temporal annotations that are costly and difficult to scale. We propose TReViS, a self-supervised video synthesis framework that enables training RAC models without any repetition labels. TReViS est…
- Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding
Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen · 22. September 2026
Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur…
- MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space
Markus Karmann, Shile Li, Christian Intern\`o, Bruno Andreis, David Klindt, Randall Balestriero, Jindong Gu, Philip Torr, Qi Zhang, Peng-Tao Jiang, Hao Zhang, Bo Li, Onay Urfalioglu · 22. September 2026
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent repre…
- Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer
Yuanzhe Li, Hounian Liu, Xiaotong Chang, Yidi Huang · 22. September 2026
The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for …
- Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
Di Wen, Ruodi Zhang, Kailun Yang, Kunyu Peng · 22. September 2026
Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captur…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
- Image Enhancement Techniques463 Papiere / 12 Monate+20 %
