Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Human Pose and Action Recognition
836 papers indexed
The study of human pose and action recognition focuses on analyzing body movements and their sequences to extract representations usable by artificial intelligence models. Recent work explores diverse approaches, such as world models capable of simulating future behaviors, transformers adapted to spatial and temporal data, or methods for action segmentation and evaluation from videos. This research relies on architectures combining computer vision, graph processing, and multimodal learning to address scenarios ranging from individual analysis to interactions between multiple agents.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China43% · 243 papers
- United States31% · 174 papers
- United Kingdom8.1% · 46 papers
- Germany7.9% · 45 papers
- South Korea6.7% · 38 papers
- Japan6.2% · 35 papers
- Singapore4.6% · 26 papers
- Canada4.1% · 23 papers
Across 567 papers on this subject with at least one lab located. 57 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video
Rikhat Akizhanov (MBZUAI), Yangsong Zhang (MBZUAI), Nikolai Kaliazin (MBZUAI), Peter Wolf (ETH Z\"urich), Yoshihiko Nakamura (MBZUAI), Pascal Fua (EPFL), Fabio Pizzati (MBZUAI), Ivan Laptev (MBZUAI) · 2 October 2026
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PAC…
- STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
Animesh Varma · 2 October 2026
Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion. W…
- NextMe-800: Anticipating Personal Behavior from Months of Egocentric Video
Zhaoxu Meng, Yiming Sun, Mingyuan Gao, Jiachang Zhang, Zhuhan Dai, Yipeng Du, Zheng Lian, Jian-Qiao Zhu · 2 October 2026
We often plan ambitiously yet act habitually and wonder, in retrospect, whether we would have planned differently had we known what we would actually do. Hindsight offers a valuable perspective on past decisions, although we often wish we could have simulated hindsight at the moment of choosing. If …
- Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan · 1 October 2026
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action…
- Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination
Xinling Xie, Haodong Wang, Jiazhi Mi, Zhiming Liu, Zicong Hong, Xiaoyi Pang, Qianli Liu, Yangjia Hu, Ying Chen, Zhengyang Yan, Song Guo · 1 October 2026
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods …
- Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation
Hongjia Zhai, Xiyu Zhang, Haoran Zhang, Zhichao Ye, Haomin Liu, Guofeng Zhang, Ian Reid, Xingxing Zuo · 1 October 2026
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person obser…
- HEIR: Learning Human-Entity Interactions with Functional Roles
Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang, Minheng Wu, Zhihang Chen, Haiwen Sun, Fei Teng, Zhiyuan Gao, Yufeng Zhang, Yuanhao Luo, Jingqi Zhang, Yufan Chen, Junwei Zheng, Ruiping Liu, Jiale Wei, Kailun Yang, Kunyu Peng · 30 September 2026
Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI…
- Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng · 30 September 2026
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose…
- Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio · 30 September 2026
Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (u…
- De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift
Mengyuan Liu, Yuhang Wen, Yi Zhang, Songtao Wu, Hong Liu, Junsong Yuan, Beichen Ding · 29 September 2026
Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which expects individuals are independent and …
- Auditing Quality Filters for Long-Tail Human Data Curation
Rishav Agarwal, Nirshal Chandra Sekar, Anirudh Vemula · 29 September 2026
Robots on construction sites must detect workers who are kneeling or bending, which we call low poses. These workers can be lost from training datasets during automatic labeling. We study a pipeline that detects people, estimates their body joints using NLF, and groups similar poses. Low poses accou…
- Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang · 28 September 2026
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled…
- Ego-Exo4D Human Meshes Dataset: 4D Human Motion Reconstruction for Ego-Exo Captures
Abhiram Maddukuri, Georgios Pavlakos · 25 September 2026
Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich resource for skill learning and assessment, procedural activity understanding, and embodied AI. However, the dataset ships with only sparse 3D human pose annotations, and reconstructing dense…
- Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models
Abu Bakar, Abdullah Aftab, Amir Hamza · 25 September 2026
Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset fr…
- DeltaWAM: Delta World Action Models for Bimanual Manipulation
Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang · 25 September 2026
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-condi…
- WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
Jerrin Bright, John Zelek · 25 September 2026
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can suppo…
- PAOLI: Pose-free Articulated Object Learning from Sparse-view Images
Jianning Deng, Kartic Subr, Hakan Bilen · 22 September 2026
We present a method for modeling articulated objects from sparse images with unknown camera poses. Existing approaches require dense multi-view observations and ground-truth camera poses shared across articulation states, whereas our approach operates with as few as four unposed views per articulati…
- AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos
Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad "Mahi'' Shafiullah · 22 September 2026
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry…
- STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation
Mena Kamel, Natalie Won, Amrut Sarangi, Sven Jager, Albert Pla Planas · 22 September 2026
Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based archite…
- Estimating Accurate Hand Pose in Camera Space with Vision Transformer
Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia · 22 September 2026
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. H…
- TReViS: Temporal Repetition Structure Aware Video Synthesis for Self-supervised Repetitive Action Counting
Fanqi Yu, Shengming Ma, Stefano Fiorini, Vito Paolo Pastore, Xuan Qi, Vittorio Murino, Cigdem Beyan · 22 September 2026
Fully supervised repetitive action counting (RAC) has achieved strong performance, but requires dense temporal annotations that are costly and difficult to scale. We propose TReViS, a self-supervised video synthesis framework that enables training RAC models without any repetition labels. TReViS est…
- Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding
Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen · 22 September 2026
Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur…
- MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space
Markus Karmann, Shile Li, Christian Intern\`o, Bruno Andreis, David Klindt, Randall Balestriero, Jindong Gu, Philip Torr, Qi Zhang, Peng-Tao Jiang, Hao Zhang, Bo Li, Onay Urfalioglu · 22 September 2026
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent repre…
- Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer
Yuanzhe Li, Hounian Liu, Xiaotong Chang, Yidi Huang · 22 September 2026
The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for …
- Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
Di Wen, Ruodi Zhang, Kailun Yang, Kunyu Peng · 22 September 2026
Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captur…
Other topics in Computer vision and pattern recognition
The topics the OpenAlex classification attaches to the same theme, most active first.
- Multimodal Machine Learning Applications8,069 papers / 12 months+191%
- Generative Adversarial Networks and Image Synthesis4,992 papers / 12 months+39%
- Advanced Neural Network Applications2,354 papers / 12 months+48%
- Advanced Vision and Imaging841 papers / 12 months+78%
- Face recognition and analysis482 papers / 12 months+88%
- Image Enhancement Techniques463 papers / 12 months+20%
