Physical Sciences › Computer Science › Human-Computer Interaction
Hand Gesture Recognition Systems
156 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States23% · 23 papers
- China20% · 20 papers
- United Kingdom13% · 13 papers
- South Korea7.9% · 8 papers
- India6.9% · 7 papers
- Türkiye6.9% · 7 papers
- Australia5.9% · 6 papers
- Italy5.9% · 6 papers
Across 101 papers on this subject with at least one lab located. 36 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval
Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden · 4 September 2026
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for cl…
- A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval
Santiago Poveda-Guti\'errez, Hideki Nakayama, Mayumi Bono · 4 September 2026
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning …
- Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks
Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden · 4 September 2026
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning …
- SignRR: Retrieve and Refine Real Motion for Sign Language Production
Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan, Gissella Bejarano · 31 August 2026
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, …
- SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting
Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin, Younggeun Choi · 27 August 2026
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learn…
- Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition
Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin · 26 August 2026
Hand gesture-based Sign Language Recognition (SLR) serves as a crucial communication bridge between deaf and non-deaf individuals. While Graph Convolutional Networks (GCNs) are common, they are limited by their reliance on fixed skeletal graphs. To overcome this, we propose the Sequential Spatio-Tem…
- JSL-DC: A Word-Level Japanese Sign Language Dataset with Linguist-Derived Descriptions for Distinguishing Confusable Signs
Ken Takaki, Asuka Ando, Misa Suzuki, Uiko Yano, Masaya Tsujimoto, Bill Neubauer, Ananay Vikram Gupta, Rose Shao, Matthias Hoppe, Sahir Shahryar, Celeste Mason, Kai Kunze, Yohei Oseki, Yoshihiro Kawahara, Thad Starner · 20 August 2026
Effective sign language (SL) acquisition is crucial for deaf children, yet 95% are born to hearing parents who often lack proficiency in SL. SL recognition can power learning tools to help parents communicate with their children. However, Japanese Sign Language (JSL) lacks large-scale, multi-signer …
- SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation
Taegang Kim, Saleh Afroogh, Junfeng Jiao · 18 August 2026
Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained hand gestures in safety-critical operational contexts remains largely unexamined. We introduce SafeGesture, a benchmark that evaluates whether a model can i…
- Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
Keren Artiaga (Victor), Yang Li (Victor), Ercan Engin Kuruoglu (Victor), Wai Kin (Victor), Chan · 18 August 2026
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100 distinct sign languages are severely lacking. In response, we present our work on sign language recognition using transfer learning and the domain adaptation me…
- Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo · 14 August 2026
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators -- global, hand, and head -- each guide a…
- UniCon-Former: Unified Convolution Transformer is All You Need for Hand Gesture Recognition
Mallika Garg, Debashis Ghosh, Pyari Mohan Pradhan · 14 August 2026
Convolutional Neural Networks (CNNs) capture local features efficiently but struggle with global context due to their limited receptive field. On the other hand, transformers effectively capture global dependencies through self-attention but suffer from high redundancy and computational costs. Thus,…
- ODE-Based Transformer Decoders for Iterative Sign Language Translation
Tu\u{g}\c{c}e K{\i}z{\i}ltepe, Hacer Yalim Keles · 13 August 2026
Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of increased computation. We propose a parameter-efficient alternative that improves expressiveness without increasing model size. Rather t…
- The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko, Youngwoo Yoon, Gustav Eje Henter · 12 August 2026
This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluat…
- A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta · 12 August 2026
Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg …
- Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
Rustem Ozakar, Eyup Gedikli · 11 August 2026
Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Point clouds obtained from depth images can be used for sign language recognition with neural networks like PointNet. In recent years, various neural ne…
- Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie · 11 August 2026
Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leadi…
- SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
Shiwei Gan, Xiao Liu, Yafeng Yin, Zhiwei Jiang, Bowen Guo, Lie Xie, Sanglu Lu, Hongkai Wen · 11 August 2026
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key is…
- TransSLR: A Lightweight Transformer for Sign Language Recognition
Lucia Yen Wanchi, Samuel Johnny, Victor Tolulope Olufemi, Emmanuel Aaron, Moise Busogi · 10 August 2026
Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we show that the common heuristic of fine-tuning hig…
- Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Saad Ahmed, Md Khalid Syfullaha · 7 August 2026
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets without expert verification and heavyweight pretrained b…
- A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition
Nitin Kumar Singh, Arie Rachmad Syulistyo, Yuichiro Tanaka, Hakaru Tamukoh · 5 August 2026
Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has achieved promising performance in SLR, its high computational cost limits deployment on edge devices. To address this challenge, we propose a lightweight reservoir…
- Attention-Steered Vision-Language Models for Sign Language Translation
Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao · 4 August 2026
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased translators: poor spatial-temporal visual grounding. In particular, we …
- Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Shiwei Gan, Lichen Wang, Xiao Liu, Yafeng Yin, Kuizhuang Liu, Sanglu Lu, Lei Xie · 31 July 2026
Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-langu…
- DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
Hongbin Zhang, Junhao Liu, Xuefeng Bai, Youcheng Pan, Yang Xiang, Kehai Chen · 31 July 2026
Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT metho…
- An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Zihan Guo, Xiaoqi Li · 24 July 2026
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based ge…
- From Sign Language Generation to Humanoid Execution: Vision-Language Guided Retargeting with Collision Mitigation
Nabeela Khan, Bowen Wu, Runwu Shi, Benjamin Yen, Takeshi Ashizawa, Carlos Toshinori Ishi, Takashi Minato, Kazuhiro Nakadai · 21 July 2026
Recent sign language generation (SLG) systems increasingly output dense 3D body representations, which better preserve full-body kinematics and geometry for downstream embodiment on humanoid robots. However, these generated motions frequently exhibit self-intersections such as hand-hand and hand-tor…
