A Gloss-driven Indian Sign Language Production System Using Learned Pose Representations

A Gloss-driven Indian Sign Language Production System Using Learned Pose Representations

Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda, Soumitra Samanta

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
AI and Social Good. Pages 7399-7407. https://doi.org/10.24963/ijcai.2026/823

Sign Language Production (SLP) system translates spoken or written language into sign language, enabling accessible communication between the deaf/hard-of-hearing and the hearing population. Being one of the most widely used sign languages globally, Indian Sign Language (ISL) is a very low-resource language and lacks such SLP systems. This paper presents a scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings. The model learns discrete pose representations (codes) by disentangling body, left-hand, and right-hand keypoints, enabling efficient pose modeling and co-articulated sign generation. The proposed system is evaluated using a Hindi movie subtitle corpus coupled with an off-the-shelf back-translation model and achieves a gloss BLEU-4 score of 47.20. The system-generated signs are evaluated by certified ISL interpreters with an average rating of 4.33/5, and a BERT precision of 0.7683 on glosses. In addition, the proposed system achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, attaining a BLEU-4 score of 10.03 and surpassing the previous best method by 0.67 points. The system with code is available at https://cs.rkmvu.ac.in/~isl/sl_gen_vqvae.
Keywords:
Computer Vision: Computer Vision
Humans and AI: Humans and AI
Multidisciplinary Topics and Applications: Multidisciplinary Topics and Applications
Natural Language Processing: Natural Language Processing