DeepL Voice: Real-Time Speech-to-Speech Translation
DeepL Voice: Real-Time Speech-to-Speech Translation
Johannes Ernesti, Peter Kaiser, Jonas Heinze, Elnaz Shafaei-Bajestan, Kristina Geißler, Weiyue Wang, Johannes Beck, Sascha Brinker, Thorben Finke
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Demo Track. Pages 8385-8388.
https://doi.org/10.24963/ijcai.2026/962
DeepL Voice is a real-time speech-to-speech translation system for global business communication, following a pragmatic incremental approach: developing a production-grade cascaded speech-to-speech-translation (S2ST) system, while exploring end-to-end solutions in parallel.
The production system (launched November 2024) achieves competitive transcription quality through proprietary real-time ASR models and eliminates translation "flickering" via stable text streaming while maintaining low latency.
Supporting 18 input languages and 30+ target languages, it offers DeepL Voice for Meetings (Microsoft Teams/Zoom integration), DeepL Voice for Conversations (mobile apps), as well as the DeepL API for Voice.
Key features include customizable formality and glossary support for business-appropriate communication, with voice cloning TTS under development.
Keywords:
AI: Humans and AI
AI: Natural Language Processing
