A Systematic Review on Audio-Visual Speech-To-Speech Translation Models
DOI:
https://doi.org/10.5753/reviews.2026.6237Keywords:
Audio-Visual Speech-to-Speech Translation, Systematic Literature Review, Self-Supervised Learning, AV-HuBERT, Discrete Units, Multimodal Machine TranslationAbstract
This systematic literature review summarizes the evolution of Audio-Visual Speech-to-Speech Translation (AV-S2ST), focusing on the shift from cascaded systems to direct end-to-end models. Key advancements utilize self-supervised learning (SSL) (e.g. AV-HuBERT) and discrete units (e.g. TransFace) to overcome data scarcity and enable textless multilingual translation. Current research tackles challenges in noise robustness, lip-synchrony, isometric translation, and speaker preservation. Techniques like cross-modal distillation show promise, driving progress towards robust, real-time AV-S2ST with enhanced zero-shot capabilities.
Downloads
References
Afouras, T., Chung, J. S., and Zisserman, A. (2018a). Deep audio-visual speech recognition. In Proceedings of the European conference on computer vision (ECCV), pages 653-668. DOI: 10.1109/tpami.2018.2889052.
Afouras, T., Chung, J. S., and Zisserman, A. (2018b). LRS3-TED: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496. DOI: 10.48550/arXiv.1809.00496.
Anwar, M., Han, H., Pino, J., Carpuat, M., Shi, B., and Wang, C. (2023). MuAViC: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation. arXiv preprint arXiv:2305.15814. DOI: 10.48550/arXiv.2303.00628.
Cheng, X., Huang, R., Li, L., Wang, Z., Jin, T., Yin, A., Feiyang, C., Duan, X., Huai, B., and Zhao, Z. (2024). TransFace: Unit-based audio-visual speech synthesizer for talking head translation. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 9973-9986, Bangkok, Thailand. Association for Computational Linguistics. DOI: 10.18653/v1/2024.findings-acl.593.
Cheng, X., Li, L., Jin, T., Huang, R., Lin, W., Wang, Z., Liu, H., Wang, Y., Yin, A., and Zhao, Z. (2023). Mixspeech: Cross-modality self-learning with audio-visual stream mixup for visual speech translation and recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15689-15699. DOI: 10.1109/iccv51070.2023.01442.
Choi, J., Park, S.-J., Kim, M., and Ro, Y. M. (2024). AV2AV: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). DOI: 10.48550/arXiv.2312.02512.
Chung, J. S., Nagrani, A., and Zisserman, A. (2018). Voxceleb2: Deep speaker recognition. In Interspeech 2018, pages 1086-1090. DOI: 10.21437/interspeech.2018-1929.
Di Gangi, M. A., Negri, M., and Turchi, M. (2019). One-to-many multilingual end-to-end speech translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 585-592. IEEE. DOI: 10.1109/asru46091.2019.9004003.
Dong, Q., Yue, F., Ko, T., Wang, M., Bai, Q., and Zhang, Y. (2022). Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation. In Interspeech 2022, pages 1781-1785. DOI: 10.21437/Interspeech.2022-10011.
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M. (2018). Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. In ACM SIGGRAPH 2018 Papers, pages 1-11. DOI: 10.1145/3197517.3201357.
Felizardo, K. R., Mendes, E., Fagan, M., and Nakagawa, E. Y. (2012). Systematic literature review using parsifal. In Proceedings of the 8th international conference on the quality of information and communications technology, pages 363-368. Book.
Goncalves, L., Mathur, P., Niu, X., Lavania, C., Houston, B., Vishnubhotla, S., Sun, L., and Ferritto, A. (2025). Improving lip-synchrony in direct audio-visual speech-to-speech translation. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1-5. DOI: 10.1109/ICASSP49660.2025.10890684.
Gupta, M., Dutta, M., and Maurya, C. K. (2024). Direct Speech-to-Speech Neural Machine Translation: A Survey. arXiv preprint arXiv:2401.14453. DOI: 10.1016/j.specom.2025.103317.
Han, H., Anwar, M., Pino, J., Hsu, W.-N., Carpuat, M., Shi, B., and Wang, C. (2024). XLAVS-R: Cross-lingual audio-visual speech representation learning for noise-robust speech perception. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12896-12911, Bangkok, Thailand. Association for Computational Linguistics. DOI: 10.18653/v1/2024.acl-long.697.
Huang, R., Liu, H., Cheng, X., Ren, Y., Li, L., Ye, Z., He, J., Zhang, L., Liu, J., Yin, X., and Zhao, Z. (2023). AV-TranSpeech: Audio-visual robust speech-to-speech translation. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8590-8604, Toronto, Canada. Association for Computational Linguistics. DOI: 10.18653/v1/2023.acl-long.479.
Inaguma, H., Kiyono, S., Duh, K., Karita, S., Yalta, N., Hayashi, T., and Watanabe, S. (2019). Multilingual end-to-end speech translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 570-577. IEEE. DOI: 10.1109/asru46091.2019.9003832.
Jia, Y., Zen, H., Rivera, J., Tran, T., Chiu, C.-K., Wu, Y., Wang, C.-I., Chen, N., Pino, J., and Wang, Y. (2022). Cvss: A large-scale multilingual speech-to-speech translation corpus. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 310-317. IEEE. DOI: 10.48550/arXiv.2201.03713.
Kitchenham, B. and Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering. Available at:[link].
Liu, Y., Luan, F., Wang, J., Wang, C., Xiao, T., and Zhu, J. (2019). End-to-end speech translation with knowledge distillation. In Interspeech 2019, pages 2020-2024. DOI: 10.21437/interspeech.2019-2582.
Marshall, C. and Brereton, P. (2015). Systematic literature reviews in software engineering-a tertiary study. Information and Software Technology, 64:1-13. DOI: 10.1016/j.infsof.2010.03.006.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311-318. DOI: 10.3115/1073083.1073135.
Post, M. (2018). A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186-191, Belgium, Brussels. Association for Computational Linguistics. Available at:[link].
Prajwal, K., Mukhopadhyay, R., Namboodiri, V. P., and Jawahar, C. (2020). A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, pages 484-492. DOI: 10.1145/3394171.3413532.
Salesky, E., Lo, J., Miculicich, L., Chen, B., Sokolov, A., Williams, A., Lane, I., Pino, J., Cherry, C., and Neubig, G. (2021). mTEDx: A multilingual transcribed and translated speech corpus of TEDx talks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3711-3725, Online. Association for Computational Linguistics. DOI: 10.18653/v1/2021.naacl-main.295.
Sarim, M., Shakeel, S., Javed, L., Jamaluddin, and Nadeem, M. (2025). Direct Speech to Speech Translation: A Review. arXiv preprint arXiv:2503.04799. DOI: 10.48550/arxiv.2503.04799.
Shintaku, K., de Souza, B., Durelli, V., Durelli, R., and Nakagawa, E. (2016). Assisting systematic literature review with a web-based text mining tool. In 2016 42th Latin American Computing Conference (CLEI), pages 1-10. IEEE. Book.
Snyder, D., Chen, G., and Povey, D. (2015). Musan: A music, speech, and noise corpus. In arXiv preprint arXiv:1510.08484. DOI: 10.48550/arxiv.1510.08484.
Wang, C., Riviere, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., Williamson, M., Pino, J., and Dupoux, E. (2021). VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 955-969, Online. Association for Computational Linguistics. DOI: 10.18653/v1/2021.acl-long.75.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Alexandre de Godoy Pereira, Renato Cordeiro Ferreira, Alfredo Goldman

This work is licensed under a Creative Commons Attribution 4.0 International License.
