A Systematic Review on Audio-Visual Speech-To-Speech Translation Models

Authors

DOI:

https://doi.org/10.5753/reviews.2026.6237

Keywords:

Audio-Visual Speech-to-Speech Translation, Systematic Literature Review, Self-Supervised Learning, AV-HuBERT, Discrete Units, Multimodal Machine Translation

Abstract

This systematic literature review summarizes the evolution of Audio-Visual Speech-to-Speech Translation (AV-S2ST), focusing on the shift from cascaded systems to direct end-to-end models. Key advancements utilize self-supervised learning (SSL) (e.g. AV-HuBERT) and discrete units (e.g. TransFace) to overcome data scarcity and enable textless multilingual translation. Current research tackles challenges in noise robustness, lip-synchrony, isometric translation, and speaker preservation. Techniques like cross-modal distillation show promise, driving progress towards robust, real-time AV-S2ST with enhanced zero-shot capabilities.

Downloads

Download data is not yet available.

References

Afouras, T., Chung, J. S., and Zisserman, A. (2018a). Deep audio-visual speech recognition. In Proceedings of the European conference on computer vision (ECCV), pages 653-668. DOI: 10.1109/tpami.2018.2889052.

Afouras, T., Chung, J. S., and Zisserman, A. (2018b). LRS3-TED: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496. DOI: 10.48550/arXiv.1809.00496.

Anwar, M., Han, H., Pino, J., Carpuat, M., Shi, B., and Wang, C. (2023). MuAViC: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation. arXiv preprint arXiv:2305.15814. DOI: 10.48550/arXiv.2303.00628.

Cheng, X., Huang, R., Li, L., Wang, Z., Jin, T., Yin, A., Feiyang, C., Duan, X., Huai, B., and Zhao, Z. (2024). TransFace: Unit-based audio-visual speech synthesizer for talking head translation. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 9973-9986, Bangkok, Thailand. Association for Computational Linguistics. DOI: 10.18653/v1/2024.findings-acl.593.

Cheng, X., Li, L., Jin, T., Huang, R., Lin, W., Wang, Z., Liu, H., Wang, Y., Yin, A., and Zhao, Z. (2023). Mixspeech: Cross-modality self-learning with audio-visual stream mixup for visual speech translation and recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15689-15699. DOI: 10.1109/iccv51070.2023.01442.

Choi, J., Park, S.-J., Kim, M., and Ro, Y. M. (2024). AV2AV: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). DOI: 10.48550/arXiv.2312.02512.

Chung, J. S., Nagrani, A., and Zisserman, A. (2018). Voxceleb2: Deep speaker recognition. In Interspeech 2018, pages 1086-1090. DOI: 10.21437/interspeech.2018-1929.

Di Gangi, M. A., Negri, M., and Turchi, M. (2019). One-to-many multilingual end-to-end speech translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 585-592. IEEE. DOI: 10.1109/asru46091.2019.9004003.

Dong, Q., Yue, F., Ko, T., Wang, M., Bai, Q., and Zhang, Y. (2022). Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation. In Interspeech 2022, pages 1781-1785. DOI: 10.21437/Interspeech.2022-10011.

Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., and Rubinstein, M. (2018). Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. In ACM SIGGRAPH 2018 Papers, pages 1-11. DOI: 10.1145/3197517.3201357.

Felizardo, K. R., Mendes, E., Fagan, M., and Nakagawa, E. Y. (2012). Systematic literature review using parsifal. In Proceedings of the 8th international conference on the quality of information and communications technology, pages 363-368. Book.

Goncalves, L., Mathur, P., Niu, X., Lavania, C., Houston, B., Vishnubhotla, S., Sun, L., and Ferritto, A. (2025). Improving lip-synchrony in direct audio-visual speech-to-speech translation. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1-5. DOI: 10.1109/ICASSP49660.2025.10890684.

Gupta, M., Dutta, M., and Maurya, C. K. (2024). Direct Speech-to-Speech Neural Machine Translation: A Survey. arXiv preprint arXiv:2401.14453. DOI: 10.1016/j.specom.2025.103317.

Han, H., Anwar, M., Pino, J., Hsu, W.-N., Carpuat, M., Shi, B., and Wang, C. (2024). XLAVS-R: Cross-lingual audio-visual speech representation learning for noise-robust speech perception. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12896-12911, Bangkok, Thailand. Association for Computational Linguistics. DOI: 10.18653/v1/2024.acl-long.697.

Huang, R., Liu, H., Cheng, X., Ren, Y., Li, L., Ye, Z., He, J., Zhang, L., Liu, J., Yin, X., and Zhao, Z. (2023). AV-TranSpeech: Audio-visual robust speech-to-speech translation. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8590-8604, Toronto, Canada. Association for Computational Linguistics. DOI: 10.18653/v1/2023.acl-long.479.

Inaguma, H., Kiyono, S., Duh, K., Karita, S., Yalta, N., Hayashi, T., and Watanabe, S. (2019). Multilingual end-to-end speech translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 570-577. IEEE. DOI: 10.1109/asru46091.2019.9003832.

Jia, Y., Zen, H., Rivera, J., Tran, T., Chiu, C.-K., Wu, Y., Wang, C.-I., Chen, N., Pino, J., and Wang, Y. (2022). Cvss: A large-scale multilingual speech-to-speech translation corpus. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 310-317. IEEE. DOI: 10.48550/arXiv.2201.03713.

Kitchenham, B. and Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering. Available at:[link].

Liu, Y., Luan, F., Wang, J., Wang, C., Xiao, T., and Zhu, J. (2019). End-to-end speech translation with knowledge distillation. In Interspeech 2019, pages 2020-2024. DOI: 10.21437/interspeech.2019-2582.

Marshall, C. and Brereton, P. (2015). Systematic literature reviews in software engineering-a tertiary study. Information and Software Technology, 64:1-13. DOI: 10.1016/j.infsof.2010.03.006.

Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311-318. DOI: 10.3115/1073083.1073135.

Post, M. (2018). A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186-191, Belgium, Brussels. Association for Computational Linguistics. Available at:[link].

Prajwal, K., Mukhopadhyay, R., Namboodiri, V. P., and Jawahar, C. (2020). A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia, pages 484-492. DOI: 10.1145/3394171.3413532.

Salesky, E., Lo, J., Miculicich, L., Chen, B., Sokolov, A., Williams, A., Lane, I., Pino, J., Cherry, C., and Neubig, G. (2021). mTEDx: A multilingual transcribed and translated speech corpus of TEDx talks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3711-3725, Online. Association for Computational Linguistics. DOI: 10.18653/v1/2021.naacl-main.295.

Sarim, M., Shakeel, S., Javed, L., Jamaluddin, and Nadeem, M. (2025). Direct Speech to Speech Translation: A Review. arXiv preprint arXiv:2503.04799. DOI: 10.48550/arxiv.2503.04799.

Shintaku, K., de Souza, B., Durelli, V., Durelli, R., and Nakagawa, E. (2016). Assisting systematic literature review with a web-based text mining tool. In 2016 42th Latin American Computing Conference (CLEI), pages 1-10. IEEE. Book.

Snyder, D., Chen, G., and Povey, D. (2015). Musan: A music, speech, and noise corpus. In arXiv preprint arXiv:1510.08484. DOI: 10.48550/arxiv.1510.08484.

Wang, C., Riviere, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., Williamson, M., Pino, J., and Dupoux, E. (2021). VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 955-969, Online. Association for Computational Linguistics. DOI: 10.18653/v1/2021.acl-long.75.

Downloads

Published

2026-08-06

How to Cite

Pereira, A. de G., Ferreira, R. C., & Goldman, A. (2026). A Systematic Review on Audio-Visual Speech-To-Speech Translation Models. SBC Computing Reviews, 5(1), 40–54. https://doi.org/10.5753/reviews.2026.6237

Issue

Section

Articles