SPL4SH: Designing a Systematic Pipeline for 4D Synthetic Humans in XR Content Creation

Authors

DOI:

https://doi.org/10.5753/jbcs.2026.8229

Keywords:

Synthetic 4D humans, XR content creation, Generative AI, Modular pipelines, Digital humans, Spatial intelligence

Abstract

XR content creation increasingly combines reconstruction, generative modeling, animation synthesis, neural rendering, and real-time engine deployment. However, creating deployable synthetic 4D humans remains fragmented because human assets must preserve body structure, appearance, motion, deformation, and scene-level plausibility across time. This paper proposes SPL4SH, a Systematic Pipeline for 4D Synthetic Humans, to investigate how contemporary AI-assisted techniques can be organized into an XR-oriented production workflow. SPL4SH is built from a stage-based taxonomy covering modeling, rigging, animation, rendering, and interaction. This taxonomy supports a comparative suitability assessment of recent techniques and guides a modular proof-of-concept implementation integrating SMPLify for parametric body modeling and rigging, Kimodo for controllable motion generation, SMPLitex for SMPL-compatible texture generation, Rokoko-based retargeting, Blender-based asset integration, and Unreal Engine deployment. The evaluation combines individual module tests, full-pipeline integration, and scene-level deployment in a meeting-room environment involving human-object, human-human, and human-scene arrangements. Results show that modular 4D human generation is technically feasible: SMPL-X geometry supports rigged deformation, generated motions can be retargeted to animated characters, texture maps improve mannequin-like bodies, and final assets can be placed in real-time XR scenes. However, persistent barriers remain, including format mismatch, manual skeleton alignment, retargeting fragility, texture discontinuities, hallucinated visual regions, motion interpenetration, stiff transitions, and scene-level validation. SPL4SH contributes a grounded framework for selecting, combining, and evaluating independent generative components as XR-ready synthetic human assets.

Downloads

Download data is not yet available.

References

Baradel, F., Armando, M., Galaaoui, S., Brégier, R., Weinzaepfel, P., Rogez, G., and Lucas, T. (2024). Multi-hmr: Multi-person whole-body human mesh recovery in a single shot. DOI: 10.48550/arXiv.2402.14654.

Bekor, Y., Harari, G. M., Perel, O., and Litany, O. (2025). Gaussian see, gaussian do: Semantic 3d motion transfer from multiview video. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers, SA Conference Papers '25, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3757377.3763992.

Bian, S., Xu, C., Xiu, Y., Grigorev, A., Liu, Z., Lu, C., Black, M. J., and Feng, Y. (2025). Chatgarment: Garment estimation, generation and editing via large language models. DOI: 10.48550/arXiv.2412.17811.

Cao, D., Sun, G., Habermann, M., and Bernard, F. (2025). Hyper diffusion avatars: Dynamic human avatar generation using network weight space diffusion. arXiv preprint. DOI: 10.48550/arXiv.2509.04145.

Cao, Y., Cao, Y.-P., Han, K., Shan, Y., and Wong, K.-Y. K. (2024). Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models. pages 958-968. DOI: 10.1109/CVPR52733.2024.00097.

Casas, D. and Comino-Trinidad, M. (2023). Smplitex: A generative model and dataset for 3d human texture estimation from single image. In British Machine Vision Conference. DOI: 10.48550/arXiv.2309.01855.

Cseke, A., Tripathi, S., Dwivedi, S. K., Lakshmipathy, A., Chatterjee, A., Black, M. J., and Tzionas, D. (2025). Pico: Reconstructing 3d people in contact with objects. DOI: 10.48550/arXiv.2504.17695.

Dai, P., Tan, F., Yu, X., Peng, Y., Wang, R., Wang, J., and Liu, Y. (2025). Go-nerf: Generating objects in neural radiance fields for virtual reality content creation. IEEE Transactions on Visualization and Computer Graphics, 31(5):3087-3097. DOI: 10.1109/TVCG.2025.3549558.

de Andrade Araujo, V. F., Costa, A. B., and Musse, S. R. (2023). Evaluating the uncanny valley effect in dark colored skin virtual humans. In 2023 36th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), pages 1-6. IEEE. DOI: 10.1109/SIBGRAPI59091.2023.10347145.

Fiche, G., Leglaive, S., Alameda-Pineda, X., and Moreno-Noguer, F. (2025). Mega: Masked generative autoencoder for human mesh recovery. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5366-5378. DOI: 10.1109/CVPR52734.2025.00505.

Gong, C., Dai, Y., Li, R., Bao, A., Li, J., Yang, J., Zhang, Y., and Li, X. (2024). Text2avatar: Text to 3d human avatar generation with codebook-driven body controllable attribute. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 16-20. DOI: 10.1109/ICASSP48485.2024.10446237.

Gonzalez Morin, D., Gonzalez-Sosa, E., Perez, P., and Villegas, A. (2023). Full body video-based self-avatars for mixed reality: from e2e system to user study. Virtual Reality, 27(3):2129-2147. DOI: 10.1007/s10055-023-00785-0.

Guo, C., Li, J., Kant, Y., Sheikh, Y., Saito, S., and Cao, C. (2025). Vid2avatar-pro: Authentic avatar from videos in the wild via universal prior. DOI: 10.48550/arXiv.2503.01610.

Habermann, M., Liu, L., Xu, W., Pons-Moll, G., Zollhoefer, M., and Theobalt, C. (2023). Hdhumans: A hybrid approach for high-fidelity digital humans. Proc. ACM Comput. Graph. Interact. Tech., 6(3). DOI: 10.1145/3606927.

Heagerty, J., Li, S., Lee, E., Bhattacharyya, S., Bista, S., Brawn, B., Feng, B. Y., Jabbireddy, S., JaJa, J., Kacorri, H., Li, D., Yarnell, D., Zwicker, M., and Varshney, A. (2024). Holocamera: Advanced volumetric capture for cinematic-quality vr applications. IEEE Transactions on Visualization and Computer Graphics, 30(5):2767-2775. DOI: 10.1109/TVCG.2024.3372123.

Hong, Y., Zhang, K., Gu, J., Bi, S., Zhou, Y., Liu, D., Liu, F., Sunkavalli, K., Bui, T., and Tan, H. (2024). Lrm: Large reconstruction model for single image to 3d. arXiv preprint. DOI: 10.48550/arXiv.2311.04400.

Hu, E., Li, M., Qian, X., Olwal, A., Kim, D., Heo, S., and Du, R. (2024a). Experiencing thing2reality: Transforming 2d content into conditioned multiviews and 3d gaussian objects for xr communication. In Adjunct Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, pages 1-3. ACM. DOI: 10.1145/3672539.3686740.

Hu, H., Fan, Z., Wu, T., Xi, Y., Lee, S., Pavlakos, G., and Wang, Z. (2024b). Expressive gaussian human avatars from monocular rgb video. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C., editors, Advances in Neural Information Processing Systems, volume 37, pages 5646-5660. Curran Associates, Inc.. DOI: 10.52202/079017-0183.

Hu, S., Hong, F., Hu, T., Pan, L., Mei, H., Xiao, W., Yang, L., and Liu, Z. (2025). Humanliff: Layer-wise 3d human diffusion model. International Journal of Computer Vision, 133(9):5938-5957. DOI: 10.1007/s11263-025-02477-5.

Kerbl, B., Kopanas, G., Leimkühler, T., and Drettakis, G. (2023). 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4). DOI: 10.1145/3592433.

Knob, P., Pinho, G., Silva, G. F., Montanha, R., Peres, V., Araujo, V., and Musse, S. R. (2024). Surveying the evolution of virtual humans expressiveness toward real humans. Computers & Graphics, 123:104034. DOI: 10.1016/j.cag.2024.104034.

Kolotouros, N., Alldieck, T., Zanfir, A., Bazavan, E. G., Fieraru, M., and Sminchisescu, C. (2023). Dreamhuman: animatable 3d avatars from text. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA. Curran Associates Inc.. DOI: 10.48550/arXiv.2306.09329.

Korban, M. and Li, X. (2022). A survey on applications of digital human avatars toward virtual co-presence. CoRR. DOI: 10.48550/arXiv.2201.04168.

Li, C., Chibane, J., He, Y., Pearl, N., Geiger, A., and Pons-moll, G. (2024). Unimotion: Unifying 3d human motion synthesis and understanding. DOI: 10.48550/arXiv.2409.15904.

Li, C., Zhang, C., Waghwase, A., Lee, L.-H., Rameau, F., Yang, Y., Bae, S.-H., and Hong, C.-S. (2023). Generative ai meets 3d: A survey on text-to-3d in aigc era. ArXiv, abs/2305.06131. DOI: 10.48550/arXiv.2305.06131.

Li, J., Cao, J., Zhang, H., Rempe, D., Kautz, J., Iqbal, U., and Yuan, Y. (2025a). Genmo: A generalist model for human motion. DOI: 10.48550/arXiv.2505.01425.

Li, K., Masuda, M., Schmidt, S., and Mori, S. (2025b). Radiance fields in xr: A survey on how radiance fields are envisioned and addressed for xr research. IEEE Transactions on Visualization and Computer Graphics. DOI: 10.1109/TVCG.2025.3616794.

Li, X., Ma, Q., Lin, T.-Y., Chen, Y., Jiang, C., Liu, M.-Y., and Xiang, D. (2025c). Articulated kinematics distillation from video diffusion models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17571-17581. DOI: 10.1109/CVPR52734.2025.01637.

Li, X., Wang, J., Cheng, Y., Zeng, Y., Ren, X., Zhu, W., Zhao, W., and Yan, Y. (2025d). Towards high-fidelity 3d talking avatar with personalized dynamic texture. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 204-214. DOI: 10.1109/CVPR52734.2025.00028.

Ling, H., Kim, S. W., Torralba, A., Fidler, S., and Kreis, K. (2024). Align your gaussians: Text-to-4d with dynamic 3d gaussians and composed diffusion models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8576-8588. DOI: 10.1109/CVPR52733.2024.00819.

Liu, X., Hou, H., Yang, Y., Li, Y.-L., and Lu, C. (2023). Revisit human-scene interaction via space occupancy. arXiv preprint arXiv:2312.02700. DOI: 10.48550/arXiv.2312.02700.

Liu, Y., Zhu, J., Tang, J., Zhang, S., Zhang, J., Cao, W., Wang, C., Wu, Y., and Huang, D. (2024). Texdreamer: Towards zero-shot high-fidelity 3d human texture generation. DOI: 10.48550/arXiv.2403.12906.

Long, X. et al. (2024). Wonder3d: Single image to 3d using cross-domain diffusion. arXiv preprint. DOI: 10.48550/arXiv.2310.15008.

Menezes, E., Leal, H., Moura, J. V., Araujo, V., and Raupp Musse, S. (2025). Evaluating skin tone biases in virtual human rendering. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Posters, pages 1-2. DOI: 10.1145/3721250.3743031.

Miao, Q., Li, K., Quan, J., Min, Z., Ma, S., Xu, Y., Yang, Y., and Luo, Y. (2025). Advances in 4d generation: A survey. arXiv preprint. DOI: 10.48550/arXiv.2503.14501.

Montanha, R., Araujo, V., Knob, P., Pinho, G., Fonseca, G., Peres, V., and Musse, S. R. (2023). Crafting realistic virtual humans: Unveiling perspectives on human perception, crowds, and embodied conversational agents. In 2023 36th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), pages 252-257. IEEE. DOI: 10.1109/SIBGRAPI59091.2023.10347175.

Nag, S., Cohen-Or, D., Zhang, H., and Amiri, A. M. (2025). In-2-4d: Inbetweening from two single-view images to 4d generation. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers, SA Conference Papers '25, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3757377.3763904.

Pan, Z., Yang, Z., Zhu, X., and Zhang, L. (2025). Efficient4d: Fast dynamic 3d object generation from a single-view video. International Journal of Computer Vision, 134(1):14. DOI: 10.1007/s11263-025-02615-z.

Pang, H. E., Liu, S., Cai, Z., Yang, L., Zhang, T., and Liu, Z. (2025). Disco4d: Disentangled 4d human generation and animation from a single image. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26331-26344. DOI: 10.1109/CVPR52734.2025.02452.

Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A. A., Tzionas, D., and Black, M. J. (2019). Expressive body capture: 3d hands, face, and body from a single image. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10967-10977. DOI: 10.1109/CVPR.2019.01123.

Peng, H.-Y., Zhang, J.-P., Guo, M.-H., Cao, Y.-P., and Hu, S.-M. (2024). Charactergen: Efficient 3d character generation from single images with multi-view pose canonicalization. ACM Trans. Graph., 43(4). DOI: 10.1145/3658217.

Rempe, D., Petrovich, M., Yuan, Y., Zhang, H., Peng, X. B., Jiang, Y., Wang, T., Iqbal, U., Minor, D., de Ruyter, M., Li, J., Tessler, C., Lim, E., Jeong, E., Wu, S., Hassani, E., Huang, M., Yu, J.-B., Chung, C., Song, L., Dionne, O., Kautz, J., Yuen, S., and Fidler, S. (2026). Kimodo: Scaling controllable human motion generation. DOI: 10.48550/arXiv.2603.15546.

Saito, J., Li, J., de Ruyter, M., Guerrero, M., Lim, E., Hassani, E., Ribera, R. B., Moon, H., Dadela, M., Lucca, M. D., Wang, Q., Li, X., Kautz, J., Yuen, S., and Iqbal, U. (2026). Soma: Unifying parametric human body models. DOI: 10.48550/arXiv.2603.16858.

Sakashita, M., Kumaravel, B. T., Marquardt, N., and Wilson, A. D. (2024). Sharednerf: Leveraging photorealistic and view-dependent rendering for real-time and remote collaboration. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM. DOI: 10.1145/3613904.3642945.

Sangeeth Chandran, J. K., Llorens Salvador, M., and Ennis, C. (2026). Virtual humans in virtual reality: a scoping review on sociability, fidelity, and expression. Frontiers in Virtual Reality, Volume 7 - 2026. DOI: 10.3389/frvir.2026.1738000.

Shao, R., Xu, Y., Shen, Y., Yang, C., Zheng, Y., Chen, C., Liu, Y., and Wetzstein, G. (2025). Interspatial attention for efficient 4d human video generation. ACM Transactions on Graphics, 44(4):1-16. DOI: 10.1145/3731165.

Shao, R. et al. (2024). 360-degree human video generation with 4d diffusion transformer. ACM Transactions on Graphics, 43(6). DOI: 10.1145/3687980.

Shaw, R., Jang, Y., Papaioannou, A., Moreau, A., Dhamo, H., Zhang, Z., and Pérez-Pellitero, E. (2026). An interactive conversational 3d virtual human. International Journal of Computer Vision, 134(4):161. DOI: 10.1007/s11263-025-02725-8.

Shen, K., Guo, C., Kaufmann, M., Zarate, J. J., Valentin, J., Song, J., and Hilliges, O. (2023). X-avatar: Expressive human avatars. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16911-16921. DOI: 10.48550/arXiv.2303.04805.

Smolic, A., Amplianitis, K., Moynihan, M., O'Dwyer, N. C., Ondrej, J., Pagés, R., Young, G. W., and Zerman, E. (2022). Volumetric video content creation for immersive xr experiences. In London Imaging Meeting 2022: Display Science, LIM 2022, pages 54-59. Society for Imaging Science and Technology. DOI: 10.2352/LIM.2022.1.1.13.

Sun, G., Chen, X., Chen, Y., Pang, A., Lin, P., Jiang, Y., Xu, L., Yu, J., and Wang, J. (2021). Neural free-viewpoint performance rendering under complex human-object interactions. In Proceedings of the 29th ACM International Conference on Multimedia, MM '21, page 4651–4660, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3474085.3475442.

Sun, G., Dabral, R., Zhu, H., Fua, P., Theobalt, C., and Habermann, M. (2025). Real-time free-view human rendering from sparse-view rgb videos using double unprojected textures. DOI: 10.48550/arXiv.2412.13183.

Taubner, F., Zhang, R., Tuli, M., and Lindell, D. B. (2025). Cap4d: Creating animatable 4d portrait avatars with morphable multi-view diffusion models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5318-5330. DOI: 10.1109/CVPR52734.2025.00501.

Tian, Z., Weng, D., Fang, H., Guo, H., and Bao, Y. (2024). 4d facial capture pipeline incorporating progressive retopology approach. In 2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), pages 741-742. DOI: 10.1109/VRW62533.2024.00167.

Uzun, Y., Riedinger, U., Biswas, S., Hetzer, S., and Oppermann, L. (2025). Enabling digital spaces: Bringing gaussian splats alive in xr. In Extended Reality, Lecture Notes in Computer Science, pages 82-92. Springer. DOI: 10.1007/978-3-031-97769-5_6.

Vilchis, C., Perez-Guerrero, C., Mendez-Ruiz, M., and Gonzalez-Mendoza, M. (2023). A survey on the pipeline evolution of facial capture and tracking for digital humans. Multimedia Syst., 29(4):1917–1940. DOI: 10.1007/s00530-023-01081-2.

Wang, R., Cao, Y., Han, K., and Wong, K.-Y. K. (2024). A survey on 3d human avatar modeling - from reconstruction to generation. ArXiv, abs/2406.04253. DOI: 10.48550/arXiv.2406.04253.

Wang, Y., Sun, Y., Patel, P., Daniilidis, K., Black, M. J., and Kocabas, M. (2025). Prompthmr: Promptable human mesh recovery. DOI: 10.48550/arXiv.2504.06397.

Weng, S. C., Chiou, Y., and Do, E. Y. (2024). Dream mesh: A speech-to-3d model generative pipeline in mixed reality. In 2024 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR), pages 345-349. IEEE. DOI: 10.1109/AIxVR59861.2024.00059.

Xu, Y., Yang, Z., and Yang, Y. (2026). Photorealistic text-to-3d avatar generation with constraints for decoupled geometry and appearance. ACM Trans. Multimedia Comput. Commun. Appl., 22(2). DOI: 10.1145/3774422.

Yin, M., Cao, Y., Peng, S., and Han, K. (2025). Splat4d: Diffusion-enhanced 4d gaussian splatting for temporally and spatially consistent content creation. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, SIGGRAPH Conference Papers '25, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3721238.3730752.

Yu, Z., Zhao, Z., Du, Y., Zheng, Y., Zuo, B., and Wang, Y. (2025). T2c: Text-guided 4d cloth generation. ACM Trans. Multimedia Comput. Commun. Appl., 21(7). DOI: 10.1145/3735642.

Zhang, J., Jiang, Z., Yang, D., Xu, H., Shi, Y., Song, G., Xu, Z., Wang, X., and Feng, J. (2022). Avatargen: A 3d generative model for animatable human avatars. In Computer Vision – ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part III, page 668–685, Berlin, Heidelberg. Springer-Verlag. DOI: 10.1007/978-3-031-25066-8_39.

Zheng, C., Xue, L., Zarate, J., and Song, J. (2025). Gaustar: Gaussian surface tracking and reconstruction. DOI: 10.48550/arXiv.2501.10283.

Zhu, S., Guo, C., Lu, Y., Yu, J., Shen, Y., Wu, Z., Hong, Y., Cai, Y., Liu, Y., Zhang, M., Xu, Y., and Xu, L. (2026). A general framework for gaussian splatting-based human-centric volumetric videos. Visual Intelligence, 4:8. DOI: 10.1007/s44267-026-00111-7.

Downloads

Published

2026-08-20

How to Cite

Cardoso, D. O. de S., Costa, W. de L., Oliveira, P. A. A. de, Teixeira, J. M. X. N., Teichrieb, V., & Lin, Q. (2026). SPL4SH: Designing a Systematic Pipeline for 4D Synthetic Humans in XR Content Creation. Journal of the Brazilian Computer Society, 32(1), 2131–2151. https://doi.org/10.5753/jbcs.2026.8229

Issue

Section

Regular Issue