Revisiting CNNs for Monocular Depth Estimation: A Synthetic-to-Real Perspective
DOI:
https://doi.org/10.5753/jbcs.2026.5874Keywords:
Depth Estimation, Convolutional Neural Networks, Synthetic Datasets, Real-World DataAbstract
Monocular depth estimation is crucial for autonomous systems, augmented reality, and scene understanding. While Vision Transformers (ViTs) excel at capturing global context, their need for large amounts of data makes them impractical for smaller datasets, leaving convolutional neural networks (CNNs) as a viable alternative. This study evaluates CNN-based architectures for monocular depth estimation, proposing a fully convolutional model optimized for efficiency and generalization. Trained solely on synthetic datasets with precise depth annotations, our model avoids real-world data limitations. We assess its performance on synthetic validation sets and real-world datasets without fine-tuning, examining its capacity for domain generalization. Results show that CNNs, with optimized architectures and training, achieve strong depth estimation and enable synthetic-to-real transfer, with an AbsREL of 0.077 on KITTI and 0.099 on NYU Depth. These findings challenge the assumption that real-world data is essential for generalization.
Downloads
References
Agarap, A. (2018). Deep Learning using Rectified Linear Units (ReLU). arXiv preprint arXiv:1803.08375, pages 1-7. DOI: 10.48550/arXiv.1803.08375.
Atapour-Abarghouei, A. and Breckon, T. P. (2018). Real-Time Monocular Depth Estimation Using Synthetic Data with Domain Adaptation via Image Style Transfer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2800-2810. DOI: 10.1109/CVPR.2018.00296.
Birkl, R., Wofk, D., and Müller, M. (2023). MiDaS v3. 1 - A Model Zoo for Robust Monocular Relative Depth Estimation. arXiv preprint arXiv:2307.14460, pages 1-14. DOI: 10.48550/arXiv.2307.14460.
Cabon, Y., Murray, N., and Humenberger, M. (2020). Virtual KITTI 2. arXiv preprint arXiv:2001.10773, pages 1-11. DOI: 10.48550/arXiv.2001.10773.
Decker, L. G. L., Campana, J. L. F., Souza, M. R., Almeida Maia, H., and Pedrini, H. (2024). Zero-Shot Synth-to-Real Depth Estimation: From Synthetic Street Scenes to Real-World Data. In 37th SIBGRAPI Conference on Graphics, Patterns and Images, pages 1-6. IEEE. DOI: 10.1109/SIBGRAPI62404.2024.10716265.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image is Worth 16×16 Words. arXiv preprint arXiv:2010.11929, 7. DOI: 10.48550/arXiv.2010.11929.
Eigen, D., Puhrsch, C., and Fergus, R. (2014). Depth Map Prediction from a Single Image Using a Multi-Scale Deep Network. Advances in Neural Information Processing Systems, 27:2366-2374. Available at:[link].
Fonder, M. and Van Droogenbroeck, M. (2019). Mid-Air: A Multi-Modal Dataset for Extremely Low Altitude Drone Flights. In Computer Vision and Pattern Recognition Workshops, pages 1-10. DOI: 10.1109/CVPRW.2019.00081.
Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018). Deep Ordinal Regression Network for Monocular Depth Estimation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2002-2011. DOI: 10.1109/CVPR.2018.00214.
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R. (2013). Vision Meets Robotics: The KITTI Dataset. The International Journal of Robotics Research, 32(11):1231-1237. DOI: 10.1177/0278364913491297.
Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016). Deeper Depth Prediction with Fully Convolutional Residual Networks. In Fourth International Conference on 3D Vision, pages 239-248. IEEE. DOI: 10.1109/3DV.2016.32.
Li, R., Xian, K., Shen, C., Cao, Z., Lu, H., and Hang, L. (2019). Deep Attention-based Classification Network for Robust Depth Prediction. In 14th Asian Conference on Computer Vision, pages 663-678, Perth, Australia. Springer. DOI: 10.1007/978-3-030-20870-7_41.
Li, Z. and Snavely, N. (2018). MegaDepth: Learning Single-View Depth Prediction from Internet Photos. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2041-2050. DOI: 10.1109/CVPR.2018.00218.
Liu, F., Shen, C., and Lin, G. (2015). Deep Convolutional Neural Fields for Depth Estimation from a Single Image. In IEEE Conference on Computer Vision and Pattern Recognition, pages 5162-5170. DOI: 10.1109/CVPR.2015.7299152.
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022). A ConvNet for the 2020s. In Computer Vision and Pattern Recognition, pages 11976-11986. DOI: 10.1109/CVPR52688.2022.01167.
Lopez-Rodriguez, A. and Mikolajczyk, K. (2023). DESC: Domain Adaptation for Depth Estimation via Semantic Consistency. International Journal of Computer Vision, 131(3):752-771. DOI: 10.1007/s11263-022-01718-1.
Loshchilov, I. (2017). Decoupled Weight Decay Regularization. arXiv preprint arXiv:1711.05101. DOI: 10.48550/arXiv.1711.05101.
Loshchilov, I. and Hutter, F. (2016). SGDR: Stochastic Gradient Descent with Warm Restarts. arXiv preprint arXiv:1608.03983. DOI: 10.48550/arXiv.1608.03983.
Masoumian, A., Rashwan, H. A., Cristiano, J., Asif, M. S., and Puig, D. (2022). Monocular Depth Estimation Using Deep Learning: A Review. Sensors, 22(14):5353. DOI: 10.3390/s22145353.
Misra, D. (2019). Mish: A Self Regularized Non-Monotonic Activation Function. arXiv preprint arXiv:1908.08681. DOI: 10.5244/C.34.191.
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2023). DINOv2: Learning Robust Visual Features without Supervision. arXiv preprint arXiv:2304.07193. DOI: 10.48550/arXiv.2304.07193.
Ranftl, R., Bochkovskiy, A., and Koltun, V. (2021). Vision Transformers for Dense Prediction. In IEEE/CVF International Conference on Computer Vision, pages 12179-12188. DOI: 10.1109/ICCV48922.2021.01196.
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V. (2020). Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer. Transactions on Pattern Aanalysis and Machine Intelligence, 44(3):1623-1637. DOI: 10.1109/TPAMI.2020.3019967.
Roberts, M., Ramapuram, J., Ranjan, A., Kumar, A., Bautista, M. A., Paczan, N., Webb, R., and Susskind, J. M. (2021). Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding. In International Conference on Computer Vision, pages 10912-10922. DOI: 10.1109/ICCV48922.2021.01073.
Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. In 18th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), pages 234-241, Munich, Germany. Springer. DOI: 10.1007/978-3-319-24574-4_28.
Ros, G., Sellart, J., Materzynska, J., Vazquez, D., and Lopez, A. M. (2016). The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3234-3243. DOI: 10.1109/CVPR.2016.352.
Roy, A. and Todorovic, S. (2016). Monocular Depth Estimation using Neural Regression Forest. In IEEE Conference on Computer Vision and Pattern Recognition, pages 5506-5514. DOI: 10.1109/CVPR.2016.594.
Shi, W., Caballero, J., Huszár, F., Totz, J., Aitken, A. P., Bishop, R., Rueckert, D., and Wang, Z. (2016). Real-Time Single Image and Video Super-Resolution using an Efficient Sub-pixel Convolutional Neural Network. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1874-1883. DOI: 10.1109/CVPR.2016.207.
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012). Indoor Segmentation and Support Inference from RGBD Images. In European Conference on Computer Vision, pages 746-760. Springer. DOI: 10.1007/978-3-642-33715-4_54.
Spencer, J., Tosi, F., Poggi, M., Arora, R. S., Russell, C., Hadfield, S., Bowden, R., Zhou, G., Li, Z., and Rao, Q. (2024). The Third Monocular Depth Estimation Challenge. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1-14. DOI: 10.1109/CVPRW63382.2024.00005.
Wang, W., Zhu, D., Wang, X., Hu, Y., Qiu, Y., Wang, C., Hu, Y., Kapoor, A., and Scherer, S. (2020). TartanAir: A Dataset to Push the Limits of Visual SLAM. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4909-4916. IEEE. DOI: 10.1109/IROS45743.2020.9341801.
Wrenninge, M. and Unger, J. (2018). Synscapes: A Photorealistic Synthetic Dataset for Street Scene Parsing. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2742-2749. Available at:[link].
Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., and Zhao, H. (2024a). Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371-10381. DOI: 10.1109/CVPR52733.2024.00987.
Yang, L., Kang, B., Huang, Z., Zhao, Z., Xu, X., Feng, J., and Zhao, H. (2024b). Depth Anything V2. 38th Conference on Neural Information Processing Systems, pages 1-37. DOI: 10.52202/079017-0688.
Yang, M., Yu, K., Zhang, C., Li, Z., and Yang, K. (2018). DenseASPP for Semantic Segmentation in Street Scenes. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3684-3692. DOI: 10.1109/CVPR.2018.00388.
Zeng, Z., Wang, D., Yang, F., Park, H., Soatto, S., Lao, D., and Wong, A. (2024). WorDepth: Variational Language Prior for Monocular Depth Estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9708-9719. DOI: 10.1109/CVPR52733.2024.00927.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Luís Gustavo Lorgus Decker, Jose Luis Flores Campana, Marcos Roberto e Souza, Helena de Almeida Maia, Helio Pedrini

This work is licensed under a Creative Commons Attribution 4.0 International License.

