On Human Perception-Guided Video Stability Assessment
DOI:
https://doi.org/10.5753/jbcs.2026.5618Keywords:
Video Stability, In-Capture Video Distortion, Assessment, Human PerceptionAbstract
User experience can be compromised if a video contains an in-capture distortion called camera instability, characterized by a typical visual effect caused by uncontrolled camera motion during recording. The process of removing this distortion is known as video stabilization. Despite recent advances in this field, few studies aim to define how video stability quality should be evaluated. In this work, we propose assessment metrics based on pixel profile derivatives and machine learning regressors. We compare them with existing measures on public datasets, correlating the results with human perception scores. Our findings indicate that a widely used measure, called Low-High Frequency Ratio (LHR), correlates poorly with human perception, achieving PLCC=0.388 on LIVE-Qualcomm and PLCC=0.538 on MIND-VQ. In contrast, the best result obtained in this work achieves PLCC=0.884 on MIND-VQ, improving LHR by 34.6 percentage points on that dataset.
Downloads
References
Adelson, E. H. and Bergen, J. R. (1985). Spatiotemporal Energy Models for the Perception of Motion. Journal of the Optical Society of America, 2(2):284-299. DOI: 10.1364/josaa.2.000284.
Bolles, R. and Baker, H. (1985). Epipolar-Plane Image Analysis: A Technique for Analyzing Motion Sequences. In 3th IEEE Workshop on Computer Vision, Representation, and Control, pages 168-178. IEEE. DOI: 10.1016/B978-0-08-051581-6.50009-X.
Ghadiyaram, D., Pan, J., Bovik, A. C., Moorthy, A. K., Panda, P., and Yang, K.-C. (2017). In-capture Mobile Video Distortions: A Study of Subjective Behavior and Objective Algorithms. IEEE Transactions on Circuits and Systems for Video Technology, 28(9):2061-2077. DOI: 10.1109/tcsvt.2017.2707479.
Grundmann, M., Kwatra, V., and Essa, I. (2011). Auto-Directed Video Stabilization with Robust L1 Optimal Camera Paths. In Conference on Computer Vision and Pattern Recognition, pages 225-232. IEEE. DOI: 10.1109/cvpr.2011.5995525.
Guan, X., Li, F., Huang, Z., and Liu, H. (2022). Study of Subjective and Objective Quality Assessment of Night-Time Videos. Transactions on Circuits and Systems for Video Technology, 32(10):6627-6641. DOI: 10.1109/tcsvt.2022.3177518.
Halperin, T., Poleg, Y., Arora, C., and Peleg, S. (2017). Egosampling: Wide View Hyperlapse From Egocentric Videos. IEEE Transactions on Circuits and Systems for Video Technology, 28(5):1248-1259. DOI: 10.1109/tcsvt.2017.2651051.
Ito, M. S. and Izquierdo, E. (2019). A Dataset and Evaluation Framework for Deep Learning Based Video Stabilization Systems. In Visual Communications and Image Processing, pages 1-4. IEEE. DOI: 10.1109/vcip47243.2019.8966057.
James, J. G., Jain, D., and Rajwade, A. (2023). GlobalFlowNet: Video Stabilization using Deep Distilled Global Motion Estimates. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5078-5087. DOI: 10.1109/wacv56688.2023.00505.
Joshi, N., Kienzle, W., Toelle, M., Uyttendaele, M., and Cohen, M. F. (2015). Real-Time Hyperlapse Creation Via Optimal Frame Selection. ACM Transactions on Graphics, 34(4):1-9. DOI: 10.1145/2766954.
Kopf, J., Cohen, M. F., and Szeliski, R. (2014). First-Person Hyper-Lapse Videos. ACM Transactions on Graphics, 33(4):1-10. DOI: 10.1145/2601097.2601195.
Kossi, K., Coulombe, S., Desrosiers, C., and Gagnon, G. (2022). No-Reference Video Quality Assessment Using Distortion Learning and Temporal Attention. IEEE Access, 10:41010-41022. DOI: 10.1109/access.2022.3167446.
Kou, T., Liu, X., Sun, W., Jia, J., Min, X., Zhai, G., and Liu, N. (2023). StableVQA: A Deep No-Reference Quality Assessment Model for Video Stability. In ACM International Conference on Multimedia, pages 1066-1076. DOI: 10.1145/3581783.3611860.
Lai, W.-S., Huang, Y., Joshi, N., Buehler, C., Yang, M.-H., and Kang, S. B. (2017). Semantic-Driven Generation Of Hyperlapse From 360 Degree Video. IEEE Transactions on Visualization and Computer Graphics, 24(9):2610-2621. DOI: 10.1109/tvcg.2017.2750671.
Lee, Y.-C., Tseng, K.-W., Chen, Y.-T., Chen, C.-C., Chen, C.-S., and Hung, Y.-P. (2021). 3D Video Stabilization with Depth Estimation by CNN-based Optimization. In IEEE Conference on Computer Vision and Pattern Recognition, pages 10621-10630. DOI: 10.1109/cvpr46437.2021.01048.
Liu, S., Yuan, L., Tan, P., and Sun, J. (2013). Bundled Camera Paths for Video Stabilization. ACM Transactions on Graphics, 32(4):1-10. DOI: 10.1145/2461912.2461995.
Liu, S., Yuan, L., Tan, P., and Sun, J. (2014). Steadyflow: Spatially Smooth Optical Flow for Video Stabilization. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4209-4216. DOI: 10.1109/cvpr.2014.536.
Liu, X., Min, X., Sun, W., Zhang, Y., Zhang, K., Timofte, R., Zhai, G., Gao, Y., Cao, Y., Kou, T., et al. (2023). NTIRE 2023 Quality Assessment of Video Enhancement Challenge. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1551-1569. DOI: 10.1109/cvprw59228.2023.00158.
Min, X., Duan, H., Sun, W., Zhu, Y., and Zhai, G. (2024). Perceptual Video Quality Assessment: A Survey. Science China Information Sciences, 67(11):211301. DOI: 10.1007/s11432-024-4133-3.
Morimoto, C. and Chellappa, R. (1997). Evaluation of Image Stabilization Algorithms. In DARPA Image Understanding Workshop, pages 295-302. DOI: 10.1109/icassp.1998.678102.
Nie, Y., Su, T., Zhang, Z., Sun, H., and Li, G. (2017). Dynamic Video Stitching via Shakiness Removing. IEEE Transactions on Image Processing, 27(1):164-178. DOI: 10.1109/tip.2017.2736603.
Nuutinen, M., Virtanen, T., Vaahteranoksa, M., Vuori, T., Oittinen, P., and Häkkinen, J. (2016). CVD2014—A Database for Evaluating No-Reference Video Quality Assessment Algorithms. IEEE Transactions on Image Processing, 25(7):3073-3086. DOI: 10.1109/tip.2016.2562513.
Peng, Z., Ye, X., Zhao, W., Liu, T., Sun, H., Li, B., and Cao, Z. (2024). 3D Multi-Frame Fusion for Video Stabilization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7507-7516. DOI: 10.1109/cvpr52733.2024.00717.
Poleg, Y., Halperin, T., Arora, C., and Peleg, S. (2015). Egosampling: Fast-Forward And Stereo For Egocentric Videos. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4768-4776. DOI: 10.1109/cvpr.2015.7299109.
Rao, Q., Yu, X., Navasardyan, S., and Shi, H. (2023). Sim2RealVS: A New Benchmark for Video Stabilization With a Strong Baseline. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5406-5415. DOI: 10.1109/wacv56688.2023.00537.
Shi, Z., Shi, F., Lai, W.-S., Liang, C.-K., and Liang, Y. (2022). Deep Online Fused Video Stabilization. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1250-1258. DOI: 10.1109/wacv51458.2022.00094.
Silva, M. M., Ramos, W. L. S., Ferreira, J. P. K., Campos, M. F. M., and Nascimento, E. R. (2016). Towards Semantic Fast-Forward And Stabilized Egocentric Videos. In European Conference on Computer Vision, pages 557-571. Springer. DOI: 10.1007/978-3-319-46604-0_40.
Souza, M. R., Almeida, H. M., and Pedrini, H. (2024a). Digital Video Stabilization: Methods, Datasets, and Evaluation. In 37th Conference on Graphics, Patterns and Images (SIBGRAPI), pages 42-48. SBC. DOI: 10.5753/sibgrapi.est.2024.31643.
Souza, M. R., Almeida Maia, H., and Pedrini, H. (2023). Rethinking Two-Dimensional Camera Motion Estimation Assessment for Digital Video Stabilization: A Camera Motion Field-based Metric. Neurocomputing, 559:126768. DOI: 10.1016/j.neucom.2023.126768.
Souza, M. R., Almeida Maia, H., Vieira, M. B., and Pedrini, H. (2020). Survey on Visual Rhythms: A Spatio-Temporal Representation for Video Sequences. Neurocomputing, 402:409-422. DOI: 10.1016/j.neucom.2020.04.035.
Souza, M. R., da Fonseca, L. F. R., and Pedrini, H. (2018). Improvement of Global Motion Estimation in Two-Dimensional Digital Video Stabilisation Methods. Image Processing, 12(12):2204-2211. DOI: 10.1049/iet-ipr.2018.5445.
Souza, M. R., Maia, H. A., and Pedrini, H. (2022). Survey on Digital Video Stabilization: Concepts, Methods, and Challenges. ACM Computing Surveys, 55(3):1-37. DOI: 10.1145/3494525.
Souza, M. R., Maia, H. A., and Pedrini, H. (2024b). NAFT and SynthStab: A RAFT-based Network and a Synthetic Dataset for Digital Video Stabilization. International Journal of Computer Vision, pages 1-26. DOI: 10.1007/s11263-024-02264-8.
Souza, M. R. and Pedrini, H. (2018a). Combination of Local Feature Detection Methods for Digital Video Stabilization. Signal, Image and Video Processing, 12(8):1513-1521. DOI: 10.1007/s11760-018-1307-8.
Souza, M. R. and Pedrini, H. (2018b). Digital Video Stabilization based on Adaptive Camera Trajectory Smoothing. EURASIP Journal on Image and Video Processing, 2018(1):37. DOI: 10.1186/s13640-018-0277-7.
Souza, M. R. and Pedrini, H. (2019). Motion Energy Image for Evaluation of Video Stabilization. The Visual Computer, 35(12):1769-1781. DOI: 10.1007/s00371-018-1572-0.
Souza, M. R. and Pedrini, H. (2020). Visual Rhythms for Qualitative Evaluation of Video Stabilization. EURASIP Journal on Image and Video Processing, 2020:1-19. DOI: 10.1186/s13640-020-00508-4.
Streijl, R. C., Winkler, S., and Hands, D. S. (2016). Mean Opinion Score (MOS) Revisited: Methods and Applications, Limitations and Alternatives. Multimedia Systems, 22(2):213-227. DOI: 10.1007/s00530-014-0446-1.
Teed, Z. and Deng, J. (2020). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In European Conference on Computer Vision, pages 402-419. Springer. DOI: 10.1007/978-3-030-58536-5_24.
Tezcan, M. O., Ishwar, P., and Konrad, J. (2021). BSUV-Net 2.0: Spatio-Temporal Data Augmentations for Video-Agnostic Supervised Background Subtraction. IEEE Access, 9:53849-53860. DOI: 10.1109/access.2021.3071163.
Tezcan, O., Ishwar, P., and Konrad, J. (2020). BSUV-Net: A Fully-Convolutional Neural Network for Background Subtraction of Unseen Videos. In IEEE Winter Conference on Applications of Computer Vision, pages 2774-2783. DOI: 10.1109/wacv45572.2020.9093464.
Wang, N., Zhou, C., Zhu, R., Zhang, B., Wang, Y., and Liu, H. (2024a). SOFT: Self-Supervised Sparse Optical Flow Transformer for Video Stabilization via Quaternion. Engineering Applications of Artificial Intelligence, 130:107725. DOI: 10.1016/j.engappai.2023.107725.
Wang, Y., Huang, Q., Liu, J., Jiang, C., and Shang, M. (2024b). Adaptive Video Stabilization Based on Feature Point Detection and Full-Reference Stability Assessment. Multimedia Tools and Applications, 83(11):32497-32524. DOI: 10.1007/s11042-023-16607-z.
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P. (2004). Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 13(4):600-612. DOI: 10.1109/tip.2003.819861.
Wu, H., Liao, L., Chen, C., Hou, J., Wang, A., Sun, W., Yan, Q., and Lin, W. (2022). Disentangling Aesthetic and Technical Effects for Video Quality Assessment of User Generated Content. arXiv preprint arXiv:2211.04894, 1:1-8. DOI: 10.48550/arXiv.2211.04894.
Zhang, L., Zheng, Q., and Huang, H. (2018a). Intrinsic Motion Stability Assessment for Video Stabilization. IEEE Transactions on Visualization and Computer Graphics, 25(4):1681-1692. DOI: 10.1109/tvcg.2018.2817209.
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. (2018b). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In IEEE Conference on Computer Vision and Pattern Recognition, pages 586-595. DOI: 10.1109/cvpr.2018.00068.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Marcos Roberto Souza, Luís Gustavo Lorgus Decker, Jose Luis Flores Campana, Helena de Almeida Maia, Helio Pedrini

This work is licensed under a Creative Commons Attribution 4.0 International License.

