Inside the Mobile Community: GUI Testing Discussions and Challenges on Stack Exchange

Authors

DOI:

https://doi.org/10.5753/jserd.2026.6604

Keywords:

Developer expertise, software testing, gui testing, stack exchange mining, testing community

Abstract

The Graphical User Interface (GUI) is increasingly becoming the focus of developers and testers. A well-designed, defect-free interface can significantly affect whether the application functions properly and provides a satisfying user experience. In particular, in the era of smartphones with varying screen sizes and operating systems, mobile developers and testers are increasingly developing and using interface tests. This article aims to characterize the practical landscape of Mobile GUI testing by analyzing discussions among practitioners across six Stack Exchange communities, identifying the main challenges, tools, and platform-specific differences reported by developers and testers. To achieve this goal, we collected and filtered data from Stack Exchange Data Dumps covering the years 2012 to 2023, combining automated keyword-based filtering with manual validation by three independent analysts, resulting in 714 posts and 203 comments for analysis. The yielded findings show that the main challenges are device compatibility, unexpected UI interference, and limitations in automation tools. Appium emerges as the most prevalent tool, mentioned in over 40% of analyzed posts, while platform-specific frameworks such as Espresso and UI Automator (Android) and XCTest and XCode UI Test (iOS) are also widely discussed. The results also show that Android offers greater testing flexibility, whereas iOS imposes stricter restrictions that require additional configuration. This study provides a practitioner-driven synthesis of real-world Mobile GUI testing practices, shedding light on the main trends and practices developed by the community.

Downloads

Download data is not yet available.

References

Adamo, D., Nurmuradov, D., Piparia, S., and Bryce, R. (2018). Combinatorial-based event sequence testing of Android applications. Information and Software Technology, 99:98–117.

Alomar, E. A., Mkaouer, M. W., Newman, C., and Ouni, A. (2021). On preserving the behavior in software refactoring: A systematic mapping study. Information and Software Technology, 140:106675.

Amalfitano, D., Amatucci, N., Memon, A. M., Tramontana, P., and Fasolino, A. R. (2017). A general framework for comparing automatic testing techniques of Android mobile apps. Journal of Systems and Software, 125:322–343.

Amalfitano, D., Fasolino, A. R., Tramontana, P., Ta, B. D., and Memon, A. M. (2014). MobiGUITAR: Automated model-based testing of mobile apps. IEEE Software, 32(5):53–59.

Ammann, P. and Offutt, J. (2016). Introduction to Software Testing. Cambridge University Press, 2nd edition.

Anvik, J., Hiew, L., and Murphy, G. C. (2006). Who should fix this bug? In Proceedings of the 28th International Conference on Software Engineering, ICSE ’06, pages 361–370, New York, NY, USA. ACM.

Aranda, J. and Venolia, G. (2009). The secret life of bugs: Going past the errors and omissions in software repositories. In Proceedings of the 31st International Conference on Software Engineering, ICSE ’09, pages 298–308, Washington, DC, USA. IEEE Computer Society.

Arnatovich, Y. L., Ngo, M. N., Kuan, T. H. B., and Soh, C. (2016). Achieving high code coverage in Android UI testing via automated widget exercising. In Proceedings of the 23rd Asia-Pacific Software Engineering Conference, APSEC ’16, pages 193–200. IEEE.

Arnatovich, Y. L. and Wang, L. (2018). A systematic literature review of automated techniques for functional GUI testing of mobile applications. arXiv preprint, arXiv:1812.11470.

Azim, T. and Neamtiu, I. (2013). Targeted and depth-first exploration for systematic testing of Android apps. ACM SIGPLAN Notices, 48(10):641–660.

Banerjee, I., Nguyen, B., Garousi, V., and Memon, A. (2013). Graphical user interface (GUI) testing: Systematic mapping and repository. Information and Software Technology, 55(10):1679–1694.

Berihun, N. G., Dongmo, C., and Van der Poll, J. A. (2023). The applicability of automated testing frameworks for mobile application testing: A systematic literature review. Computers, 12(5):97.

Bernardi, M. L., Canfora, G., Di Lucca, G. A., Di Penta, M., and Distante, D. (2012). Do developers introduce bugs when they do not communicate? The case of Eclipse and Mozilla. In Proceedings of the 16th European Conference on Software Maintenance and Reengineering, CSMR ’12, pages 139–148, Washington, DC, USA. IEEE Computer Society.

Bettenburg, N., Just, S., Schröter, A., Weiß, C., Premraj, R., and Zimmermann, T. (2007). Quality of bug reports in Eclipse. In Proceedings of the 2007 OOPSLA Workshop on Eclipse Technology eXchange, ETX ’07, pages 21–25, New York, NY, USA. ACM.

Bettenburg, N., Just, S., Schröter, A., Weiss, C., Premraj, R., and Zimmermann, T. (2008). What makes a good bug report? In Proceedings of the 16th ACM SIGSOFT International Symposium on Foundations of Software Engineering, FSE ’08, pages 308–318, New York, NY, USA. ACM.

Boehm, B. W. (1981). Software Engineering Economics. Prentice-Hall, Englewood Cliffs, NJ, USA.

Braun, V. and Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2):77–101.

Bryce, R. C., Sampath, S., and Memon, A. M. (2010). Developing a single model and test prioritization strategies for event-driven software. IEEE Transactions on Software Engineering, 37(1):48–64.

Cavalcanti, Y. C., da Mota Silveira Neto, P. A., Carmo Machado, I., Vale, T. F., de Almeida, E. S., and de Lemos Meira, S. R. (2014). Challenges and opportunities for software change request repositories: A systematic mapping study. Journal of Software: Evolution and Process, 26(7):620–653.

Chaudhary, N. and Sangwan, O. (2016). Metrics for event-driven software. International Journal of Advanced Computer Science and Applications, 7(1).

Choi, W., Necula, G., and Sen, K. (2013). Guided GUI testing of Android apps with minimal restart and approximate learning. ACM SIGPLAN Notices, 48(10):623–640.

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1):37–46.

Cruzes, D. S. and Dybå, T. (2011). Recommended steps for thematic synthesis in software engineering. In Proceedings of the 5th International Symposium on Empirical Software Engineering and Measurement (ESEM ’11), pages 275–284. IEEE.

Deng, L., Offutt, J., Ammann, P., and Mirzaei, N. (2017). Mutation operators for testing Android apps. Information and Software Technology, 81:154–168.

Farley, D. (2021). Modern Software Engineering: Doing What Works to Build Better Software Faster. Addison-Wesley Professional, Boston, MA, USA.

Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5):378–382.

Genc-Nayebi, N. and Abran, A. (2017). A systematic literature review: Opinion mining studies from mobile app store user reviews. Journal of Systems and Software, 125:207–219.

Gomes, F., Santos, E. P. d., Freire, S., Mendonça, M., Mendes, T. S., and Spínola, R. (2022). Investigating the point of view of project management practitioners on technical debt: A preliminary study on Stack Exchange. In Proceedings of the International Conference on Technical Debt, TechDebt ’22, pages 31–40.

Harrold, M. J. (2000). Testing: A roadmap. In Proceedings of the Conference on the Future of Software Engineering, ICSE ’00, pages 61–72, New York, NY, USA. ACM.

Herbold, S. and Harms, P. (2013). AutoQUEST: Automated quality engineering of event-driven software. In Proceedings of the 6th IEEE International Conference on Software Testing, Verification and Validation Workshops, ICSTW ’13, pages 134–139. IEEE.

Herzig, K., Just, S., and Zeller, A. (2013). It's not a bug, it's a feature: How misclassification impacts bug prediction. In Proceedings of the 35th International Conference on Software Engineering, ICSE ’13, pages 392–401, Washington, DC, USA. IEEE Computer Society.

Humble, J. and Farley, D. (2010). Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation. Addison-Wesley Professional.

Indika, A., Lee, C., Wang, H., Lisoway, J., Peruma, A., and Kazman, R. (2024). Exploring accessibility trends and challenges in mobile app development: A study of Stack Overflow questions. arXiv preprint, arXiv:2409.07945.

International Data Corporation (IDC) (2025). Worldwide Quarterly Mobile Phone Tracker — Q4 2025. [link]. Accessed: April 2026.

ISO/IEC/IEEE (2022). Software and Systems Engineering — Software Testing — Part 1: General Concepts.

Júnior, J., Boechat, G., and Machado, I. (2021). Label it be! A large-scale study of issue labeling in modern open-source repositories. arXiv preprint, arXiv:2110.01328.

Júnior, J. M. and Machado, I. (2020). Avaliação empírica de termos técnicos em issues de projetos open source. Revista Eletrônica de Iniciação Científica em Computação, 18(3).

Junior, N., Costa, H. A. X., Karita, L., Machado, I., and Soares, L. (2021). Experiences and practices in GUI functional testing: A software practitioners’ view. In Vasconcellos, C. D., Roggia, K. G., Collere, V., and Bousfield, P., editors, 35th Brazilian Symposium on Software Engineering (SBES 2021), Joinville, Santa Catarina, Brazil, 27 September–1 October 2021, pages 195–204. ACM.

Kanwal, J. and Maqbool, O. (2012). Bug prioritization to facilitate bug report triage. Journal of Computer Science and Technology, 27(2):397–412.

Kim, S. and Whitehead, E. J. (2006). How long did it take to fix bugs? In Proceedings of the 2006 International Workshop on Mining Software Repositories, MSR ’06, pages 173–174, New York, NY, USA. ACM.

Kousar, S. and Javed, A. (2025). A systematic literature review on graphical user interface testing through software patterns. IET Software, 2025:9140693.

Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1):159–174.

Li, A., Qin, Z., Chen, M., and Liu, J. (2014). ADAutomation: An activity diagram based automated GUI testing framework for smartphone applications. In Proceedings of the 8th International Conference on Software Security and Reliability, SERE ’14, pages 68–77. IEEE.

Lientz, B. P., Swanson, E. B., and Tompkins, G. E. (1978). Characteristics of application software maintenance. Communications of the ACM, 21(6):466–471.

Linares-Vásquez, M., Dit, B., and Poshyvanyk, D. (2013). An exploratory analysis of mobile development issues using Stack Overflow. In Proceedings of the 10th Working Conference on Mining Software Repositories, MSR ’13, pages 93–96. IEEE Press.

Mahajan, R. and Shneiderman, B. (1997). Visual and textual consistency checking tools for graphical user interfaces. IEEE Transactions on Software Engineering, 23(11):722–735.

Mahmood, R., Mirzaei, N., and Malek, S. (2014). EvoDroid: Segmented evolutionary testing of Android apps. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE ’14), pages 599–609, New York, NY, USA. ACM.

Mao, K., Harman, M., and Jia, Y. (2016). Sapienz: Multi-objective automated testing for Android applications. In Proceedings of the 25th International Symposium on Software Testing and Analysis, ISSTA ’16, pages 94–105, New York, NY, USA. ACM.

Martins, L., Campos, D., Santana, R., Junior, J. M., Costa, H., and Machado, I. (2023). Hearing the voice of experts: Unveiling Stack Exchange communities’ knowledge of test smells. In Proceedings of the 16th IEEE/ACM International Conference on Cooperative and Human Aspects of Software Engineering, CHASE ’23, pages 80–91. IEEE.

Mazuera-Rozo, A., Trubiani, C., Linares-Vásquez, M., and Bavota, G. (2020). Investigating types and survivability of performance bugs in mobile apps. Empirical Software Engineering, 25:1644–1686.

McConnell, S. (1996). Daily build and smoke test. IEEE Software, 13(4):144.

McMinn, P. (2011). Search-based software testing: Past, present and future. In Proceedings of the 4th IEEE International Conference on Software Testing, Verification and Validation Workshops, ICSTW ’11, pages 153–163, Washington, DC, USA. IEEE Computer Society.

Memon, A. M. (2002). GUI testing: Pitfalls and process. IEEE Computer, 35(8):87–88.

Memon, A. M., Pollack, M. E., and Soffa, M. L. (2001). Hierarchical GUI test case generation using automated planning. IEEE Transactions on Software Engineering, 27(2):144–155.

Mockus, A. (2010). Organizational volatility and its effects on software defects. In Proceedings of the 18th ACM SIGSOFT International Symposium on Foundations of Software Engineering, FSE ’10, pages 117–126, New York, NY, USA. ACM.

Myers, B. A. (1995). User interface software tools. ACM Transactions on Computer-Human Interaction, 2(1):64–103.

Nascimento, L. P. G., Santos, A., Machado, I., et al. (2024). Issue labeling dynamics in open-source projects: A comprehensive analysis. In Anais do Simpósio Brasileiro de Componentes, Arquiteturas e Reutilização de Software, SBCARS ’24, pages 51–60. SBC.

Nie, L., Said, K. S., Ma, L., Zheng, Y., and Zhao, Y. (2023). A systematic mapping study for graphical user interface testing on mobile apps. IET Software, 17(3):249–267.

Palomba, F., Di Nucci, D., Panichella, A., Zaidman, A., and De Lucia, A. (2019). On the impact of code smells on the energy consumption of mobile applications. Information and Software Technology, 105:43–55.

Posnett, D., Warburg, E., Devanbu, P., and Filkov, V. (2012). Mining Stack Exchange: Expertise is evident from initial contributions. In Proceedings of the 2012 International Conference on Social Informatics, SocialInformatics ’12, pages 199–204. IEEE.

Rajlich, V. (2014). Software evolution and maintenance. In Future of Software Engineering Proceedings, FOSE ’14, pages 133–144, New York, NY, USA. ACM.

Rosen, C. and Shihab, E. (2016). What are mobile developers asking about? A large-scale study using Stack Overflow. Empirical Software Engineering, 21:1192–1223.

Sadeghi, A., Jabbarvand, R., and Malek, S. (2017). PATDroid: Permission-aware GUI testing of Android. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering, ESEC/FSE ’17, pages 220–232, New York, NY, USA. ACM.

Seaman, C. B. (1999). Qualitative methods in empirical studies of software engineering. IEEE Transactions on Software Engineering, 25(4):557–572.

Silva, G. and Santos, R. S. (2023). Comparing mobile testing tools using documentary analysis. In 2023 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), pages 1–6. IEEE.

Song, W., Qian, X., and Huang, J. (2017). EHBDroid: Beyond GUI testing for Android applications. In Proceedings of the 32nd IEEE/ACM International Conference on Automated Software Engineering, ASE ’17, pages 27–37. IEEE Press.

Su, T., Meng, G., Chen, Y., Wu, K., Yang, W., Yao, Y., Pu, G., Liu, Y., and Su, Z. (2017). Guided, stochastic model-based GUI testing of Android apps. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering, ESEC/FSE ’17, pages 245–256, New York, NY, USA. ACM.

Tahir, A., Dietrich, J., Counsell, S., Licorish, S., and Yamashita, A. (2020). A large-scale study on how developers discuss code smells and anti-patterns in Stack Exchange sites. Information and Software Technology, 125:106333.

Tassey, G. (2002). The Economic Impacts of Inadequate Infrastructure for Software Testing. Planning Report 02-3, National Institute of Standards and Technology (NIST), Gaithersburg, MD, USA. U.S. Department of Commerce.

Villanes, I. K., Ascate, S. M., Gomes, J., and Dias-Neto, A. C. (2017). What are software engineers asking about Android testing on Stack Overflow? In Proceedings of the 31st Brazilian Symposium on Software Engineering, SBES ’17, pages 104–113, New York, NY, USA. ACM.

Wang, J. and Wu, J. (2019). Research on mobile application automation testing technology based on Appium. In 2019 International Conference on Virtual Reality and Intelligent Systems (ICVRIS), pages 247–250. IEEE.

Wu, X., Jiang, Y., Xu, C., Cao, C., Ma, X., and Lu, J. (2016). Testing Android apps via guided gesture event generation. In Proceedings of the 23rd Asia-Pacific Software Engineering Conference (APSEC ’16), pages 201–208. IEEE.

Xuan, J., Jiang, H., Ren, Z., and Zou, W. (2012). Developer prioritization in bug repositories. In Proceedings of the 34th International Conference on Software Engineering, ICSE ’12, pages 25–35, Washington, DC, USA. IEEE Computer Society.

Yu, S., Fang, C., Tuo, Z., Zhang, Q., Chen, C., Chen, Z., and Su, Z. (2025). Vision-based mobile app GUI testing: A survey. ACM Computing Surveys, 58(6):1–46.

Yuan, X. and Memon, A. M. (2007). Using GUI run-time state as feedback to generate test cases. In Proceedings of the 29th International Conference on Software Engineering, ICSE ’07, pages 396–405, Washington, DC, USA. IEEE Computer Society.

Yuan, X. and Memon, A. M. (2009). Generating event sequence-based test cases using GUI run-time state feedback. IEEE Transactions on Software Engineering, 36(1):81–95.

Zein, S., Salleh, N., and Grundy, J. (2016). A systematic mapping study of mobile application testing techniques. Journal of Systems and Software, 117:334–356.

Zubrow, D. (2009). IEEE Standard Classification for Software Anomalies. IEEE Std 1044-2009.

Downloads

Published

2026-08-06

How to Cite

Marçal, M. C. P., Mota Junior, J., Nascimento, L. P. G., Sant’anna, C., & Machado, I. (2026). Inside the Mobile Community: GUI Testing Discussions and Challenges on Stack Exchange. Journal of Software Engineering Research and Development, 14(1), 259–278. https://doi.org/10.5753/jserd.2026.6604

Issue

Section

Research Article