More than one million test smells: how are Dart projects and their sentiments?
DOI:
https://doi.org/10.5753/jserd.2026.7537Keywords:
Software testing, Test smells, Dart, Flutter, Sentiment analysisAbstract
Objective: This study investigates the quality of tests and the associated sentiments in Dart, the primary language for mobile application development with the Flutter framework. Methods: The study begins by using the DNose tool to detect 14 types of test smells in code written in the Dart language, along with their associated sentiments. Next, we evaluate the tool in terms of precision, recall, accuracy, and F1-score. Using this tool, we conduct a detailed analysis of tests in open-source projects extracted from the language’s central repository. Results: The study starts with a dataset of 5,410 Dart projects, from which 4,154 repositories were successfully cloned after processing. Based on the cloned projects, we generated a dataset containing 1,115,938 occurrences of test smells and their associated sentiments. The analysis allowed us to characterize the most frequently encountered test smells and identify their causes. We observed the presence of test smells in 77% of test files. Another key characteristic observed in the analyzed projects was the scarcity of tests: 1,873 projects had one or no tests, which led us to expand the analysis to a broader base. In addition, we analyzed the relationship between developers’ sentiments and the identified test smells, enabling us to classify test smells according to sentiment. Conclusion: This research makes a significant contribution by providing in-depth insights into the quality of tests and associated sentiments in projects from Dart’s official repository, as well as offering an open-source tool for detecting 14 types of test smells and their associated sentiments.
Downloads
References
Afonso, J., & Campos, J. (2023). Automatic generation of smell-free unit tests. In 2023 IEEE/ACM International Workshop on Search-Based and Fuzz Testing (SBFT) (pp. 9–16). IEEE.
Ahmed, T., Bosu, A., Iqbal, A., & Rahimi, S. (2017). SentiCR: A customized sentiment analysis tool for code review interactions. In 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE) (pp. 106–111). IEEE.
Aljedaani, W., Mkaouer, M. W., Peruma, A., & Ludi, S. (2023). Do the test smells Assertion Roulette and Eager Test impact students’ troubleshooting and debugging capabilities? In Proceedings of the 45th International Conference on Software Engineering: Software Engineering Education and Training (ICSE-SEET ’23) (pp. 29–39). IEEE Press.
Aljedaani, W., Peruma, A., Aljohani, A., Alotaibi, M., Mkaouer, M. W., Ouni, A., Newman, C. D., Ghallab, A., & Ludi, S. (2021). Test smell detection tools: A systematic mapping study. In Proceedings of the 25th International Conference on Evaluation and Assessment in Software Engineering (EASE ’21) (pp. 170–180). Association for Computing Machinery.
Bavota, G., Qusef, A., Oliveto, R., De Lucia, A., & Binkley, D. (2015). Are test smells really harmful? An empirical study. Empirical Software Engineering, 20, 1052–1094.
Calefato, F., Lanubile, F., Maiorano, F., & Novielli, N. (2018). Sentiment polarity detection for software development. Empirical Software Engineering, 23, 1352–1382.
Camara, B., Silva, M., Endo, A., & Vergilio, S. (2021). On the use of test smells for prediction of flaky tests. In Proceedings of the 6th Brazilian Symposium on Systematic and Automated Software Testing (pp. 46–54).
Campos, D., Rocha, L., & Machado, I. (2021). Developers’ perception on the severity of test smells: An empirical study. In 24th Iberoamerican Conference on Software Engineering (CIbSE 2021) (pp. 192–205). Curran Associates. Also available as arXiv preprint arXiv:2107.13902.
Clark, B. (2013). Cellular phones as a primary communications device: What are the implications for a global community? Global Media Journal, 12.
De Stefano, M., Pecorelli, F., Di Nucci, D., & De Lucia, A. (2022). A preliminary evaluation on the relationship among architectural and test smells. In 2022 IEEE 22nd International Working Conference on Source Code Analysis and Manipulation (SCAM) (pp. 66–70).
Dey, K., Singh, J., Cao, H., & Palma, F. (2025). Pushing feelings: Emotion and sentiment in software commit messages. In IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE.
Ding, J., Fan, G., Yu, H., & Huang, Z. (2022). Automatic identification of high-impact bug report by product and test code quality. International Journal of Software Engineering and Knowledge Engineering, 32(6), 893–916.
Fard, A. M., & Mesbah, A. (2013). JSNose: Detecting JavaScript code smells. In 2013 IEEE 13th International Working Conference on Source Code Analysis and Manipulation (SCAM) (pp. 116–125).
Fernandes, D., Machado, I., & Maciel, R. (2021). Handling test smells in Python: Results from a mixed-method study. In Proceedings of the XXXV Brazilian Symposium on Software Engineering (pp. 84–89).
Fernandes, D., Machado, I., & Maciel, R. (2022). TemPy: Test smell detector for Python. In Proceedings of the XXXVI Brazilian Symposium on Software Engineering (SBES ’22) (pp. 214–219). Association for Computing Machinery.
Garousi, V., & Küçük, B. (2018). Smells in software test code: A survey of knowledge in industry and academia. Journal of Systems and Software, 138, 52–81.
Guzman, E., Azócar, D., & Li, Y. (2014). Sentiment analysis of commit comments in GitHub: An empirical study. In Proceedings of the 11th Working Conference on Mining Software Repositories (MSR 2014) (pp. 352–355). ACM.
Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34.
Islam, M. R., & Zibran, M. F. (2018). SentiStrength-SE: Exploiting domain specificity for improved sentiment analysis in software engineering text. Journal of Systems and Software, 145, 125–146. Elsevier.
Jorge, D., Machado, P., & Andrade, W. (2021a). Investigating test smells in JavaScript test code. In Anais do VI Simpósio Brasileiro de Testes de Software Sistemático e Automatizado (pp. 36–45). SBC.
Jorge, D., Machado, P., & Andrade, W. (2021b). Investigating test smells in JavaScript test code. In Proceedings of the 6th Brazilian Symposium on Systematic and Automated Software Testing (SAST ’21) (pp. 36–45). Association for Computing Machinery.
Junior, N. S., Martins, L., Rocha, L., Costa, H., & Machado, I. (2021). How are test smells treated in the wild? A tale of two empirical studies. Journal of Software Engineering Research and Development, 9, 9–1.
Junior, N. S., Rocha, L., Martins, L. A., & Machado, I. (2020). A survey on test practitioners’ awareness of test smells. In Proceedings of the XXIII Iberoamerican Conference on Software Engineering (CIbSE 2020) (pp. 462–475). Curran Associates. Also available as arXiv preprint arXiv:2003.05613.
Kaur, R., Chahal, K., & Saini, M. (2022). Analysis of factors influencing developers’ sentiments in commit logs: Insights from applying sentiment analysis. e-Informatica Software Engineering Journal, 16(1).
Kim, D. J., Chen, T.-H., & Yang, J. (2021). The secret life of test smells: An empirical study on test smell evolution and maintenance. Empirical Software Engineering, 26.
Lin, B., Zampetti, F., Bavota, G., Di Penta, M., Lanza, M., & Oliveto, R. (2018). Sentiment analysis for software engineering: How far can we go? In Proceedings of the 40th International Conference on Software Engineering (ICSE) (pp. 94–104). ACM.
Martins, L., Bezerra, C., Costa, H., & Machado, I. (2021). Smart prediction for refactorings in the software test code. In Proceedings of the XXXV Brazilian Symposium on Software Engineering (SBES ’21) (pp. 115–120). Association for Computing Machinery.
Martins, L., Costa, H., & Machado, I. (2024). On the diffusion of test smells and their relationship with test code quality of Java projects. Journal of Software: Evolution and Process, 36(4), e2532.
Miller, C., Widder, D. G., Kästner, C., & Vasilescu, B. (2019). Why do people give up flossing? A study of contributor disengagement in open source. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Society (ICSE-SEIS ’19) (pp. 116–127). IEEE.
Nielsen, F. Å. (2011). A new ANEW: Evaluation of a word list for sentiment analysis in microblogs. arXiv preprint, arXiv:1103.2903.
Paula, E., & Bonifácio, R. (2022). TestAxe: Automatically refactoring test smells using JUnit 5 features. In Anais Estendidos do XIII Congresso Brasileiro de Software: Teoria e Prática (pp. 89–98). SBC.
Peruma, A., Almalki, K., Newman, C. D., Mkaouer, M. W., Ouni, A., & Palomba, F. (2019a). On the distribution of test smells in open source Android applications: An exploratory study. In Proceedings of the 29th Annual International Conference on Computer Science and Software Engineering (CASCON ’19) (pp. 193–202). IBM Corp.
Peruma, A., Almalki, K., Newman, C. D., Mkaouer, M. W., Ouni, A., & Palomba, F. (2019b). On the distribution of test smells in open source Android applications: An exploratory study. In Proceedings of the 29th Annual International Conference on Computer Science and Software Engineering (CASCON ’19) (pp. 193–202). IBM Corp.
Peruma, A., Almalki, K., Newman, C. D., Mkaouer, M. W., Ouni, A., & Palomba, F. (2020a). TSDetect: An open source test smells detection tool. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2020) (pp. 1650–1654). Association for Computing Machinery.
Peruma, A., & Newman, C. D. (2021). On the distribution of "simple stupid bugs" in unit test files: An exploratory study. In 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR) (pp. 525–529). IEEE.
Peruma, A., Newman, C. D., Mkaouer, M. W., Ouni, A., & Palomba, F. (2020b). An exploratory study on the refactoring of unit test files in Android applications. In Proceedings of the IEEE/ACM 42nd International Conference on Software Engineering Workshops (ICSEW ’20) (pp. 350–357). Association for Computing Machinery.
Pontillo, V., Amoroso d’Aragona, D., Pecorelli, F., Di Nucci, D., Ferrucci, F., & Palomba, F. (2024). Machine learning-based test smell detection. Empirical Software Engineering, 29(2), 55.
Raman, N., Cao, M., Tsvetkov, Y., Kästner, C., & Vasilescu, B. (2020). Stress and burnout in open source: Toward finding, understanding, and mitigating unhealthy interactions. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER ’20) (pp. 57–60). ACM.
Santana, R., Fernandes, D., Campos, D., Soares, L., Maciel, R., & Machado, I. (2021). Understanding practitioners’ strategies to handle test smells: A multi-method study. In Proceedings of the XXXV Brazilian Symposium on Software Engineering (pp. 49–53).
Santana, R., Martins, L., Rocha, L., Virgínio, T., Cruz, A., Costa, H., & Machado, I. (2020a). RAIDe: A tool for Assertion Roulette and Duplicate Assert identification and refactoring. In Proceedings of the XXXIV Brazilian Symposium on Software Engineering (SBES ’20) (pp. 374–379). Association for Computing Machinery.
Santana, R., Martins, L., Rocha, L., Virgínio, T., Cruz, A., Costa, H., & Machado, I. (2020b). RAIDe: A tool for Assertion Roulette and Duplicate Assert identification and refactoring. In Proceedings of the XXXIV Brazilian Symposium on Software Engineering (pp. 374–379).
Santana, R., Martins, L., Virgínio, T., Rocha, L., Costa, H., & Machado, I. (2024). An empirical evaluation of RAIDe: A semi-automated approach for test smells detection and refactoring. Science of Computer Programming, 231, 103013.
Santana, R., Martins, L., Virgínio, T., Soares, L., Costa, H., & Machado, I. (2022). Refactoring Assertion Roulette and Duplicate Assert test smells: A controlled experiment. In Anais do XXV Congresso Ibero-Americano em Engenharia de Software (pp. 263–277). SBC.
Sinha, V., Lazar, A., & Sharber, B. (2016). Analyzing developer sentiment in commit logs. In Proceedings of the 13th International Conference on Mining Software Repositories (MSR ’16) (pp. 520–523). ACM.
Soares, E., Aranda, M., Oliveira, N., Ribeiro, M., Gheyi, R., Souza, E., Machado, I., Santos, A., Fonseca, B., & Bonifácio, R. (2023a). Manual tests do smell! Cataloging and identifying natural language test smells. In 2023 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (pp. 1–11).
Soares, E., III, M. A., Romão, D., & Ribeiro, M. (2023b). The Open Catalog of Test Smells. Available at [link].
Soares, E., Ribeiro, M., Amaral, G., Gheyi, R., Fernandes, L., Garcia, A., Fonseca, B., & Santos, A. (2020). Refactoring test smells: A perspective from open-source developers. In Proceedings of the 5th Brazilian Symposium on Systematic and Automated Software Testing (SAST ’20) (pp. 50–59). Association for Computing Machinery.
Soares, E., Ribeiro, M., Gheyi, R., Amaral, G., & Santos, A. (2023c). Refactoring test smells with JUnit 5: Why should developers keep up-to-date? IEEE Transactions on Software Engineering, 49(3), 1152–1170.
Spadini, D., Palomba, F., Zaidman, A., Bruntink, M., & Bacchelli, A. (2018). On the relation of test smells to software code quality. In 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME) (pp. 1–12).
Statista. (2024). Cross-platform mobile frameworks used by software developers worldwide from 2019 to 2023. [link]. Accessed: 2024-07-15.
van Deursen, A., Moonen, L., van den Bergh, A., & Kok, G. (2001). Refactoring test code. In M. Marchesi & G. Succi (Eds.), Proceedings of the 2nd International Conference on Extreme Programming and Flexible Processes in Software Engineering (XP2001).
Veloso, V., & Hora, A. (2022). Characterizing high-quality test methods: A first empirical study. In 2022 IEEE/ACM 19th International Conference on Mining Software Repositories (MSR) (pp. 265–269).
Virgínio, T., Martins, L., Rocha, L., Santana, R., Cruz, A., Costa, H., & Machado, I. (2020). JNose: Java test smell detector. In Proceedings of the XXXIV Brazilian Symposium on Software Engineering (SBES ’20) (pp. 564–569). Association for Computing Machinery.
Virgínio, T., Ribeiro, M., & Machado, I. (2025a). DNose: Dart test smell detector. In Anais do XXXIX Simpósio Brasileiro de Engenharia de Software (pp. 976–982). SBC.
Virgínio, T., Ribeiro, M., & Machado, I. (2025b). On the prevalence of test smells in mobile development. In Anais do X Simpósio Brasileiro de Testes de Software Sistemático e Automatizado (pp. 84–93). SBC.
Virgínio, T., Ribeiro, M., & Machado, I. (2026). Dataset: More than one million test smells. https://doi.org/10.5281/zenodo.18436532.
Virgínio, T., Santana, R., Martins, L. A., Soares, L. R., Costa, H., & Machado, I. (2019). On the influence of test smells on test coverage. In Proceedings of the XXXIII Brazilian Symposium on Software Engineering (pp. 467–471).
Wang, T., Golubev, Y., Smirnov, O., Li, J., Bryksin, T., & Ahmed, I. (2022). PyNose: A test smell detector for Python. In Proceedings of the ACM/IEEE International Conference on Automated Software Engineering (ASE ’21) (pp. 593–605). IEEE Press.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Tássio Virgínio, Márcio Ribeiro, Ivan Machado

This work is licensed under a Creative Commons Attribution 4.0 International License.

