An Empirical Study of Gemini 3 for Detecting Natural Language Test Smells in Manual Test Cases
DOI:
https://doi.org/10.5753/jserd.2026.7576Keywords:
Manual Testing, Test Smells, Large Language Models, DetectionAbstract
Manual testing, in which testers follow natural language instructions to validate system behavior, remains essential for uncovering issues that are difficult to capture with automation. However, manual test cases often contain test smells, quality issues such as ambiguity, redundancy, or missing checks that reduce reliability, maintainability, and reproducibility. Existing detection approaches largely depend on manually engineered rules and thus struggle to generalize and scale across heterogeneous test suites. In our previous work, we assessed the feasibility of using Small Language Models (SLMs) for test smell detection by evaluating Gemma-3-4B, Llama-3.2-3B, and Phi-4-14B on test steps from 143 real-world Ubuntu test cases, covering seven smell types. Phi-4-14B achieved the best performance. In this article, we investigate whether a contemporary Large Language Model (Gemini-3-Pro-Preview) available at the time of the study can identify test smells in natural language manual test cases using a prompt-based, whole-test-case analysis strategy. Unlike approaches that analyze individual test steps in isolation, our approach evaluates complete test cases, enabling the model to consider relationships and dependencies among test steps. We evaluate the approach on 100 Ubuntu test cases covering seven test smell types and compare its performance against previously evaluated SLMs, including Gemma-3-4B, Llama-3.2-3B, and Phi-4-14B. Our results show that Gemini-3-Pro-Preview outperforms the SLMs, while producing actionable explanations that can help practitioners revise manual test cases for greater clarity and consistency. We also find that test smells are pervasive in practice, with nearly one detected test smell per step on average, highlighting the need for scalable and automated quality support for manual testing artifacts.
Downloads
References
Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., et al. (2024). Phi-4 technical report. arXiv preprint arXiv:2412.08905.
Aranda, M., Oliveira, N., Soares, E., Ribeiro, M., Romão, D., Patriota, U., Gheyi, R., Souza, E., and Machado, I. (2024). A catalog of transformations to remove smells from natural language tests. In International Conference on Evaluation and Assessment in Software Engineering, pages 7–16. ACM.
Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., and Baldwin, B. (2025). Non-determinism of “deterministic” LLM settings.
Basili, V. R., Caldiera, G., and Rombach, H. D. (1994). The Goal Question Metric Approach.
Bavota, G., Qusef, A., Oliveto, R., De Lucia, A., and Binkley, D. (2015). Are test smells really harmful? An empirical study. Empirical Software Engineering, 20(4):1052–1094.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems.
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M. I., Gonzalez, J. E., and Stoica, I. (2024). Chatbot arena: an open platform for evaluating LLMs by human preference. In International Conference on Machine Learning, ICML’24. JMLR.org.
DAIR.AI (2024). Prompt Engineering Guide. [link].
Hauptmann, B. (2016). Reducing system testing effort by focusing on commonalities in test procedures. PhD thesis, Technische Universität München.
Hauptmann, B., Junker, M., Eder, S., Heinemann, L., Vaas, R., and Braun, P. (2013). Hunting for smells in natural language tests. In International Conference on Software Engineering, pages 1217–1220. IEEE Computer Society.
Hou, Y., Dong, H., Wang, X., Li, B., and Che, W. (2022). Metaprompting: Learning to learn better prompts. arXiv preprint arXiv:2209.11486.
Juhnke, K., Nikic, A., and Tichy, M. (2021). Clustering natural language test case instructions as input for deriving automotive testing DSLs. J. Object Technol., 20(3):5–1.
Junior, N. S., Martins, L., Rocha, L., Costa, H., and Machado, I. (2021). How are test smells treated in the wild? A tale of two empirical studies. Journal of Software Engineering Research and Development, 9:9–1.
Junior, N. S., Rocha, L., Martins, L. A., and Machado, I. (2020). A survey on test practitioners’ awareness of test smells. In Iberoamerican Conference on Software Engineering, pages 462–475. Curran Associates.
Lucas, K., Gheyi, R., Ribeiro, M., Palomba, F., Martins, L., and Soares, E. (2025). Investigating the performance of small language models in detecting test smells in manual test cases. In Proceedings of the Brazilian Symposium on Software Engineering, pages 783–789, Porto Alegre, RS, Brasil. SBC.
Lucas, K., Gheyi, R., Ribeiro, M., Palomba, F., Martins, L., and Soares, E. (2026). An empirical study of Gemini 3 for detecting natural language test smells in manual test cases (artifacts). https://doi.org/10.5281/zenodo.20549059.
Lucas, K., Gheyi, R., Soares, E., Ribeiro, M., and Machado, I. (2024). Evaluating large language models in detecting test smells. In Brazilian Symposium on Software Engineering, pages 672–678.
Melo, R., Simoes, P., Gheyi, R., d’Amorim, M., Ribeiro, M., Soares, G., Almeida, E., and Soares, E. (2026). Agentic LMs: Hunting Down Test Smells. IEEE Software.
Peixoto, M., Baia, D., Nascimento, N., Alencar, P., Fonseca, B., and Ribeiro, M. (2025). On the Effectiveness of LLMs for Manual Test Verifications. In 2025 IEEE/ACM International Workshop on Deep Learning for Testing and Testing for Deep Learning (DeepTest), pages 45–52, Los Alamitos, CA, USA. IEEE Computer Society.
Rajkovic, K. and Enoiu, E. (2022). NALABS: Detecting bad smells in natural language requirements and test specifications. arXiv preprint arXiv:2202.05641.
Sallou, J., Durieux, T., and Panichella, A. (2024). Breaking the silence: the threats of using LLMs in software engineering. In International Conference on Software Engineering: New Ideas and Emerging Results, pages 102–106. ACM.
Soares, E., Aranda, M., Oliveira, N., Ribeiro, M., Gheyi, R., Souza, E., Machado, I., Santos, A. L. M., Fonseca, B., and Bonifácio, R. (2023a). Manual tests do smell! Cataloging and identifying natural language test smells. In International Symposium on Empirical Software Engineering and Measurement, pages 1–11. IEEE.
Soares, E., Ribeiro, M., Amaral, G., Gheyi, R., Fernandes, L., Garcia, A., Fonseca, B., and Santos, A. (2020). Refactoring test smells: A perspective from open-source developers. In Brazilian Symposium on Systematic and Automated Software Testing, pages 50–59.
Soares, E., Ribeiro, M., Gheyi, R., Amaral, G., and Santos, A. L. M. (2023b). Refactoring test smells with JUnit 5: Why should developers keep up-to-date? IEEE Transactions on Software Engineering, 49(3):1152–1170.
Soares, G., Santos, V., Ribeiro, M., Martins, L., Pontillo, V., III, M. A., Gheyi, R., Machado, I., and Palomba, F. (2025). On the harmfulness of test smells in manual system testing: A controlled experiment. In International Symposium on Empirical Software Engineering and Measurement. IEEE.
spaCy (2026). Industrial-strength natural language processing in Python. [link].
Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ramé, A., Rivière, M., et al. (2025). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
Ubuntu (2026). Manual tests. [link].
Veizaga, A., Shin, S. Y., and Briand, L. C. (2024). Automated smell detection and recommendation in natural language requirements. IEEE Transactions on Software Engineering, 50(4):695–720.
Wu, D., Mu, F., Shi, L., Guo, Z., Liu, K., Zhuang, W., Zhong, Y., and Zhang, L. (2024). iSMELL: Assembling LLMs with Expert Toolsets for Code Smell Detection and Refactoring. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, pages 1345–1357.
Yang, Y., Hu, X., Xia, X., and Yang, X. (2024). The lost world: Characterizing and detecting undiscovered test smells. ACM Transactions on Software Engineering and Methodology, 33(3).
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Keila Lucas, Rohit Gheyi, Márcio Ribeiro, Fabio Palomba, Luana Martins, Elvys Soares

This work is licensed under a Creative Commons Attribution 4.0 International License.

