Revisiting the Energy Footprint of Small Language Models for Unit Test Generation: A Bayesian Comparative Study

Authors

DOI:

https://doi.org/10.5753/jserd.2026.7689

Keywords:

Unit test generation, Energy consumption, Language model, Sustainable Software Engineering

Abstract

Context. Manual unit test creation is a cognitively intensive and time-consuming activity, prompting researchers and practitioners to increasingly adopt automated testing tools. Recent advancements in language models have expanded automation possibilities, including unit test generation, yet these models raise substantial sustainability concerns due to their energy consumption compared to conventional, specialized tools.
Goal. Our research investigates whether the energy overhead associated with employing small language models (SLMs) for unit test generation is justified compared to a conventional, lightweight testing tool. We compare and analyze the energy consumption incurred during test suite generation, as well as the fault-finding effectiveness of the resulting test suites, for two SLMs (Phi-3.1 Mini 128k and Qwen2.5-Coder 1.5B) and Pynguin, a purpose-built tool for unit test generation.
Method. We posed three research questions: (i) Is the additional energy cost incurred by employing SLMs for unit test generation justified when compared to a lightweight, purpose-built alternative?; (ii) To what extent does the energy cost associated with SLM-based unit test generation vary across different SLM architectures?; and (iii) To what extent do unit test suites generated by Phi, Qwen, and Pynguin differ in their fault-finding effectiveness across a range of faulty programs? To rigorously address the first two research questions, we employed Bayesian Data Analysis. For the third research question, we conducted a complementary empirical analysis using descriptive statistics.
Results. The results from our Bayesian analysis provide robust evidence indicating that Phi and Qwen consistently consume significantly more energy than Pynguin during test suite generation.
Conclusions. These findings underscore significant sustainability concerns associated with employing even SLMs for routine Software Engineering tasks such as unit test generation. The results challenge the assumption of universal energy efficiency benefits from smaller-scale models and emphasize the necessity for careful energy consumption evaluations in the adoption of automated software testing approaches.

Downloads

Download data is not yet available.

References

Abdullin, A., Derakhshanfar, P., and Panichella, A. (2025). Test wars: A comparative study of SBST, symbolic execution, and LLM-based approaches to unit test generation. In IEEE Conference on Software Testing, Verification and Validation (ICST), pages 221–232. New York, NY, USA: ACM.

Abril-Pla, O., Andreani, V., Carroll, C., Dong, L., Fonnesbeck, C. J., Kochurov, M., Kumar, R., Lao, J., Luhmann, C. C., Martin, O. A., Osthege, M., Vieira, R., Wiecki, T., and Zinkov, R. (2023). PyMC: A modern and comprehensive probabilistic programming framework in Python. PeerJ Computer Science, 9:e1516.

Ali, S., Briand, L. C., Hemmati, H., and Panesar-Walawege, R. K. (2010). A systematic review of the application and empirical investigation of search-based test case generation. IEEE Transactions on Software Engineering, 36(6):742–762.

Anand, S., Burke, E. K., Chen, T. Y., Clark, J., Cohen, M. B., Grieskamp, W., Harman, M., Harrold, M. J., and McMinn, P. (2013). An orchestrated survey of methodologies for automated software test case generation. Journal of Systems and Software, 86(8):1978–2001.

Anthropic (2025). Claude. [link]. Large language model.

Boehm, B., Rombach, H., and Zelkowitz, M. (2005). Foundations of Empirical Software Engineering: The Legacy of Victor R. Basili. Springer.

Castaño, J., Martínez-Fernández, S., Franch, X., and Bogner, J. (2023). Exploring the carbon footprint of Hugging Face's ML models: A repository mining study. In ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), pages 1–12. New York, NY, USA: ACM.

Dienes, Z. and McLatchie, N. (2018). Four reasons to prefer Bayesian analyses over significance testing. Psychonomic Bulletin & Review, 25(1):207–218.

Ding, Y. and Shi, T. (2024). Sustainable LLM serving: Environmental implications, challenges, and opportunities. In IEEE 15th International Green and Sustainable Computing Conference (IGSC), pages 37–38. IEEE.

Durelli, R., Endo, A., and Durelli, V. (2025). On the energy footprint of using a small language model for unit test generation. In Anais do X Simpósio Brasileiro de Testes de Software Sistemático e Automatizado, pages 55–64. Porto Alegre, RS, Brasil: SBC.

Fontes, A., Gay, G., de Oliveira Neto, F. G., and Feldt, R. (2023). Automated support for unit test generation. In Romero, J. R., Medina-Bulo, I., and Chicano, F., editors, Optimising the Software Development Process with Artificial Intelligence, Natural Computing Series, pages 179–219. Springer.

Fraser, G. and Arcuri, A. (2013). Whole test suite generation. IEEE Transactions on Software Engineering, 39(2):276–291.

Furia, C. A., Feldt, R., and Torkar, R. (2021). Bayesian data analysis in empirical software engineering research. IEEE Transactions on Software Engineering, 47(9):1786–1810.

Furia, C. A., Torkar, R., and Feldt, R. (2022). Applying Bayesian analysis guidelines to empirical software engineering data: The case of programming languages and code quality. ACM Transactions on Software Engineering and Methodology (TOSEM), 31(3).

Gelman, A. and Loken, E. (2014). The statistical crisis in science: Data-dependent analysis—a "garden of forking paths"—explains why many statistically significant comparisons don’t hold up. American Scientist, 102(6):460–466.

Gigerenzer, G. (2004). Mindless Statistics. The Journal of Socio-Economics, 33(5):587–606.

Godefroid, P., Klarlund, N., and Sen, K. (2005). DART: Directed automated random testing. ACM SIGPLAN Notices, 40(6):213–223.

Jovanović, M. and Campbell, M. (2024). Compacting AI: In search of the small language model. Computer, 57(8):96–100.

Kifetew, F., Prandi, D., and Susi, A. (2025). On the energy consumption of test generation. In IEEE Conference on Software Testing, Verification and Validation (ICST), pages 360–370.

Kruschke, J. K. (2013). Bayesian estimation supersedes the t test. Journal of Experimental Psychology: General, 142(2):573–603.

Kruschke, J. K. (2015). Doing Bayesian Data Analysis: A Tutorial with R, JAGS, and Stan. 2nd revised edition. Academic Press.

Lakhotia, K., McMinn, P., and Harman, M. (2010). An empirical investigation into branch coverage for C programs using CUTE and Austin. Journal of Systems and Software, 83(12):2379–2391.

Lin, D., Koppel, J., Chen, A., and Solar-Lezama, A. (2017). QuixBugs: A multi-lingual program repair benchmark set based on the Quixey challenge. In Proceedings Companion of the ACM SIGPLAN International Conference on Systems, Programming, Languages, and Applications: Software for Humanity (SPLASH), pages 55–56. New York, NY, USA: ACM.

Luccioni, A. S., Viguier, S., and Ligozat, A.-L. (2023). Estimating the carbon footprint of BLOOM, a 176B parameter language model. The Journal of Machine Learning Research, 24(1).

Lukasczyk, S. and Fraser, G. (2022). Pynguin: Automated unit test generation for Python. In IEEE/ACM 44th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion), pages 168–172.

Lukasczyk, S., Kroiß, F., and Fraser, G. (2023). An empirical study of automated unit test generation for Python. Empirical Software Engineering, 28(2):36.

McMinn, P. (2004). Search-based software test data generation: A survey. Software Testing, Verification & Reliability, 14(2):105–156.

Menzies, T. and Shepperd, M. (2019). "Bad smells" in software analytics papers. Information and Software Technology, 112:35–47.

Myers, G. J., Sandler, C., and Badgett, T. (2011). The Art of Software Testing. 3rd edition. Wiley.

Nickerson, R. S. (2000). Null hypothesis significance testing: A review of an old and continuing controversy. Psychological Methods, 5(2):241–301.

Noureddine, A. (2022). PowerJoular and JoularJX: Multi-platform software power monitoring tools. In 18th International Conference on Intelligent Environments, Biarritz, France.

OpenAI (2025). ChatGPT. [link]. Large language model.

Schäfer, M., Nadi, S., Eghbali, A., and Tip, F. (2024). An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering, 50(1):85–105.

van de Schoot, R., Depaoli, S., King, R., Kramer, B., Märtens, K., Tadesse, M. G., Vannucci, M., Gelman, A., Veen, D., Willemsen, J., and Yau, C. (2021). Bayesian statistics and modelling. Nature Reviews Methods Primers, 1.

VanderStoep, S. W. and Johnson, D. D. (2008). Research Methods for Everyday Life: Blending Qualitative and Quantitative Approaches. Jossey-Bass.

Wang, J., Huang, Y., Chen, C., Liu, Z., Wang, S., and Wang, Q. (2024). Software testing with large language models: Survey, landscape, and vision. IEEE Transactions on Software Engineering, 50(4):911–936.

Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., and Wesslén, A. (2012). Experimentation in Software Engineering. Springer.

Ye, H., Martinez, M., Durieux, T., and Monperrus, M. (2021). A comprehensive study of automatic program repair on the QuixBugs benchmark. Journal of Systems and Software, 171.

Downloads

Published

2026-08-27

How to Cite

Durelli, V. H. S., Durelli, R. S., & Endo, A. T. (2026). Revisiting the Energy Footprint of Small Language Models for Unit Test Generation: A Bayesian Comparative Study. Journal of Software Engineering Research and Development, 14(1), 482–497. https://doi.org/10.5753/jserd.2026.7689

Issue

Section

Research Article

Most read articles by the same author(s)