Assessing the Effectiveness of Large Language Models in Detecting Semantic Conflicts
DOI:
https://doi.org/10.5753/jserd.2026.7565Abstract
Semantic conflicts occur when a developer introduces changes to a codebase that unintentionally affect the behavior of changes integrated in parallel by other developers. Since merge tools used in practice cannot detect this type of conflict, complementary tools have been proposed, such as SAM (SemAntic Merge), based on the SMAT approach and relies on the generation and execution of unit tests in Java. Despite showing good conflict detection capabilities, SAM presents a high rate of false negatives (existing conflicts not signaled by it). Part of this problem is due to the natural limitations of unit test generation tools, specifically Randoop and EvoSuite. To understand if these limitations can be overcome by large language models (LLMs), this work proposes, and integrates into SMAT, LUCIA (LLM-based Unit-tests for Conflict Identification & Analysis), a new test generation tool based on the Ollama framework to interface with various local LLMs, including the Llama, Gemma, DeepSeek, and Qwen families. We then explore these models' capability to generate tests, using different interaction strategies, prompts with different contents, and different model parameter configurations. We evaluate the results with two distinct samples: a benchmark with simpler systems, used in related work, and a more significant sample based on complex systems used in practice. Finally, we evaluate the effectiveness of LUCIA in detecting conflicts, comparing the selected LLMs with each other and against the original test generation tools used by SAM: Randoop, Randoop Clean, EvoSuite, and Differential EvoSuite. Our evaluation shows that LLMs effectively detect semantic conflicts, with Llama and Qwen outperforming the other models. While no single model could surpass traditional tools, multiple models combined achieved superior performance in conflict detection. Overall, the new extension identified three additional conflicts undetected by SAM, representing a 200% improvement in identifying unique conflicts compared to the single-model approach (Code Llama 70B) used in our previous work. These results reinforce previous findings that LLMs can generate unit tests that are effective in detecting semantic conflicts, while demonstrating the superior effectiveness of multi-model approaches.
Downloads
References
Barbosa, N. (2025). SMAT with LUCIA tool integration. [link].
Barbosa, N., Borba, P., and Da Silva, L. (2026). Supplementary data for: Assessing the effectiveness of large language models in detecting semantic conflicts. https://doi.org/10.5281/zenodo.20938400.
Barbosa, N., Borba, P., and Silva, L. (2025). Detecção de conflitos semânticos com testes gerados por LLM. In Anais do X Simpósio Brasileiro de Testes de Software Sistemático e Automatizado, pages 18–27, Porto Alegre, RS, Brasil. SBC.
Brun, Y., Holmes, R., Ernst, M. D., and Notkin, D. (2011). Crystal: Precise and unobtrusive conflict warnings. In Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering, ESEC/FSE ’11, pages 444–447, New York, NY, USA. Association for Computing Machinery.
Da Silva, L., Borba, P., Maciel, T., Mahmood, W., Berger, T., Moisakis, J., Gomes, A., and Leite, V. (2024). Detecting semantic conflicts with unit tests. Journal of Systems and Software, 214:112070.
Da Silva, L., Borba, P., and Pires, A. (2022). Build conflicts in the wild. Journal of Software: Evolution and Process, 34(4):e2441.
De Jesus, G. S., Borba, P., Bonifácio, R., and De Oliveira, M. B. (2024). Lightweight semantic conflict detection with static analysis. In Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, ICSE-Companion ’24, pages 343–345, New York, NY, USA. Association for Computing Machinery.
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N. (2020). Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping.
Eclipse Foundation (2020). Eclipse Cargo Tracker: Applied Domain-Driven Design Blueprints for Jakarta EE. [link].
Fraser, G. and Arcuri, A. (2011). EvoSuite: Automatic test suite generation for object-oriented software. In Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering, pages 416–419.
Guo, D., Yang, D., Zhang, H., et al. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633–638.
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. (2020). The curious case of neural text degeneration. In 8th International Conference on Learning Representations (ICLR 2020), Addis Ababa, Ethiopia, April 26–30, 2020. OpenReview.net.
Jain, K. and Le Goues, C. (2025). TestForge: Feedback-driven, agentic test suite generation.
Just, R., Jalali, D., and Ernst, M. D. (2014). Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 International Symposium on Software Testing and Analysis, ISSTA 2014, pages 437–440, New York, NY, USA. Association for Computing Machinery.
Maciel, T., Borba, P., Silva, L., and Burity, T. (2024). Explorando a detecção de conflitos semânticos nas integrações de código em múltiplos métodos. In Anais do XXXVIII Simpósio Brasileiro de Engenharia de Software, pages 181–191, Porto Alegre, RS, Brasil. SBC.
Maddila, C., Nagappan, N., Bird, C., Gousios, G., and van Deursen, A. (2021). CONE: A concurrent edit detection tool for large-scale software development. ACM Transactions on Software Engineering and Methodology, 31(2).
Meta AI (2025). The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. [link].
Moraes, A., Borba, P., and Silva, L. (2024). Semantic conflict detection via dynamic analysis. In Anais do XXVIII Simpósio Brasileiro de Linguagens de Programação, pages 53–61, Porto Alegre, RS, Brasil. SBC.
Ollama (2023). Ollama. [link].
OpenLiberty (2018). DayTrader8 Sample. [link].
Pacheco, C. and Ernst, M. D. (2007). Randoop: Feedback-directed random testing for Java. In Companion to the 22nd ACM SIGPLAN Conference on Object-Oriented Programming Systems and Applications Companion, pages 815–816.
An, R., Kim, M., Krishna, R., Pavuluri, R., and Sinha, S. (2025). ASTER: Natural and multi-language unit test generation with LLMs. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pages 413–424, Los Alamitos, CA, USA. IEEE Computer Society.
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Sauvestre, R., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G. (2024). Code Llama: Open foundation models for code.
Sarma, A., Redmiles, D. F., and van der Hoek, A. (2012). Palantir: Early detection of development conflicts arising from parallel code changes. IEEE Transactions on Software Engineering, 38(4):889–908.
Schäfer, M., Nadi, S., Eghbali, A., and Tip, F. (2024). An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering, 50(1):85–105.
Silva, E., Coelho, R., and Silva, L. (2025). LLMs as test generators: A comparative benchmarking study. In Anais do XXXIX Simpósio Brasileiro de Engenharia de Software, pages 25–36, Porto Alegre, RS, Brasil. SBC.
Silva, L. D., Borba, P., Mahmood, W., Berger, T., and Moisakis, J. (2020). Detecting semantic conflicts via automated behavior change detection. In 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), pages 174–184.
Sousa, M., Dillig, I., and Lahiri, S. K. (2018). Verified three-way program merge. Proceedings of the ACM on Programming Languages, 2(OOPSLA).
Team, G., Kamath, A., Ferret, J., et al. (2025). Gemma 3 Technical Report.
Tree-sitter (2024). Tree-sitter: A parser generator tool and incremental parsing library. [link].
Wang, Z., Liu, K., Li, G., and Jin, Z. (2024). HiTS: High-coverage LLM-based unit test generation via method slicing. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, ASE ’24, pages 1258–1268, New York, NY, USA. Association for Computing Machinery.
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with ChatGPT. In Proceedings of the 30th Conference on Pattern Languages of Programs, PLoP ’23, USA. The Hillside Group.
Yang, A., Li, A., Yang, B., et al. (2025). Qwen3 Technical Report.
Yuan, Z., Liu, M., Ding, S., Wang, K., Chen, Y., Peng, X., and Lou, Y. (2024). Evaluating and improving ChatGPT for unit test generation. Proceedings of the ACM on Software Engineering, 1(FSE).
Zhang, J., Kaufman, M., Mytkowicz, T., Piskac, R., and Lahiri, S. (2022). Using pre-trained language models to resolve textual and semantic merge conflicts (experience paper). In ISSTA 2022: Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis. ACM.
Zhang, Y., Lu, Q., Liu, K., Dou, W., Zhu, J., Qian, L., Zhang, C., Lin, Z., and Wei, J. (2026). CityWalk: Enhancing LLM-based C++ unit test generation via project-dependency awareness and language-specific knowledge. ACM Transactions on Software Engineering and Methodology, 35(5).
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nathalia Barbosa, Paulo Borba, Léuson Da Silva

This work is licensed under a Creative Commons Attribution 4.0 International License.

