Metamorphic Fairness Testing of Retrieval-Augmented Generation: Diagnosing Retriever Bias and Evaluating Graph-based Mitigation

Authors

DOI:

https://doi.org/10.5753/jserd.2026.7713

Keywords:

metamorphic testing, fairness testing, retrieval-augmented generation, graph-based retrieval, demographic bias, small language models

Abstract

Fairness is an under-tested quality attribute in Artificial Intelligence (AI)-enabled software systems. Retrieval-Augmented Generation (RAG) pipelines in production software pose a critical testing challenge: the retriever, as an upstream component, can exhibit demographic sensitivity that propagates defects across downstream stages. Yet most test suites target only relevance and factuality, leaving fairness systematically untested. Applying metamorphic testing (MT) is well-suited here since, as an oracle-free technique, it checks whether semantically neutral input transformations such as demographic perturbations preserve system outputs. This exposes fairness violations without requiring ground-truth labels. This paper presents a two-stage empirical study that applies MT to RAG pipelines as a component-level software testing problem. In the Bias Diagnosis stage, we treat the retriever as a first-class test artifact and apply 21 controlled demographic metamorphic relations across four categories to three Small Language Models (SLMs). In the Mitigation Assessment stage, we perform regression testing to determine whether graph-enhanced retrieval affects relevance and fairness. At this stage, the results show that graph-based retrieval is fairness-neutral. Only the 3B model degrades significantly under graph reranking (p = 0.0005), while larger models remain unaffected (p > 0.14), a model-specific regression effect invisible to aggregate testing. The change in Attack Success Rate (ASR) between flat retrieval and graph reranking is ΔASR = +-0.0008: the architectural change passes relevance regression but leaves the demographic-bias defect unmitigated, neither correcting nor worsening fairness. Together, these test stages yield implications for software testing practice: (i) RAG fairness testing must be component-level, not end-to-end only; (ii) retrieval enhancements such as graph reranking require independent fairness regression testing per model; (iii) discard retrieval stability as a promising lightweight, Graphics Processing Unit (GPU)-free signal for fairness-oriented monitoring in Continuous Integration / Continuous Deployment (CI/CD) workflows; and (iv) fault mitigation should target the embedding stage, since downstream architectural change evaluated here did not affect fairness. We provide a reusable MT test framework and grounded thresholds to help practitioners integrate fairness testing into RAG development workflows.

Downloads

Download data is not yet available.

References

Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., and 2019, T. Z. (2019). Software engineering for machine learning: A case study.

Baeza-Yates, R. (2018). Bias on the web.

Basili, V. R., Caldiera, G., and 1994, H. D. R. (1994). The goal question metric approach.

Bender, E. M., Gebru, T., McMillan-Major, A., and 2021, S. S. (2021). On the dangers of stochastic parrots: Can language models be too big?

Cantini, R., Orsino, A., Ruggiero, M., and Talia, D. (2025). Benchmarking adversarial robustness to bias elicitation in large language models: Scalable automated assessment with llm-as-a-judge. Machine Learning, 114(11):249.

Chen, J., Dong, H., Wang, X., Feng, F., Wang, M., and 2023, X. H. (2023). Bias and fairness in information retrieval systems.

Chen, T. Y., Kuo, F.-C., Liu, H., Poon, P.-L., Towey, D., Tse, T. H., and 2018, Z. Q. Z. (2018). Metamorphic testing: A review of challenges and opportunities.

Chen, Z., Zhang, J. M., Hort, M., Harman, M., and 2024, F. S. (2024). Fairness testing: A comprehensive survey and analysis of trends.

Cruz, A. F., Saleiro, P., Belém, C., Soares, C., and Bizarro, P. (2021). Promoting fairness through hyperparameter optimization. In 2021 IEEE international conference on data mining (ICDM), pages 1036–1041. IEEE.

Es, S., James, J., Anke, L. E., and Schockaert, S. (2024). Ragas: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th conference of the european chapter of the association for computational linguistics: system demonstrations, pages 150–158.

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., and 2023, H. W. (2023). A comprehensive survey of retrieval-augmented generation (rag): Evolution, current landscape and future directions.

Giramata, S., Srinivasan, M., Gudivada, V. N., and Kanewala, U. (2025). Efficient fairness testing in large language models: Prioritizing metamorphic relations for bias detection. In 2025 IEEE International Conference on Artificial Intelligence Testing (AITest), pages 191–200. IEEE.

Hu, M., Wu, H., Guan, Z., Zhu, R., Guo, D., Qi, D., and 2024, S. L. (2024). No free lunch: Retrieval-augmented generation undermines fairness in llms, even for vigilant users.

Hyun, S., Guo, M., and 2024, M. A. B. (2024). METAL: Metamorphic testing framework for analyzing large-language model qualities.

Ji, Z., Yu, T., Xu, Y., Lee, N., Ishii, E., and 2023, P. F. (2023). Towards mitigating LLM hallucination via self reflection.

Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and tau Yih. 2020, W. (2020). Dense passage retrieval for open-domain question answering.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., tau Yih, W., Rocktäschel, T., and et al. 2020 (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks.

Li, M., Zhang, Z., Cao, T., Liu, F., Zhang, C., and 2024, K. M. (2024). Trustworthiness in retrieval-augmented generation systems: A survey.

Li, Y., Du, M., Song, R., Wang, X., and 2023, Y. W. (2023). A survey on fairness in large language models.

Li, Z., Guo, Q., Shao, J., Song, L., Bian, J., Zhang, J., and Wang, R. (2025). Graph neural network enhanced retrieval for question answering of large language models.

Liang, L., Bo, Z., Gui, Z., Zhu, Z., Zhong, L., Zhao, P., Sun, M., Zhang, Z., Zhou, J., Chen, W., et al. (2025). Kag: Boosting llms in professional domains via knowledge augmented generation. In Companion Proceedings of the ACM on Web Conference 2025, pages 334–343.

Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Guo, R., Cheng, H., Klochkov, Y., Taufiq, M. F., and 2023, H. L. (2023). Trustworthy llms: A survey and guideline for evaluating large language models’ alignment.

Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A. (2025). A comprehensive overview of large language models. ACM Transactions on Intelligent Systems and Technology, 16(5):1–72.

Oliveira, M., Silva, J., and Fontão, A. (2025). Fairness testing in retrieval-augmented generation: How small perturbations reveal bias in small language models. In Simpósio Brasileiro de Qualidade de Software (SBQS), pages 595–605. SBC.

Reddy, H., Srinivasan, M., and Kanewala, U. (2025). Metamorphic testing for fairness evaluation in large language models: Identifying intersectional bias in llama and gpt. In 2025 IEEE/ACIS 23rd International Conference on Software Engineering Research, Management and Applications (SERA), pages 239–246. IEEE.

Shrestha, R., Zou, Y., Chen, Q., Li, Z., Xie, Y., and Deng, S. (2024). Fairrag: Fair human generation via fair retrieval augmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11996–12005.

Tukey, J. W. et al. (1977). Exploratory data analysis, volume 2. Addison-Wesley Reading, MA.

Van Nguyen, C., Shen, X., Aponte, R., Xia, Y., Basu, S., Hu, Z., Chen, J., Parmar, M., Kunapuli, S., Barrow, J., et al. (2025). A survey on small language models. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing-Natural Language Processing in the Generative AI Era, pages 807–821.

Wang, C., Wan, Z., Kang, H., Chen, E., Xie, Z., Krishna, T., Janapa Reddi, V., and Du, Y. (2026). Slm-mux: Orchestrating small language models for reasoning. In International Conference on Learning Representations, volume 2026, pages 73697–73723.

Wang, W., Huang, J.-t., Wu, W., Zhang, J., Huang, Y., Li, S., He, P., and Lyu, M. R. (2023). Mttm: Metamorphic testing for textual content moderation software. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pages 2387–2399. IEEE.

Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., and 2012, A. W. (2012). Experimentation in Software Engineering.

Xiong, L., Xiong, C., Li, Y., Tang, K.-F., Liu, J., Bennett, P., Ahmed, J., and 2020, A. O. (2020). Approximate nearest neighbor negative contrastive learning for dense text retrieval.

Yi, P., Liang, L., Zhang, D., Chen, Y., Zhu, J., Liu, X., Tang, K., Chen, J., Lin, H., Qiu, L., et al. (2024). Kgfabric: A scalable knowledge graph warehouse for enterprise data interconnection. Proceedings of the VLDB Endowment, 17(12):3841–3854.

Yu, Yue, Xiong, Wei, Li, Zihan, Kong, Lingpeng, Wang, Qi, Chen, Qinyuan, Wang, and Yang, W. (2024). Rankrag: Unifying context ranking with retrieval-augmented generation in llms.

Zhang, J. M., Harman, M., Ma, L., and 2022, Y. L. (2022). Machine learning testing: Survey, landscapes and horizons.

Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., and et al. 2023 (2023). A survey of large language models.

Zheng, Z., Ren, D., Liu, H., Chen, T. Y., and Li, T. (2025). Identifying the failure-revealing test cases in metamorphic testing: A statistical approach. ACM Transactions on Software Engineering and Methodology, 34(2):1–26.

Downloads

Published

2026-08-04

How to Cite

Oliveira, M., Vergilio, B. J., Sobrinho, R. R., Silva, J., & Fontão, A. de L. (2026). Metamorphic Fairness Testing of Retrieval-Augmented Generation: Diagnosing Retriever Bias and Evaluating Graph-based Mitigation. Journal of Software Engineering Research and Development, 14(1), 179–208. https://doi.org/10.5753/jserd.2026.7713

Issue

Section

Research Article