Evaluating Incremental Context Strategy in RAG-Based Chatbots for Hospital Document Retrieval: A Trade-off Between Quality and Cost
DOI:
https://doi.org/10.5753/jis.2026.8133Keywords:
Retrieval-Augmented Generation, Chatbots, Large Language Models, Incremental Context, Token EfficiencyAbstract
Applications in the healthcare field where access to contextualized information anchored in reliable documents has been crucial have increasingly driven the adoption of Retrieval-Augmented Generation (RAG) systems. In this context, this work presents the development and evaluation of an RAG-based chatbot for retrieving normative documents from a university hospital. Traditional RAG approaches typically rely on fixed-size contexts, which may lead to excessive token consumption and degradation in response quality due to the inclusion of irrelevant information. Building upon a previously optimized retrieval pipeline based on the intfloat/multilingual-e5-small embedding model and Reciprocal Rank Fusion (RRF), this study investigates a novel incremental context strategy that progressively expands the input to the language model's during response generation. Experiments were conducted on a hybrid dataset of 192 question-answer pairs, combining synthetic queries and expert-generated questions. Four large language models (Gemini Flash 2.5, GPT-4.1-mini, Sabiá 3.1, and DeepSeek 3.2) were evaluated under fixed and incremental context settings, using BERTScore, LLM-as-a-Judge metrics, and token consumption as evaluation criteria. The results show that the incremental approach significantly reduces token usage, achieving up to 88.68% savings, while maintaining or improving semantic quality in most cases. These findings demonstrate that incremental context construction is an effective strategy for improving efficiency and response quality in RAG-based chatbots, while mitigating context dilution effects. The proposed approach is model-agnostic and readily applicable to real-world scenarios, particularly in resource-sensitive and high-stakes domains such as healthcare.
Downloads
References
Abo El-Enen, M., Saad, S., and Nazmy, T. (2025). A survey on retrieval-augmentation generation (rag) models for healthcare applications. Neural Computing and Applications, 37(33):28191–28267. DOI: https://doi.org/10.1007/s00521-025-11666-9.
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. (2023). Self-rag: Self-reflective retrieval augmented generation. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
Ayala, O. and Bechard, P. (2024). Reducing hallucination in structured outputs via retrieval-augmented generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), pages 228–238. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2024.naacl-industry.19.
Baur, D., Ansorg, J., Heyde, C.-E., and Voelker, A. (2025). Development and evaluation of a retrieval-augmented generation chatbot for orthopedic and trauma surgery patient education: Mixed-methods study. JMIR AI, 4:e75262. DOI: https://doi.org/10.2196/75262.
Brown, T. B. et al. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
Carraro, D. and Bridge, D. (2025). Enhancing recommendation diversity by re-ranking with large language models. ACM Transactions on Recommender Systems, 4(2):1–40. DOI: https://doi.org/10.1145/3700604.
Cederlund, O., Alawadi, S., and Awaysheh, F. M. (2024). Llmrag: An optimized digital support service using llm and retrieval-augmented generation. In Proceedings of the 9th International Conference on Fog and Mobile Edge Computing (FMEC 2024), pages 54–62, Malmö, Sweden. IEEE. DOI: https://doi.org/10.1109/FMEC62297.2024.10710181.
Cormack, G. V., Clarke, C. L., and Buettcher, S. (2009). Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 758–759. ACM. DOI: https://doi.org/10.1145/1571941.1572114.
da Cunha, M. V., Silveira, M. R., Santana, B. S., Freitas, L. A., and Corrêa, U. B. (2025). Optimizing and evaluating a retrieval-augmented generation system for normative document retrieval in hospital settings. In Anais do XXXI Simpósio Brasileiro de Sistemas Multimídia e Web (WebMedia), pages 385–393, Porto Alegre, Brasil. SBC. DOI: https://doi.org/10.5753/webmedia.2025.16029.
Devi, S., Dhar, G., Bharadwaj, C., and M, A. (2024). Retrieval augmented medlm. In Proceedings of the 2024 IEEE Conference on Artificial Intelligence (CAI 2024), pages 1220–1221, Singapore. IEEE. DOI: https://doi.org/10.1109/CAI59869.2024.00217.
Empresa Brasileira de Serviços Hospitalares (EBSERH) (2025). Estrutura administrativa — hu-ufsc. [link]. Accessed on 1 September 2026.
Fan, W., Ding, Y., Ning, L., Wang, S., Li, H., Yin, D., Chua, T.-S., and Li, Q. (2024). A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 6491–6501, Barcelona, Spain. ACM. DOI: https://doi.org/10.1145/3637528.3671470.
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Guo, Q., Wang, M., et al. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997. DOI: https://doi.org/10.48550/arXiv.2312.10997.
Gomez-Cabello, C. A., Prabha, S., Haider, S. A., Genovese, A., Collaco, B. G., Wood, N. G., Bagaria, S., and Forte, A. J. (2025). Comparative evaluation of advanced chunking for retrieval-augmented generation in large language models for clinical decision support. Bioengineering, 12(11):1194. DOI: https://doi.org/10.3390/bioengineering12111194.
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al. (2026). A survey on llm-as-a-judge. The Innovation, 6(9):101253. DOI: https://doi.org/10.1016/j.xinn.2025.101253.
Gummadi, V., Udayaraju, P., Sarabu, V. R., Ravulu, C., Seelam, D. R., and Venkataramana, S. (2024). Enhancing communication and data transmission security in rag using large language models. In Proceedings of the 4th International Conference on Sustainable Expert Systems (ICSES 2024), pages 612–617. IEEE. DOI: https://doi.org/10.1109/ICSES63445.2024.10763024.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023a). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38. DOI: https://doi.org/10.1145/3571730.
Ji, Z., Yu, T., Xu, Y., Lee, N., Ishii, E., and Fung, P. (2023b). Towards mitigating llm hallucination via self reflection. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1827–1843. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/2023.findings-emnlp.123.
Joshi, P., Gupta, A., Kumar, P., and Sisodia, M. (2024). Robust multi model rag pipeline for documents containing text, table & images. In Proceedings of the 3rd International Conference on Applied Artificial Intelligence and Computing (ICAAIC 2024), pages 993–999, Salem, India. IEEE. DOI: https://doi.org/10.1109/ICAAIC60222.2024.10574972.
Kalyan, K. S. and Sangeetha, S. (2020). Secnlp: A survey of embeddings in clinical natural language processing. Journal of Biomedical Informatics, 101:103323. DOI: https://doi.org/10.1016/j.jbi.2019.103323.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc.
Maryamah, M., Irfani, M. M., Tri Raharjo, E. B., Rahmi, N. A., Ghani, M., and Raharjana, I. K. (2024). Chatbots in academia: A retrieval-augmented generation approach for improved efficient information access. In Proceedings of the 16th International Conference on Knowledge and Smart Technology (KST 2024), pages 259–264, Krabi, Thailand. IEEE. DOI: https://doi.org/10.1109/KST61284.2024.10499652.
Medeiros, L. S. F. and de Oliveira, H. T. A. (2025). Comparação de modelos de embeddings e llms para geração aumentada por recuperação em português. In Anais do LII Seminário Integrado de Software e Hardware (SEMISH), pages 429–440. SBC. DOI: https://doi.org/10.5753/semish.2025.9027.
Nai, R., Sulis, E., Fatima, I., and Meo, R. (2024). Large language models and recommendation systems: A proof-of-concept study on public procurements. In Natural Language Processing and Information Systems (NLDB 2024), Part II, pages 280–290, Turin, Italy. Springer. DOI: https://doi.org/10.1007/978-3-031-70242-6_27.
Obaid, S. and Bawany, N. Z. (2024). Seerahgpt: Retrieval augmented generation based large language model. In Proceedings of the 18th International Conference on Open Source Systems and Technologies (ICOSST 2024), pages 1–7, Lahore, Pakistan. IEEE. DOI: https://doi.org/10.1109/ICOSST64562.2024.10871159.
Oliveira, S. S. T., Fazzioni, D., and Ferreira, D. O. C. (2025). Grandes modelos de linguagem. In Kudo, T. N. et al. (Org.). Cegraf UFG, Goiânia. E-book (254 p.). ISBN 978-85-495-1096-9. Accessed on 1 September 2026.
Perković, G., Drobnjak, A., and Botički, I. (2024). Hallucinations in llms: Understanding and addressing challenges. In Proceedings of the 47th MIPRO ICT and Electronics Convention (MIPRO 2024), pages 2084–2088, Opatija, Croatia. IEEE. DOI: https://doi.org/10.1109/MIPRO60963.2024.10569238.
Puthenputhussery, A., Kang, C., Magnani, A., Zhang, T., Shang, H., Yadav, N., Chandran, P., Madhani, B., Fu, Y.-T., Wang, H., et al. (2025). Large scale deployment of bert based cross encoder model for re-ranking in walmart search engine. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 4365–4369. ACM. DOI: https://doi.org/10.1145/3726302.3731965.
Ramos, J. et al. (2003). Using tf-idf to determine word relevance in document queries. In Proceedings of the First Instructional Conference on Machine Learning, volume 242, pages 29–48, New Jersey, USA.
Raschka, S. (2024). Build a Large Language Model (From Scratch). Manning Publications.
Raza, M., Jahangir, Z., Riaz, M. B., Saeed, M. J., and Sattar, M. A. (2025). Industrial applications of large language models. Scientific Reports, 15:13755. DOI: https://doi.org/10.1038/s41598-025-98483-1.
Reimers, N. and Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992. Association for Computational Linguistics. DOI: https://doi.org/10.18653/v1/D19-1410.
Son, N., Kang, I., Kim, I., Lee, K., Nam, S., and Lee, D. (2025). Development and evaluation of a retrieval-augmented generation-based electronic medical record chatbot system. Healthcare Informatics Research, 31(3):218–225. DOI: https://doi.org/10.4258/hir.2025.31.3.218.
Taguchi, C., Maekawa, S., and Bhutani, N. (2025). Efficient context selection for long-context qa: No tuning, no iteration, just adaptive-k. arXiv preprint arXiv:2506.08479. DOI: https://doi.org/10.48550/arXiv.2506.08479.
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research. Survey Certification.
Wijaya, O. C. and Purwarianti, A. (2024). An interactive question-answering system using large language model and retrieval-augmented generation in an intelligent tutoring system on the programming domain. In Proceedings of the 2024 11th International Conference on Advanced Informatics: Concept, Theory and Application (ICAICTA), pages 1–6, Singapore. IEEE. DOI: https://doi.org/10.1109/ICAICTA63815.2024.10763263.
Yu, H., Gan, A., Zhang, K., Tong, S., Liu, Q., and Liu, Z. (2025). Evaluation of retrieval-augmented generation: A survey. In Zhu, W., Xiong, H., Cheng, X., Cui, L., Dou, Z., Dong, J., Pang, S., Wang, L., Kong, L., and Chen, Z., editors, Big Data (BigData 2024), Communications in Computer and Information Science, pages 102–120, Singapore. Springer. DOI: https://doi.org/10.1007/978-981-96-1024-2_8.
Zhang, D., Du, H., Wang, X., Zhu, M., Pang, X., Wei, D., and Wang, X. (2025). Cmedragbot: A chinese medical chatbot based on graph rag and large language models. Interdisciplinary Sciences: Computational Life Sciences. DOI: https://doi.org/10.1007/s12539-025-00715-5.
Zhao, W. X. et al. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223. DOI: https://doi.org/10.48550/arXiv.2303.18223.
Zhou, Y., Dai, S., Cao, Z., Zhang, X., and Xu, J. (2024). Length-induced embedding collapse in transformer-based models. arXiv preprint arXiv:2410.24200. DOI: https://doi.org/10.48550/arXiv.2410.24200.
Zhu, Z., Wang, Y., Liu, H., Chen, X., and Zhang, M. (2024). Enhancing large language models with knowledge graphs for robust question answering. In Proceedings of the 30th IEEE International Conference on Parallel and Distributed Systems (ICPADS 2024), pages 262–269, Belgrade, Serbia. IEEE. DOI: https://doi.org/10.1109/ICPADS63350.2024.00042.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Murilo Vargas da Cunha, Marilia Rosa Silveira, Brenda Salenave Santana, Larissa Astrogildo Freitas, Ulisses Brisolara Corrêa

This work is licensed under a Creative Commons Attribution 4.0 International License.
JIS is free of charge for authors and readers, and all papers published by JIS follow the Creative Commons Attribution 4.0 International (CC BY 4.0) license.


