Constructing Knowledge Graphs from Text Using Large Language Models: Scoping Review
DOI:
https://doi.org/10.5753/reviews.2026.6738Keywords:
Knowledge Graphs, Large Language Models, Scoping ReviewAbstract
Constructing Knowledge Graphs from unstructured textual sources poses significant challenges due to the inherent ambiguity of natural language and the high cost and limited scalability of manual knowledge modeling. Recently, Large Language Models (LLMs) have emerged as a promising alternative for automating knowledge extraction and structuring, resulting in a rapidly expanding body of research. This article aims to provide a scoping review of methods using LLMs to construct Knowledge Graphs from text, map existing approaches, identify methodological patterns, and analyze their strengths and limitations, thereby synthesizing the state of the art. The review employs a systematic protocol consistent with PRISMA guidelines and examines 126 primary studies. The literature is categorized into four methodological groups: Ontology-Based, Prompt-Based, RAG-Based, and Hybrid Pipelines. Each category is analyzed with respect to the role of LLMs within the construction pipeline, the degree of semantic formalization, and the architectural strategies implemented. Comparative analysis is performed using five evaluation criteria: exactness, scalability, adaptability, reproducibility, and ease of implementation. The findings demonstrate that no single approach comprehensively addresses all challenges inherent to LLM-based Knowledge Graph construction. Ontology-Based methods provide robust semantic guarantees but demand considerable manual effort. Prompt-Based approaches facilitate rapid deployment yet exhibit variability and limited reproducibility. RAG-Based methods enhance grounding through external evidence but introduce reliance on retrieval mechanisms. Hybrid Pipelines yield higher extraction accuracy, though at the expense of increased architectural complexity. The review identifies substantial heterogeneity in evaluation practices and a lack of standardized quantitative metrics. These results underscore the necessity for consistent evaluation frameworks to enhance comparability, reliability, and the advancement of research in LLM-assisted Knowledge Graph construction.
Downloads
References
Alammar, J. and Grootendorst, M. (2024). Hands-On Large Language Models. O'Reilly Media. Book.
Berryman, J. and Ziegler, A. (2024). Prompt Engineering for LLMs. O'Reilly Media. Book.
Cao, L., Sun, J., Cross, A., et al. (2024). An automatic and end-to-end system for rare disease knowledge graph construction based on ontology-enhanced large language models: Development study. JMIR Medical Informatics, 12(1):e60665. DOI: 10.2196/60665.
Cauter, Z. and Yakovets, N. (2024). Ontology-guided knowledge graph construction from maintenance short texts. In Proceedings of the 1st Workshop on Knowledge Graphs and Large Language Models (KaLLM 2024), pages 75-84. DOI: 10.18653/v1/2024.kallm-1.8.
Fang, Y., Chen, Y., Jiang, Z., Xiao, J., and Ge, Y. (2024). Effective and reliable domain-specific knowledge question answering. In 2024 IEEE International Conference on e-Business Engineering (ICEBE), pages 238-243. IEEE. DOI: 10.1109/icebe62490.2024.00044.
Gao, T., Zhai, X., Yang, C., Lv, L., and Wang, H. (2024). Joint extraction of entity and relation based on fine-tuning bert for long biomedical literatures. Bioinformatics Advances, 4(1):vbae194. DOI: 10.1093/bioadv/vbae194.
Geng, H., Shi, C., Jiang, X., Kong, Z., and Liu, S. (2024). An entity relation extraction framework based on large language model and multi-tasks iterative prompt engineering. In 2024 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 4763-4769. IEEE. DOI: 10.1109/smc54092.2024.10831494.
Ghanem, H. and Cruz, C. (2024). Fine-tuning vs. prompting: evaluating the knowledge graph construction with llms. In 3rd International Workshop on Knowledge Graph Generation from Text (Text2KG) Co-located with the Extended Semantic Web Conference (ESWC 2024), volume 3747, page 7. DOI: 10.3389/fdata.2025.1505877.
Ji, L., Du, S., Qiu, Y., Xu, H., and Guo, X. (2024). Constructing a medical domain functional knowledge graph with large language models. In 2024 IEEE 6th International Conference on Power, Intelligent Computing and Systems (ICPICS), pages 491-496. IEEE. DOI: 10.1109/icpics62053.2024.10796042.
Jurafsky, D. and Martin, J. H. (2025). Speech and language processing. DOI: 10.4324/9780203461891-3.
Kejriwal, M., Knoblock, C. A., and Szekely, P. (2021). Knowledge graphs: Fundamentals, techniques, and applications. MIT Press. Book.
Li, D. and Xu, F. (2024). The deep integration of knowledge graphs and large language models: Advancements, challenges, and future directions. In 2024 IEEE 2nd International Conference on Sensors, Electronics and Computer Engineering (ICSECE). IEEE. DOI: 10.1109/ICSECE61636.2024.10729340.
Li, H., Xia, C., Hou, Y., Hu, S., Liu, Y., and Jiang, Q. (2024a). Tcmrd-kg: Design and development of a rheumatism knowledge graph based on ancient chinese literature. In Proceedings of the 2024 IEEE International Conference on Medical Artificial Intelligence (MedAI), pages 588-593. IEEE. DOI: 10.1109/MedAI62885.2024.00083.
Li, Z., Wei, Q., Huang, L.-C., Li, J., Hu, Y., Chuang, Y.-S., He, J., Das, A., Keloth, V. K., Yang, Y., et al. (2024b). Ensemble pretrained language models to extract biomedical knowledge from literature. Journal of the American Medical Informatics Association, 31(9):1904-1911. DOI: 10.1093/jamia/ocae061.
Liang, X., Wang, Z., Li, M., and Yan, Z. (2024). A survey of llm-augmented knowledge graph construction and application in complex product design. Procedia CIRP, 128:870-875. DOI: 10.1016/j.procir.2024.07.069.
Mao, Y., He, J., and Chen, C. (2025). From prompts to templates: A systematic prompt template analysis for real-world llmapps. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering (FSE Companion). ACM. DOI: 10.1145/3696630.3728533.
Mihindukulasooriya, N., Tiwari, S. M., Enguix, C. F., and Lata, K. (2023). Text2kgbench: A benchmark for ontology-driven knowledge graph generation from text. arXiv preprint arXiv:2308.02357. DOI: 10.1007/978-3-031-47243-5_14.
Muscolino, H., Machado, A., Vesset, D., and Rydning, J. (2023). Untapped value: What every executive needs to know about unstructured data. White paper, IDC, Framingham, MA. Available at:[link]. Sponsored by Box.
Negro, A., Kus, V., Futia, G., and Montagna, F. (2022). Knowledge Graphs and LLM in Action. Manning. Book.
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., et al. (2021). The prisma 2020 statement: an updated guideline for reporting systematic reviews. bmj, 372. DOI: 10.1136/bmj.n71.
Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., and Wu, X. (2024). Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering, 36(7):3580-3599. DOI: 10.1109/tkde.2024.3352100.
Peters, M. D. J., Godfrey, C. M., McInerney, P., Soares, C. B., Khalil, H., and Parker, D. (2015). The joanna briggs institute reviewers' manual 2015: methodology for jbi scoping reviews. Available at:[link].
Phoenix, J. and Taylor, M. (2024). Prompt Engineering for Generative AI. O'Reilly Media. Book.
Queiroz, J., Jaculli, M. A., Junior, N. C., Silveira, I. M., Mendes, J. R. P., Penteado, B. E., Guilherme, I. R., and Perrout, S. R. (2023). An ontology of well engineering entities to extract and structure text data from daily reports. In Proceedings of the 28th International Conference on Engineering Applications of Artificial Intelligence (EAAI), pages 1-8. Elsevier. Available as preprint. DOI: 10.2118/223801-ms.
Reynolds, L. and McDonell, K. (2021). Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. DOI: 10.1145/3411763.3451760.
Schulhoff, S., Ilie, M., Balepur, N., et al. (2024). The prompt report: A systematic survey of prompting techniques. arXiv preprint arXiv:2406.06608. DOI: 10.48550/arXiv.2406.06608.
Shim, M., Choi, H., Koo, H., Um, K., Lee, K.-H., and Lee, S. (2025). Omega ($ømega$): Ontology-based information extraction framework for constructing task-centric knowledge graph from manufacturing documents with large language model. Advanced Engineering Informatics, 64:103001. DOI: 10.1016/j.aei.2024.103001.
Tian, X., Xu, J., Wang, S., and Zhou, Y. (2025). Construction and application of a multi-modal knowledge graph integrated with large language models in the field of manufacturing processes. In 2025 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), pages 0288-0292. IEEE. DOI: 10.1109/icaiic64266.2025.10920784.
Wang, Q., Li, C., Zhang, Y., Liu, Y., Xiong, H., and Shen, T. (2024). Knowledge graph construction: State-of-the art techniques and future opportunities. In 2024 12th International Conference on Information Systems and Computing Technology (ISC Tech). IEEE. DOI: 10.1109/ISCTech63666.2024.10845648.
Wei, Z. and Esquivel, J. A. (2024). Enhancing knowledge graph completion with retrieval-augmented generation using large language models. In 2024 6th International Conference on Frontier Technologies of Information and Computer (ICFTIC), pages 828-831. IEEE. DOI: 10.1109/icftic64248.2024.10912943.
Zhong, L., Wu, J., Li, Q., Peng, H., and Wu, X. (2023). A comprehensive survey on automatic knowledge graph construction. ACM Computing Surveys, 56(4):1-62. DOI: 10.1145/3618295.
Zhu, Y., Wang, X., Chen, J., Qiao, S., Ou, Y., Yao, Y., Deng, S., Chen, H., and Zhang, N. (2024). Llms for knowledge graph construction and reasoning: Recent capabilities and future opportunities. World Wide Web, 27(5):58. DOI: 10.1007/s11280-024-01297-w.
Zong, H., Wu, R., Cha, J., Feng, W., Wu, E., Li, J., Shao, A., Tao, L., Li, Z., Tang, B., et al. (2024). Advancing chinese biomedical text mining with community challenges. Journal of Biomedical Informatics, page 104716. DOI: 10.1016/j.jbi.2024.104716.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Giovanna Borges Bottino, José de Jesus Pérez Alcazár

This work is licensed under a Creative Commons Attribution 4.0 International License.
