Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts

Authors

DOI:

https://doi.org/10.5753/jbcs.2026.7717

Keywords:

Narrative Consolidation, Task Formulation, Benchmark Resource, Reference Baselines, Multi-Document Summarization, Temporal Coherence, Temporal Alignment Event Graph

Abstract

Processing overlapping narrative documents, such as legal testimonies or historical accounts, often aims not for compression but for a unified, coherent, and chronologically sound text. Standard Multi-Document Summarization (MDS), with its focus on conciseness, fails to preserve narrative flow. This paper formally defines this challenge as a new NLP task, Narrative Consolidation, whose central objectives are chronological integrity, completeness, and the fusion of complementary details, and establishes the resources needed to study it: a formal task definition, an evaluation paradigm that includes a selection-level metric, the Gospel Consolidation Language Resource --- a benchmark built from the four Biblical Gospels, with 169 canonical events, cross-document alignments, and a manually created reference consolidation --- and a suite of reference systems, ranging from a timeline-agnostic centrality method to timeline-aware heuristics and the Temporal Alignment Event Graph (TAEG), a multi-relational graph that explicitly models chronology and event alignment. Benchmarking these systems yields three findings that characterize the task. First, the explicit temporal backbone is the dominant factor: every system granted the canonical timeline raises ROUGE-L F1 from 0.206 to at least 0.81, whereas content-selection sophistication accounts for a far smaller margin. Second, on a fusion-style reference, a simple length heuristic --- selecting the longest available account of each event --- is remarkably strong (0.947 ROUGE-L F1; 0.648 selection accuracy), outperforming the graph-based selection (0.846; 0.489), which is nevertheless clearly above the random floor. Third, ablations show that the discriminative signal resides in the temporal edges, while intra-cluster lexical similarity is uninformative --- a negative result that constrains the design of future models. Together, these results establish Narrative Consolidation as a task distinct from summarization, provide the first reference points for it, and pose an explicit open challenge: to surpass that heuristic with a principled selection mechanism, and to move beyond version selection towards the fusion of complementary details.

Downloads

Download data is not yet available.

References

Aschmann, R. (2022). Chronology of the four gospels (harmony of the gospels). Available at:[link].

Banerjee, S. and Lavie, A. (2005). METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proc. ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for MT and Summarization, pages 65-72. Available at:[link].

Barzilay, R., McKeown, K. R., and Elhadad, M. (1999). Information fusion in the context of multi-document summarization. In Proc. Annual Meeting of the Assoc. for Computational Linguistics (ACL), pages 550-557. DOI: 10.3115/1034678.1034762.

Biblica, I. (2011). The Holy Bible, New International Version. Book. Copyright © 1973, 1978, 1984, 2011 by Biblica, Inc.™ Used by permission for academic research.

Bollegala, D., Okazaki, N., and Ishizuka, M. (2010). A bottom-up approach to sentence ordering for multi-document summarization. Inf. Process. Manage., 46(1):89-109. DOI: 10.1016/j.ipm.2009.07.004.

Cunha, A. (2025). Semana da paixão unificada: uma harmonização moderna dos evangelhos. In Anais do IV Congr. Bras. de Humanismo Solidário na Ciência. SBCC. Book.

Erkan, G. and Radev, D. R. (2004). Lexrank: Graph-based lexical centrality as salience in text summarization. J. Artif. Intell. Res., 22:457-479. DOI: 10.1613/jair.1523.

Eusebius (1999). The Church History. Kregel Publications. Book.

Finger, R. A., Cortes, E. G., Rigo, S. J., and Ramos, G. d. O. (2026). Neurosymbolic narrative consolidation: Grounding abstractive MDS and LLMs with temporal event graphs. In Proc. IEEE Int. Joint Conf. on Neural Networks (IJCNN). IEEE. Available at:[link].

Jiao, P., Chen, H., Guo, X., Zhao, Z., He, D., and Jin, D. (2025). A survey on temporal interaction graph representation learning: Progress, challenges, and opportunities. DOI: 10.48550/arXiv.2505.04461.

Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81-93. DOI: 10.1093/biomet/30.1-2.81.

Lin, C.-Y. (2004). ROUGE: a package for automatic evaluation of summaries. In Proc. Workshop on Text Summarization Branches Out, pages 74-81. Available at:[link].

Ma, C., Zhang, W. E., Guo, M., Wang, H., and Sheng, Q. Z. (2023). Multi-document summarization via deep learning techniques: A survey. ACM Comput. Surv., 55(5):1-37. DOI: 10.1145/3529754.

Mamidala, K. K. and Sanampudi, S. K. (2021). A novel framework for multi-document temporal summarization (MDTS). Emerg. Sci. J., 5(2):184-190. DOI: 10.28991/esj-2021-01268.

Martschat, S. and Markert, K. (2017). Improving ROUGE for timeline summarization. In Proc. Conf. of the European Chapter of the ACL (EACL), pages 285-290. DOI: 10.18653/v1/E17-2046.

Petersen, W. L. (1994). Tatian's Diatessaron: Its Creation, Dissemination, Significance, and History in Scholarship. Brill. DOI: 10.2307/3266926.

Rossi, E., Chamberlain, B., Frasca, F., Eynard, D., Monti, F., and Bronstein, M. (2020). Temporal graph networks for deep learning on dynamic graphs. DOI: 10.48550/arXiv.2006.10637.

Santana, B., Campos, R., Amorim, E., Jorge, A., Silvano, P., and Nunes, S. (2023). A survey on narrative extraction from textual data. Artif. Intell. Rev., 56(8):8393-8435. DOI: 10.1007/s10462-022-10338-7.

Song, J., Akhter, M. E., Atzil-Slonim, D., and Liakata, M. (2025). Temporal reasoning for timeline summarisation in social media. In Proc. Annual Meeting of the Assoc. for Computational Linguistics (ACL), pages 28085-28101. DOI: 10.18653/v1/2025.acl-long.1362.

Xiao, W., Beltagy, I., Carenini, G., and Cohan, A. (2022). PRIMERA: pyramid-based masked sentence pre-training for multi-document summarization. In Proc. Annual Meeting of the Assoc. for Computational Linguistics (ACL), pages 4016-4036. DOI: 10.18653/v1/2022.acl-long.276.

Yasunaga, M., Zhang, R., Meelu, K., Pareek, A., Srinivasan, K., and Radev, D. (2017). Graph-based neural multi-document summarization. In Proc. Conf. on Computational Natural Language Learning (CoNLL), pages 452-462. DOI: 10.18653/v1/K17-1045.

Yu, Y., Jatowt, A., Doucet, A., Sugiyama, K., and Yoshikawa, M. (2021). Multi-TimeLine summarization (MTLS): Improving timeline summarization by generating multiple summaries. In Proc. Annual Meeting of the Assoc. for Computational Linguistics and Int. Joint Conf. on NLP (ACL-IJCNLP), pages 377-387. DOI: 10.18653/v1/2021.acl-long.32.

Zhang, J., Zhao, Y., Saleh, M., and Liu, P. J. (2020a). PEGASUS: pre-training with extracted gap-sentences for abstractive summarization. In Proc. Int. Conf. on Machine Learning (ICML), pages 11328-11339. PMLR. DOI: 10.48550/arXiv.1912.08777.

Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020b). BERTScore: evaluating text generation with BERT. In Proc. Int. Conf. on Learning Representations (ICLR). DOI: 10.48550/arXiv.1904.09675.

Downloads

Published

2026-08-28

How to Cite

Finger, R. A., Cortes, E. G., Rigo, S. J., & Ramos, G. de O. (2026). Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts. Journal of the Brazilian Computer Society, 32(1), 2203–2215. https://doi.org/10.5753/jbcs.2026.7717

Issue

Section

Regular Issue