Aggregating Neighbor Embedding Projection and Rank-Based Manifold Learning for Image Retrieval
DOI:
https://doi.org/10.5753/jbcs.2026.5685Keywords:
Image Retrieval, Manifold Learning, Rank Aggregation, Projection, Re-RankingAbstract
Content-based image retrieval (CBIR) has evolved significantly with the advent of deep learning models, yet effectively ranking similar images remains a challenging task, particularly in high-dimensional feature spaces where pairwise distance measures often fail to capture complex contextual relationships and the semantic gap between visual features and high-level concepts persists. In this scenario, manifold learning and rank-based refinement methods have emerged as complementary strategies, respectively improving feature representations and exploiting contextual information embedded in ranked lists, such as neighborhood relationships among images. However, combining projection-based and rank-based strategies to exploit their complementary properties remains a challenging research problem. To address this, this paper proposes a framework that combines neighbor embedding projections with rank-based manifold learning through rank aggregation. Specifically, Uniform Manifold Approximation and Projection (UMAP) is used to generate alternative low-dimensional feature representations, while ranked lists obtained from UMAP projections and rank-based re-ranking methods are combined using the Borda Count aggregation strategy. The experimental evaluation was conducted on several public datasets using deep learning features extracted from ResNet152, Swin Transformer, and DINOv2 models. The results show that the proposed approach can improve retrieval effectiveness in several scenarios, particularly when the baseline representation struggles to achieve high precision values. In addition, the aggregation strategy often improves the quality of the top-ranked positions, leading to competitive Mean Average Precision (MAP) and Precision values across different datasets and feature extractors. These findings suggest that combining projection-based and rank-based manifold learning strategies through rank aggregation can provide complementary contextual information for image retrieval tasks.
Downloads
References
Allaoui, M., Kherfi, M. L., and Cheriet, A. (2020). Considerably improving clustering algorithms using umap dimensionality reduction technique: a comparative study. In International conference on image and signal processing, pages 317-325. Springer. DOI: 10.1007/978-3-030-51935-3_34.
Bai, S., Tang, P., Torr, P. H., and Latecki, L. J. (2019). Re-ranking via metric fusion for object retrieval and person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 740-749. DOI: 10.1109/CVPR.2019.00083.
Barz, B. and Denzler, J. (2021). Content-based image retrieval and the semantic gap in the deep learning era. In International Conference on Pattern Recognition, pages 245-260. Springer. DOI: 10.1007/978-3-030-68790-8_20.
Becht, E., McInnes, L., Healy, J., et al. (2019). Dimensionality reduction for visualizing single-cell data using umap. Nature Biotechnology, 37:38-44. DOI: 10.1038/nbt.4314.
Belalia, A., Belloulata, K., and Redaoui, A. (2025). Enhanced image retrieval using multiscale deep feature fusion in supervised hashing. Journal of Imaging, 11(1):20. DOI: 10.3390/jimaging11010020.
Belkin, M. and Niyogi, P. (2001). Laplacian eigenmaps and spectral techniques for embedding and clustering. Advances in neural information processing systems, 14. DOI: 10.7551/mitpress/1120.003.0080.
Chen, Y., Wang, J. Z., and Krovetz, R. (2003). An unsupervised learning approach to content-based image retrieval. In Seventh International Symposium on Signal Processing and Its Applications, 2003. Proceedings., volume 1, pages 197-200. IEEE. DOI: 10.1109/ISSPA.2003.1224674.
Coifman, R. R. and Lafon, S. (2006). Diffusion maps. Applied and computational harmonic analysis, 21(1):5-30. DOI: 10.1016/j.acha.2006.04.006.
Cormack, G. V., Clarke, C. L., and Buettcher, S. (2009). Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pages 758-759. DOI: 10.1145/1571941.1572114.
Diaz-Papkovich, A., Anderson-Trocmé, L., Ben-Eghan, C., and Gravel, S. (2019). Umap reveals cryptic population structure and phenotype heterogeneity in large genomic cohorts. PLoS genetics, 15(11):e1008432. DOI: 10.1371/journal.pgen.1008432.
Donoser, M. and Bischof, H. (2013). Diffusion processes for retrieval revisited. In 2013 IEEE Conference on Computer Vision and Pattern Recognition, pages 1320-1327. DOI: 10.1109/CVPR.2013.174.
Dorrity, M. W., Saunders, L. M., Queitsch, C., Fields, S., and Trapnell, C. (2020). Dimensionality reduction by umap to visualize physical and genetic interactions. Nature communications, 11(1):1537. DOI: 10.1038/s41467-020-15351-4.
Emerson, P. (2013). The original borda count and partial voting. Social Choice and Welfare, 40(2):353-358. DOI: 10.1007/s00355-011-0603-9.
Espadoto, M., Martins, R. M., Kerren, A., Hirata, N. S. T., and Telea, A. C. (2021). Toward a quantitative survey of dimension reduction techniques. IEEE Transactions on Visualization and Computer Graphics, 27(3):2153-2173. DOI: 10.1109/TVCG.2019.2944182.
He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. In CVPR, pages 770-778. DOI: 10.1109/CVPR.2013.174.
Hinton, G. E. and Roweis, S. (2002). Stochastic neighbor embedding. Advances in neural information processing systems, 15. Available at:[link].
Jiang, J., Wang, B., and Tu, Z. (2011). Unsupervised metric learning by self-smoothing operator. In 2011 International Conference on Computer Vision, pages 794-801. IEEE. DOI: 10.1109/ICCV.2011.6126318.
Jing, P., Su, Y., Xu, C., and Zhang, L. (2018). Hyperssr: A hypergraph based semi-supervised ranking method for visual search reranking. Neurocomputing, 274:50-57. DOI: 10.1016/j.neucom.2016.05.085.
Kawai, V. A. S., Leticio, G. R., and Pedronette, D. C. G. (2025). Semi-supervised image retrieval through particle competition and cooperation combined with manifold learning. In 2025 38th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), pages 1-6. IEEE. DOI: 10.1109/SIBGRAPI67909.2025.11223376.
Kawai, V. A. S., Leticio, G. R., Valem, L. P., and Pedronette, D. C. G. (2024). Neighbor embedding projection and rank-based manifold learning for image retrieval. In 2024 37th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), pages 1-6. DOI: 10.1109/SIBGRAPI62404.2024.10716269.
Khosla, A., Jayadevaprakash, N., Yao, B., and Fei-Fei, L. (2011). Novel dataset for fine-grained image categorization. In Workshop on Fine-Grained Visual Categorization, CVPR. Available at:[link].
Kibriya, A. M. and Frank, E. (2007). An empirical comparison of exact nearest neighbour algorithms. In 11th European Conference on Principles and Practice of Knowledge Discovery in Databases, ECMLPKDD'07, page 140–151. DOI: 10.1007/978-3-540-74976-9_16.
Kouiroukidis, N. and Evangelidis, G. (2011). The effects of dimensionality curse in high dimensional knn search. In 2011 15th Panhellenic Conference on Informatics, pages 41-45. IEEE. DOI: 10.1109/PCI.2011.45.
Lee, J. H. (1995). Combining multiple evidence from different properties of weighting schemes. In Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval, pages 180-188. DOI: 10.1145/215206.215358.
Leticio, G. R., Kawai, V. A. S., Valem, L. P., Pedronette, D. C. G., and Torres, R. d. S. (2024). Manifold information through neighbor embedding projection for image retrieval. Pattern Recognition Letters, 183:17-25. DOI: 10.1016/j.patrec.2024.04.022.
Leticio, G. R., Valem, L. P., Lopes, L. T., and Pedronette, D. C. G. (2023). pyudlf: A python framework for unsupervised distance learning tasks. In Proceedings of the 31st ACM International Conference on Multimedia, MM '23, page 9680–9684, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3581783.3613466.
Liu, G.-H. and Yang, J.-Y. (2013). Content-based image retrieval using color difference histogram. Pattern Recognition, 46(1):188 - 198. DOI: 10.1016/j.patcog.2012.06.001.
Liu, Y.-T., Liu, T.-Y., Qin, T., Ma, Z.-M., and Li, H. (2007). Supervised rank aggregation. In Proceedings of the 16th international conference on World Wide Web, pages 481-490. DOI: 10.1145/1242572.1242638.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. ICCV. DOI: 10.1109/ICCV48922.2021.00986.
Long, F., Zhang, H., and Feng, D. D. (2003). Fundamentals of content-based image retrieval. In Multimedia Information Retrieval and Management: Technological Fundamentals and Applications, pages 1-26. Springer. DOI: 10.1007/978-3-662-05300-3_1.
McInnes, L., Healy, J., and Melville, J. (2018). Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. DOI: 10.21105/joss.00861.
Milošević, D., Medeiros, A. S., Piperac, M. S., Cvijanović,, D., Soininen, J., Milosavljević,, A., and Predić, B. (2022). The application of uniform manifold approximation and projection (umap) for unconstrained ordination and classification of biological indicators in aquatic ecology. Science of The Total Environment, 815:152365. DOI: 10.1016/j.scitotenv.2021.152365.
Müller, H., Michoux, N., Bandon, D., and Geissbuhler, A. (2004). A review of content-based image retrieval systems in medical applications—clinical benefits and future directions. International journal of medical informatics, 73(1):1-23. DOI: 10.1016/j.ijmedinf.2003.11.024.
Nilsback, M.-E. and Zisserman, A. (2006). A visual vocabulary for flower classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 2, pages 1447-1454. DOI: 10.1109/CVPR.2006.42.
Oquab, M., Darcet, T., Moutakanni, T., et al. (2023). Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193. DOI: 10.48550/arXiv.2304.07193.
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. (2012). Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498-3505. IEEE. DOI: 10.1109/CVPR.2012.6248092.
Pedronette, D. C. G., Gonçalves, F. M. F., and Guilherme, I. R. (2018). Unsupervised manifold learning through reciprocal knn graph and connected components for image retrieval tasks. Pattern Recognition, 75:161-174. DOI: 10.1016/j.patcog.2017.05.009.
Pedronette, D. C. G., Penatti, O. A., and Torres, R. d. S. (2014). Unsupervised manifold learning using reciprocal knn graphs in image re-ranking and rank aggregation tasks. Image and Vision Computing, 32(2):120-130. DOI: 10.1016/j.imavis.2013.12.009.
Pedronette, D. C. G. and Torres, R. d. S. (2011). Exploiting contextual information for rank aggregation. In 2011 18th IEEE International Conference on Image Processing, pages 97-100. IEEE. DOI: 10.1109/ICIP.2011.6116726.
Pedronette, D. C. G. and Torres, R. d. S. (2012). Combining re-ranking and rank aggregation methods. In Progress in Pattern Recognition, Image Analysis, Computer Vision, and Applications: 17th Iberoamerican Congress, CIARP 2012, Buenos Aires, Argentina, September 3-6, 2012. Proceedings 17, pages 170-178. Springer. DOI: 10.1007/978-3-642-33275-3_21.
Pedronette, D. C. G. and Torres, R. d. S. (2013). Image re-ranking and rank aggregation based on similarity of ranked lists. Pattern Recognition, 46(8):2350-2360. DOI: 10.1016/j.patcog.2013.01.004.
Pedronette, D. C. G., Valem, L. P., Almeida, J., and Torres, R. d. S. (2019). Multimedia retrieval through unsupervised hypergraph-based manifold ranking. IEEE Transactions on Image Processing, 28(12):5824-5838. DOI: 10.1109/TIP.2019.2920526.
Pereira-Ferrero, V. H., Lewis, T. G., Valem, L. P., Ferrero, L., Pedronette, D. C., and Latecki, L. J. (2024). Unsupervised affinity learning based on manifold analysis for image retrieval: A survey. Computer Science Review, 53:100657. DOI: 10.1016/j.cosrev.2024.100657.
Rafieian, B., Hermosilla, P., and Vázquez, P.-P. (2023). Improving dimensionality reduction projections for data visualization. Applied Sciences, 13(17):9967. DOI: 10.3390/app13179967.
Roweis, S. T. and Saul, L. K. (2000). Nonlinear Dimensionality Reduction by Locally Linear Embedding. Science, 290(5500):2323-2326. DOI: 10.1126/science.290.5500.2323.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015). ImageNet Large Scale Visual Recognition Challenge. Int. J. Computer Vision, 115(3):211-252. DOI: 10.1007/s11263-015-0816-y.
Sabahi, F., Ahmad, M. O., and Swamy, M. (2024). Refinerhash: a new hashing-based re-ranking technique for image retrieval. Multimedia Systems, 30(3):119. DOI: 10.1007/s00530-024-01296-x.
Sánchez-Rico, M., Hoertel, N., and Alvarado, J. (2023). Combination of cluster analysis with dimensionality reduction techniques for pattern recognition studies in healthcare data: Comparing pca, t-sne and umap. DOI: 10.31234/osf.io/zxvf2.
Sculley, D. (2007). Rank aggregation for similar items. In Proceedings of the 2007 SIAM international conference on data mining, pages 587-592. SIAM. DOI: 10.1137/1.9781611972771.66.
Tenenbaum, J. B., Silva, V. d., and Langford, J. C. (2000). A global geometric framework for nonlinear dimensionality reduction. science, 290(5500):2319-2323. DOI: 10.1126/science.290.5500.2319.
Trozzi, F., Wang, X., and Tao, P. (2021). Umap as a dimensionality reduction tool for molecular dynamics simulations of biomacromolecules: a comparison study. The Journal of Physical Chemistry B, 125(19):5022-5034. DOI: 10.1021/acs.jpcb.1c02081.
Valem, L. P., Kawai, V. A. S., Pereira-Ferrero, V. H., and Pedronette, D. C. G. (2022). A novel rank correlation measure for manifold learning on image retrieval and person re-id. In 2022 IEEE International Conference on Image Processing (ICIP), pages 1371-1375. IEEE. DOI: 10.1109/ICIP46576.2022.9898060.
Valem, L. P., Pedronette, D. C. G., and Almeida, J. (2018). Unsupervised similarity learning through cartesian product of ranking references. Pattern Recognition Letters, 114:41-52. DOI: 10.1016/j.patrec.2017.10.013.
Valem, L. P., Pedronette, D. C. G., and Latecki, L. J. (2023). Rank flow embedding for unsupervised and semi-supervised manifold learning. IEEE Transactions on Image Processing. DOI: 10.1109/TIP.2023.3268868.
van der Maaten, L. and Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9:2579-2605. Available at:[link].
Vharkate, M. N. and Musande, V. B. (2022). Fusion based feature extraction and optimal feature selection in remote sensing image retrieval. Multimedia Tools and Applications, 81(22):31787-31814. DOI: 10.1007/s11042-022-11997-y.
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011). The caltech-ucsd birds-200-2011 dataset. Technical report. Available at:[link].
Wang, Y., Huang, H., Rudin, C., and Shaposhnik, Y. (2021). Understanding how dimension reduction tools work: an empirical approach to deciphering t-sne, umap, trimap, and pacmap for data visualization. Journal of Machine Learning Research, 22(201):1-73. Available at:[link].
Xiao, Z., Suma, P., Sachdeva, A., Wang, H.-J., Kordopatis-Zilos, G., Tolias, G., and Ordonez, V. (2025). Locore: Image re-ranking with long-context sequence modeling. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9580-9590. DOI: 10.1109/CVPR52734.2025.00895.
Yang, X., Koknar-Tezel, S., and Latecki, L. J. (2009). Locally constrained diffusion process on locally densified distance spaces with applications to shape retrieval. In 2009 IEEE conference on computer vision and pattern recognition, pages 357-364. IEEE. DOI: 10.1109/CVPR.2009.5206844.
Zhou, Y., Zeng, D., Zhang, S., and Tian, Q. (2015). Augmented feature fusion for image retrieval system. In Proceedings of the 5th ACM on International Conference on Multimedia Retrieval, pages 447-450. DOI: 10.1145/2671188.2749288.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Vinicius Atsushi Sato Kawai, Gustavo Rosseto Leticio, Lucas Pascotti Valem, Daniel Carlos Guimarães Pedronette

This work is licensed under a Creative Commons Attribution 4.0 International License.

