Computação Paralela Usando Arquiteturas ARM: Uma Análise Experimental com Raspberry Pi e Algoritmos de Ordenação de Dados
DOI:
https://doi.org/10.5753/reic.2026.7792Keywords:
Computação de Alto Desempenho, Algoritmos de Ordenação, Arquitetura ARM, Paralelismo, Raspberry PiAbstract
A Computação de Alto Desempenho (High-Performance Computing -- HPC) combina hardware e software para explorar paralelismo e distribuição de tarefas, sendo essencial para problemas computacionalmente intensivos. Este trabalho apresenta uma análise experimental do desempenho de oito algoritmos de ordenação em arquiteturas ARM (Advanced RISC Machines) baseadas em Raspberry Pi, considerando abordagens sequenciais, OpenMP (Open Multi-Processing) e MPI (Message Passing Interface). Foram avaliados cinco algoritmos clássicos (Insertion Sort, Merge Sort, Quick Sort, Heap Sort e Radix Sort) e três orientados à execução paralela (Odd-Even Bubble Sort, Bitonic Sort e Sample Sort). Os resultados indicam que o MPI favorece algoritmos divisíveis, como Merge Sort e Sample Sort, enquanto o OpenMP favorece algoritmos com paralelismo intra-nó, como Quick Sort e Heap Sort. O Insertion Sort, por sua natureza sequencial, não foi avaliado quanto ao speedup, enquanto o Radix Sort apresentou ganhos limitados, evidenciando a influência do algoritmo, paralelismo e hardware, conforme a Lei de Amdahl.
Downloads
Referências
Ala’anzy, M. A., Mazhit, Z., Ala’anzy, A. F., Algarni, A., Akhmedov, R., and Bauyrzhan, A. (2024). Comparative analysis of sorting algorithms: A review. In Proceedings of the 11th International Conference on Soft Computing & Machine Intelligence (ISCMI), pages 88–100, Melbourne, Australia. DOI: 10.1109/ISCMI63661.2024.10851593.
Andrews, G. R. (2001). Foundations of Multithreaded, Parallel, and Distributed Programming. Addison-Wesley, Boston, MA, 2 edition.
Aung, H. H. (2019). Analysis and comparative of sorting algorithms. International Journal of Trend in Scientific Research and Development, 3(5):1049–1053.
Axtmann, M., Witt, S., Ferizovic, D., and Sanders, P. (2017). In-place parallel super scalar samplesort (ips4o). In Proceedings of the 25th European Symposium on Algorithms (ESA 2017), volume 87 of Leibniz International Proceedings in Informatics (LIPIcs). DOI: 10.4230/LIPIcs.ESA.2017.9.
Batcher, K. E. (1968). Sorting networks and their applications. In Proceedings of the AFIPS Spring Joint Computer Conference. DOI: 10.1145/1468075.1468121.
Bilardi, G. and Nicolau, A. (1989). Adaptive bitonic sorting: An optimal parallel algorithm for shared-memory machines. SIAM Journal on Computing, 18(2). DOI: 10.1137/0218014.
Blelloch, G. E., Leiserson, C. E., Maggs, B. M., et al. (1998). A comparison of sorting algorithms for the connection machine cm-2. Theory of Computing Systems. DOI: 10.1145/113379.113380.
Chandra, R., Dagum, L., Kohr, D., Maydan, D., McDonald, J., and Menon, R. (2000). Parallel Programming in OpenMP. Morgan Kaufmann, San Francisco, CA.
Colichio, H., Pimenta, Y., Borges, P., Menecucci, M., Barros, L., and Guardia, H. (2024). Ordenação paralela com openmp e cuda. In Escola Regional de Alto Desempenho de São Paulo (ERAD-SP), pages 9–12, Rio Claro, SP, Brasil. DOI: 10.5753/eradsp.2024.239919.
Cormen, T. H., Leiserson, C. E., Rivest, R. L., and Stein, C. (2024). Algoritmos: Teoria e Prática. GEN LTC, 4 edition.
Esmaeilzadeh, H., Blem, E., St. Amant, R., Sankaralingam, K., and Burger, D. (2011). Dark silicon and the end of multicore scaling. In Proceedings of the 38th Annual International Symposium on Computer Architecture (ISCA). DOI: 10.1145/2000064.2000108.
Frazer, W. D. and McKellar, A. C. (1970). Samplesort: A sampling approach to minimal storage tree sorting. Journal of the ACM, 17(3). DOI: 10.1145/321592.321600.
Google Cloud (2025). Bigquery: Managed data warehouse for analytics.
Gropp, W., Lusk, E., and Skjellum, A. (1994). Using MPI: Portable Parallel Programming with the Message Passing Interface. MIT Press, Cambridge, MA.
Gustafson, J. L. (1988). Reevaluating amdahl’s law. Communications of the ACM, 31(5). DOI: 10.1145/42411.42415.
Góes, L. F. W. et al. (2023). Challenges in high-performance computing. Journal of the Brazilian Computer Society, 29(1). DOI: 10.5753/jbcs.2023.2219.
Habermann, A. N. (1972). Parallel neighbor sort (or the glory of the induction principle). Technical Report CS-72-112, Carnegie-Mellon University.
Hager, G. and Wellein, G. (2010). Introduction to High Performance Computing for Scientists and Engineers. CRC Press, Boca Raton, FL.
Hiremath, S. K. et al. (2024). A systematic comprehensive review and survey on high-performance computing systems approaches in scientific applications. Journal of Computer Science Engineering and Software Testing.
Hoefler, T. and Belli, R. (2015). Scientific benchmarking of parallel computing systems: Twelve ways to tell the masses when reporting performance results. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’15), Austin, TX. DOI: 10.1145/2807591.2807644.
Ignacio, A. L. J. and Dias, W. R. A. (2023). Análise do desempenho computacional de algoritmos paralelizados com openmp e mpi executados em raspberry pi. In Workshop de Iniciação Científica - Simpósio em Sistemas Computacionais de Alto Desempenho (SSCAD), pages 41–48, Porto Alegre, RS, Brasil. DOI: 10.5753/wscad_estendido.2023.235967.
Knuth, D. E. (1998). The Art of Computer Programming, Volume 3: Sorting and Searching. Addison-Wesley.
Kumar, D., Rajput, M. A., Kumar, P., Ahmed, S., Bhatti, M. S., and Tsetse, A. (2023). Performance evaluation of ARM-based versus x86-based processors in high performance computing clusters. Journal of Independent Studies and Research Computing (JISR-C), 21(2). DOI: 10.31645/JISRC.23.21.2.6.
Lakshmivarahan, S., Yang, S. K., and Dhall, S. K. (1984). Parallel sorting algorithms. Advances in Computers, 23. DOI: 10.1016/S0065-2458(08)60467-2.
Lima, F. A., Moreno, E. D., and Dias, W. R. A. (2016). Performance analysis of a low cost cluster with parallel applications and arm processors. IEEE Latin America Transactions, 14(11). DOI: 10.1109/TLA.2016.7795834.
Moore, G. E. (1965). Cramming more components onto integrated circuits. Electronics, 38(8). DOI: 10.1109/N-SSC.2006.4785860.
MPI Forum (2025). Mpi: A message-passing interface standard version 2.1.
OpenMP Architecture Review Board (2025). The openmp api specification for parallel programming.
Ou, Z., Pang, B., Deng, Y., Nurminen, J. K., Ylä-Jääski, A., and Hui, P. (2012). Energy- and cost-efficiency analysis of arm-based clusters. In Proceedings of the IEEE International Conference on Cluster Computing. DOI: 10.1109/CCGrid.2012.84.
Pacheco, P. S. (2011). An Introduction to Parallel Programming. Morgan Kaufmann.
Parhami, B. (2006). Introduction to Parallel Processing: Algorithms and Architectures. Kluwer Academic Publishers, New York.
Peters, H., Schulz-Hildebrandt, O., and Luttenberger, N. (2010). Fast in-place sorting with cuda based on bitonic sort. In Parallel Processing and Applied Mathematics. PPAM 2009. Springer. DOI: 10.1007/978-3-642-14390-8_42.
Rajagopal, D. and Thilakavalli, K. (2016). Different sorting algorithm’s comparison based upon the time complexity. International Journal of u- and e-Service, Science and Technology, 9(8). DOI: 10.14257/IJUNESST.2016.9.8.24.
Rauber, T. and Rünger, G. (2013). Parallel Programming for Multicore and Cluster Systems. Springer, 2 edition.
Regi, G., Kunnappally, J. J., and Yunus, R. (2024). Comparative analysis of energy efficiency between ARM and x86 architectures in mobile devices and its implications for human-computer interaction. Kristu Jayanti Journal of Computational Sciences, 4. DOI: 10.59176/kjcs.v4i1.2434.
Sedgewick, R. (1983). Algorithms. Addison-Wesley.
Simakov, N. A., DeLeon, R. L., White, J. P., Jones, M. D., and Furlani, T. R. (2023). Are we ready for broader adoption of ARM in the HPC community: Performance and energy efficiency analysis of benchmarks and applications executed on high-end ARM systems. In International Conference on High Performance Computing in Asia-Pacific Region Workshops (HPCASIAWORKSHOP 2023). DOI: 10.1145/3581576.3581618.
Xavier, F., Camargo, E. T., and Duarte Jr., E. (2019). Uma implementação mpi tolerante a falhas do algoritmo paralelo de ordenação quickmerge. In Simpósio em Sistemas Computacionais de Alto Desempenho (SSCAD), Campo Grande, MS, Brasil. DOI: 10.5753/wscad.2019.8675.
Yue, Y. (2015). Performance evaluation of sorting algorithms in raspberry pi and personal computer. Master’s thesis, University of Vaasa.
Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., and Stoica, I. (2016). Apache spark: A unified engine for big data processing. Communications of the ACM, 59(11). DOI: 10.1145/2934664.
Downloads
Published
Como Citar
Issue
Section
Licença
Copyright (c) 2026 Os autores

Este trabalho está licenciado sob uma licença Creative Commons Attribution 4.0 International License.
