Learning Parallel Computing with Multilayer Perceptron Neural Networks

Authors

DOI:

https://doi.org/10.5753/reic.2026.8178

Keywords:

Parallel Programming, Performance Evaluation, Multilayer Perceptron Neural Network, Parallelism Education

Abstract

As multi-core architectures become the standard for computational efficiency, parallel programming has evolved into a primary avenue for optimization and resource management. However, an integrated design-and-evaluation method is often difficult for beginners to envision, leading to incorrect intuition about the relationship between theory and practice. This article proposes an implementation-driven methodological approach using a Multilayer Perceptron (MLP) to bridge this gap, comprising three main components: (1) establishing a sequential baseline, (2) applying parallelization and data containerization strategies to the MLP, and (3) evaluating through key metrics such as speedup and hardware efficiency. By addressing data dependencies and synchronization challenges inherent to backpropagation, the proposed trajectory provides a framework for comprehending data integrity in parallel systems through an exemplified parallel implementation. Our approach led to successful optimization of the feedforward and backpropagation routines, resulting in a speedup, e.g., of 4 times the baseline execution time using 8 hardware threads. Furthermore, the results indicate that outer-loop nesting and data containerization should be considered for managing structures with high data dependencies.

Downloads

Não há dados estatísticos.

Referências

Abdurhaman, A., Singh, A., Hossain, A., and Ahmed, K. (2024). A hands-on approach to teaching parallel and heterogeneous computing. In 2024 IEEE 31st International Conference on High Performance Computing, Data and Analytics Workshop (HiPCW), pages 9–16. DOI: 10.1109/HiPCW63042.2024.00012.

Akinsola, J. E. T., Olatunbosun, M. A., Olaniyi, I. M., Adeagbo, M. A., Olajubu, E. A., and Aderounmu, G. A. (2025). Application of artificial intelligence on mnist dataset for handwritten digit classification for evaluation of deep learning models. Acadlore Transactions on AI and Machine Learning. DOI: 10.56578/ataiml040305.

Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. In Proceedings of the April 18-20, 1967, Spring Joint Computer Conference, AFIPS ’67 (Spring), page 483–485, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/1465482.1465560.

Ayguade, E., Martorell, X., Labarta, J., Gonzalez, M., and Navarro, N. (1999). Exploiting multiple levels of parallelism in openmp: a case study. In Proceedings of the 1999 International Conference on Parallel Processing, pages 172–180. DOI: 10.1109/ICPP.1999.797402.

Bindi, L., D’Amico, S., Mencagli, G., and Torquati, M. (2026). Enabling pinning strategies for stream processing applications on multicores. International Journal of Parallel Programming, 54(2). DOI: 10.1007/s10766-026-00813-x.

Carneiro Neto, J. A., Alves Neto, A. J., and Moreno, E. D. (2022). A systematic review on teaching parallel programming. In Proceedings of the 11th Euro American Conference on Telematics and Information Systems, EATIS ’22, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/3544538.3544659.

Chadwick, A. W., Erdős, M., Bora, U., Bhosale, A., Lytton, B., Guo, Y., Cooper, R., Gabrielli, G., and Jones, T. M. (2025). The future of instruction-level parallelism (ilp). In 2025 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), pages 350–352. DOI: 10.1109/ISPASS64960.2025.00040.

Che, H. and Nguyen, M. (2014). Amdahl’s law for multithreaded multicore processors. Journal of Parallel and Distributed Computing, 74(10):3056–3069. DOI: 10.1016/j.jpdc.2014.06.012.

Conte, D. J., de Souza, P. S. L., Martins, G., and Bruschi, S. M. (2020). Teaching parallel programming for beginners in computer science. In 2020 IEEE Frontiers in Education Conference (FIE), pages 1–9. DOI: 10.1109/FIE44824.2020.9274155.

de Freitas, H. C. (2012). Introducing parallel programming to traditional undergraduate courses. In 2012 Frontiers in Education Conference Proceedings, pages 1–6. DOI: 10.1109/FIE.2012.6462263.

Gustafson, J. L. (1988). Reevaluating amdahl’s law. Commun. ACM, 31(5):532–533. DOI: 10.1145/42411.42415.

Hill, M. D. and Marty, M. R. (2008). Amdahl’s law in the multicore era. Computer, 41(7):33–38. DOI: 10.1109/MC.2008.209.

Huang, K.-T., Lin, T.-Y., Cheng, P.-W., and Chen, P.-S. (2024). Enhancing tvm vta simulator performance through simd vectorization. In 2024 10th International Conference on Applied System Innovation (ICASI), pages 418–420. DOI: 10.1109/ICASI60819.2024.10547748.

Indrakumari, R., Poongodi, T., and Singh, K. (2021). Introduction to Deep Learning, pages 1–22. Springer International Publishing, Cham. DOI: 10.1007/978-3-030-66519-7_1.

Islam, M., Chen, G., and Jin, S. (2019). An overview of neural network. American Journal of Neural Networks and Applications, 5:05. DOI: 10.11648/j.ajnna.20190501.12.

Iudean, B. (2024). Experience report on teaching parallel and distributed programming through storytelling. In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering - Companion (SANER-C), pages 1–9. DOI: 10.1109/SANER-C62648.2024.00019.

Kumar, K., Pappu, A., Kumar, K., and Sanyal, S. (2006). Hybrid approach for parallelization of sequential code with function level and block level parallelization. In International Symposium on Parallel Computing in Electrical Engineering (PARELEC’06), pages 161–166. DOI: 10.1109/PARELEC.2006.44.

Kumar, S. A. (2017). Achieving software parallelism through alternative process models. In 2017 International Conference on Inventive Systems and Control (ICISC), pages 1–4. DOI: 10.1109/ICISC.2017.8068596.

Lara Soares, F. A., Neri Nobre, C., and Cota de Freitas, H. (2019). Parallel programming in computing undergraduate courses: a systematic mapping of the literature. IEEE Latin America Transactions, 17(08):1371–1381. DOI: 10.1109/TLA.2019.8932371.

Lupo, C., Wood, Z. J., and Victorino, C. (2012). Cross teaching parallelism and ray tracing: a project-based approach to teaching applied parallel computing. In Proceedings of the 43rd ACM Technical Symposium on Computer Science Education, SIGCSE ’12, page 523–528, New York, NY, USA. Association for Computing Machinery. DOI: 10.1145/2157136.2157288.

Ma, Y., Chen, X., Guo, H., Li, F., and Liu, X. (2024). Multilevel load balancing algorithm for domestic heterogeneous manycore architecture. In 2024 IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA), pages 1931–1936. DOI: 10.1109/ISPA63168.2024.00263.

Machado, P. R. D. S., Souza, M. A., and Freitas, H. C. D. (2026). Evaluation of thread scalability and compiler optimization in parallel job execution using unity’s job system. IEEE Access, 14:5012–5025. DOI: 10.1109/ACCESS.2026.3651527.

Macukow, B. (2016). Neural Networks – State of Art, Brief History, Basic Models and Architecture. In Computer Information Systems and Industrial Management, pages 3–14, Cham. Springer International Publishing. DOI: 10.1007/978-3-319-45378-1_1.

Magalhães, J., Gonçalves, A., Nobre, C., and Freitas, H. (2024). Avaliação de desempenho e escalabilidade do algoritmo de otimização de colônia de formigas em c++ e python. In Anais do XXV Simpósio em Sistemas Computacionais de Alto Desempenho, pages 97–108, Porto Alegre, RS, Brasil. SBC. DOI: 10.5753/sscad.2024.244375.

Markidis, S. (2024). What is quantum parallelism, anyhow? In ISC High Performance 2024 Research Paper Proceedings (39th International Conference), pages 1–12. DOI: 10.23919/ISC.2024.10528926.

OpenMP Architecture Review Board (2025). Technical report 14: Public comment draft of openmp api version 6.1. Technical Report TR14, OpenMP Architecture Review Board. Available at: [link].

Pawar, P., Mehta, P., Boggarapu, N., and Grange, L. (2014). Enhanced automated data dependency analysis for functionally correct parallel code. In Proceedings of the 2014 International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA’14), pages 253–259. WorldComp. Available at: [link].

Ribeiro, C. P., Castro, M., Marangozova-Martin, V., Mehaut, J.-F., Freitas, H. C., and Martins, C. A. P. S. (2012). Evaluating cpu and memory affinity for numerical scientific multithreaded benchmarks on multi-cores. IADIS International Journal on Computer Science and Information Systems, 7:79–93. Available at: [link].

Rosenblatt, F. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychological Review, 65(6):386–408. DOI: 10.1037/h0042519.

Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323:533–536. DOI: 10.1038/323533a0.

Strazdins, P. E. (2025). A simple tiled approach to teaching parallel computing. In 2025 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pages 385–394. DOI: 10.1109/IPDPSW66978.2025.00064.

Vargas-Pérez, S. (2024). Teaching performance metrics in parallel computing courses. In 2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pages 385–390. DOI: 10.1109/IPDPSW63119.2024.00086.

Vasconcelos, L. B. A., Soares, F. A. L., Penna, P. H. M. M., Machado, M. V., Góes, L. F. W., Martins, C. A. P. S., and Freitas, H. C. (2019). Teaching parallel programming to freshmen in an undergraduate computer science program. In 2019 IEEE Frontiers in Education Conference (FIE), pages 1–8. DOI: 10.1109/FIE43999.2019.9028566.

Downloads

Published

2026-08-07

Como Citar

Diniz, M. H. M., & Cota de Freitas, H. (2026). Learning Parallel Computing with Multilayer Perceptron Neural Networks. Revista Eletrônica De Iniciação Científica Em Computação, 24(1), 582–593. https://doi.org/10.5753/reic.2026.8178

Issue

Section

Artigos