On the Effectiveness of Feedforward Neural Networks for Performance Prediction in Highly Configurable Systems: The Case of Linux Kernel Size
DOI:
https://doi.org/10.5753/jserd.2026.6399Keywords:
Software Product Lines, Configurable Systems, Deep Learning, Performance PredictionAbstract
Highly configurable software systems, such as the Linux kernel, expose thousands of configuration options whose combinations can drastically influence non-functional properties. In embedded or resource-constrained environments, kernel binary size is particularly important because it directly affects memory footprint, boot time, storage usage, and overall system responsiveness. Accurately predicting binary size from configuration options enables developers to evaluate trade-offs early in the development cycle, reducing the need for costly trial-and-error compilation and measurement processes. This study investigates the use of Feedforward Neural Networks (FNNs) for predicting Linux kernel binary size and compares their performance against strong classical Machine Learning (ML) baselines, including Random Forest and Gradient Boosting Trees. In addition, we analyze the impact of feature selection (FS) strategies on predictive accuracy and computational cost. We evaluate five FS approaches, including supervised tree-based importance rankings and a semantic, label-free method based on Word2Vec embeddings extracted from Linux kernel documentation. Our experiments were conducted on a dataset containing 9,670 Linux kernel configuration options. The results show that classical tree-based ensemble methods outperformed FNNs in predictive accuracy, with Gradient Boosting Trees achieving the best overall results (MAPE of 5.21%). Although FNNs combined with feature selection achieved reasonable accuracy (best MAPE of 8.26%) and benefited from reduced training times, they did not surpass the traditional ML baselines. These findings provide important empirical evidence that more complex neural architectures do not necessarily yield superior performance for SPL prediction tasks, even in high-dimensional configuration spaces. We also show that the Word2Vec-based semantic FS method offers a practical label-free alternative for early-stage scenarios where measured NFP data are unavailable or expensive to obtain, although it generally underperforms supervised feature selection strategies in both accuracy and efficiency. Overall, our findings provide practical guidance for SPL practitioners and researchers in selecting prediction models and preprocessing strategies that balance accuracy, interpretability, and computational cost in highly configurable systems.
Downloads
References
Acher, M., Martin, H., Lesoil, L., Blouin, A., Jézéquel, J.-M., Khelladi, D. E., Barais, O., and Pereira, J. A. (2022). Feature subset selection for learning huge configuration spaces: The case of Linux kernel size. In Proceedings of the 26th ACM International Systems and Software Product Line Conference (SPLC), pages 85–96. ACM.
Acher, M., Martin, H., Pereira, J. A., Blouin, A., Jézéquel, J.-M., Khelladi, D. E., Lesoil, L., and Barais, O. (2019). Learning very large configuration spaces: What matters for Linux kernel sizes. Technical Report RR-9286, Inria.
Baldi, P. and Sadowski, P. J. (2013). Understanding dropout. Advances in Neural Information Processing Systems, 26:2814–2822.
Barron, J. T. (2019). A general and adaptive robust loss function. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4331–4339. IEEE.
Bergstra, J. and Bengio, Y. (2012). Random search for hyperparameter optimization. Journal of Machine Learning Research, 13:281–305.
Bessa, J. M., Cavalcanti, M., Acher, M., Endler, M., and Pereira, J. A. (2025). Unveiling the impact of sampling on feature selection for performance prediction in configurable systems. In Proceedings of the International Conference on Software and Systems Reuse (ICSR), pages 1–12. IEEE.
Bessa, J. M. and Pereira, J. A. (2023). A unified repository of product lines measurements. GitHub repository. [link]. Accessed: 15 Jun. 2023.
Biau, G. and Scornet, E. (2016). A random forest guided tour. TEST, 25(2):197–227.
Borges, H., Pereira, J. A., Khelladi, D. E., and Acher, M. (2025). Linux kernel configurations at scale: A dataset for performance and evolution analysis. In Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering (EASE), pages 1–10. ACM.
Chai, T. and Draxler, R. R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? Arguments against avoiding RMSE in the literature. Geoscientific Model Development, 7(3):1247–1250.
Chen, T. and Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 785–794. ACM.
Cheng, J., Gao, C., and Zheng, Z. (2023). HINNPerf: Hierarchical interaction neural network for performance prediction of configurable systems. ACM Transactions on Software Engineering and Methodology, 32(2):1–30.
Clevert, D.-A., Unterthiner, T., and Hochreiter, S. (2016). Fast and accurate deep network learning by exponential linear units (ELUs). In Proceedings of the 4th International Conference on Learning Representations (ICLR).
Gong, J. and Chen, T. (2025). Deep configuration performance learning: A systematic survey and taxonomy. ACM Transactions on Software Engineering and Methodology, 34(1):1–52.
Guo, J., Czarnecki, K., Apel, S., Siegmund, N., and Wąsowski, A. (2013). Variability-aware performance prediction: A statistical learning approach. In Proceedings of the 28th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 301–311. IEEE.
Ha, H. and Zhang, H. (2019a). DeepPerf: Performance prediction for configurable software with deep sparse neural network. In Proceedings of the 41st IEEE/ACM International Conference on Software Engineering (ICSE), pages 1095–1106. IEEE.
Ha, H. and Zhang, H. (2019b). Performance-influence model for highly configurable software with Fourier learning and Lasso regression. In Proceedings of the 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME), pages 470–480. IEEE.
He, K., Zhang, X., Ren, S., and Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 1026–1034. IEEE.
He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778. IEEE.
Huang, L., Zhang, B., and Liu, Y. (2021). A systematic study of early stopping for deep learning. IEEE Access, 9:120501–120512.
Kim, M., Notkin, D., and Grossman, D. (2007). Automatic inference of structural changes for matching across program versions. In Proceedings of the 29th International Conference on Software Engineering (ICSE), pages 333–343. IEEE.
Kingma, D. P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint, arXiv:1412.6980.
Kotsiantis, S. B. (2013). Decision trees: A recent overview. Artificial Intelligence Review, 39(4):261–283.
Li, Z., Zhang, Z., Xu, R., Ding, Z., and Tao, D. (2022). DeepFS: An end-to-end deep feature selection model for tabular data.
Loshchilov, I. and Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations (ICLR).
Ma, Y., Huang, Q., Wang, J., Xu, K., Goldblum, M., and Goldstein, T. (2022). Overfitting in adversarially robust deep learning.
Maas, A. L., Hannun, A. Y., and Ng, A. Y. (2013). Rectifier nonlinearities improve neural network acoustic models. In Proceedings of the 30th International Conference on Machine Learning (ICML), Workshop on Deep Learning for Audio, Speech and Language Processing.
Martin, H., Acher, M., Pereira, J. A., Lesoil, L., Jézéquel, J.-M., and Khelladi, D. E. (2022). Transfer learning across variants and versions: The case of Linux kernel size. IEEE Transactions on Software Engineering, 48(11):4274–4290.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26 (NeurIPS), pages 3111–3119.
Nair, V. and Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML), pages 807–814.
Nasrabadi, N. M. (2007). Pattern recognition and machine learning. Journal of Electronic Imaging, 16(4):049901.
Natekin, A. and Knoll, A. (2013). Gradient boosting machines, a tutorial. Frontiers in Neurorobotics, 7:21.
Ochoa, L., Gonzalez-Rojas, O., Pereira, J. A., Castro, H., and Saake, G. (2018). A systematic literature review on the semi-automatic configuration of extended product lines. Journal of Systems and Software, 144:511–532.
Ochoa, L., Pereira, J. A., González-Rojas, O., Castro, H., and Saake, G. (2017). A survey on scalability and performance concerns in extended product lines configuration. In Proceedings of the 11th International Workshop on Variability Modelling of Software-Intensive Systems (VaMoS), pages 5–12. ACM.
Pereira, J. A., Acher, M., Martin, H., and Jézéquel, J.-M. (2020a). Sampling effect on performance prediction of configurable systems: A case study. In Proceedings of the ACM/SPEC International Conference on Performance Engineering (ICPE), pages 277–288. ACM.
Pereira, J. A., Acher, M., Martin, H., Jézéquel, J.-M., Botterweck, G., and Ventresque, A. (2021). Learning software configuration spaces: A systematic literature review. Journal of Systems and Software, 182:111044.
Pereira, J. A., Constantino, K., and Figueiredo, E. (2014). A systematic literature review of software product line management tools. In Proceedings of the 13th International Conference on Software Reuse (ICSR), pages 73–89. Springer.
Pereira, J. A., Krieter, S., Meinicke, J., Schröter, R., Saake, G., and Leich, T. (2016a). FeatureIDE: Scalable product configuration of variable systems. In Proceedings of the 15th International Conference on Software Reuse (ICSR), pages 397–401. Springer.
Pereira, J. A., Maciel, L., Noronha, T. F., and Figueiredo, E. (2017). Heuristic and exact algorithms for product configuration in software product lines. International Transactions in Operational Research, 24(6):1285–1306.
Pereira, J. A., Martin, H., Temple, P., and Acher, M. (2020b). Machine learning and configurable systems: A gentle introduction. In Proceedings of the 24th ACM International Systems and Software Product Line Conference (SPLC), pages 1–6. ACM.
Pereira, J. A., Matuszyk, P., Krieter, S., Spiliopoulou, M., and Saake, G. (2016b). A feature-based personalized recommender system for product-line configuration. In Proceedings of the 2016 ACM SIGPLAN International Conference on Generative Programming: Concepts and Experiences (GPCE), pages 120–131. ACM.
Pereira, J. A., Matuszyk, P., Krieter, S., Spiliopoulou, M., and Saake, G. (2018a). Personalized recommender systems for product-line configuration processes. Computer Languages, Systems & Structures, 54:451–471.
Pereira, J. A., Schulze, S., Krieter, S., Ribeiro, M., and Saake, G. (2018b). A context-aware recommender system for extended software product line configurations. In Proceedings of the 12th International Workshop on Variability Modelling of Software-Intensive Systems (VaMoS), pages 97–104. ACM.
Pereira, J. A., Souza, C., Figueiredo, E., Abilio, R., Vale, G., and Costa, H. A. X. (2013). Software variability management: An exploratory study with two feature modeling tools. In Proceedings of the 7th Brazilian Symposium on Software Components, Architectures and Reuse (SBCARS), pages 20–29. IEEE.
Raschka, S., Liu, Y., and Mirjalili, V. (2022). Machine Learning with PyTorch and Scikit-Learn. Packt Publishing.
Rokach, L. (2015). Data Mining with Decision Trees: Theory and Applications. World Scientific, 2nd edition.
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088):533–536.
Smith, L. N. (2018). A disciplined approach to neural network hyper-parameters: Part 1 — Learning rate, batch size, momentum, and weight decay.
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958.
Wohlin, C., Runeson, P., Höst, M., Ohlsson, M., Regnell, B., and Wesslén, A. (2012). Experimentation in Software Engineering. Springer.
Zhang, C., Bengio, S., and Singer, Y. (2021). Understanding deep learning requires rethinking generalization. Communications of the ACM, 64(3):107–115.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 João Marcello Bessa, Heraldo Borges, Pedro Lopes, Lucas Lopes, Mathieu Acher, Juliana Alves Pereira

This work is licensed under a Creative Commons Attribution 4.0 International License.

