Interpretable by Design: MH-AutoML for Transparent and Efficient Android Malware Detection without Compromising Performance
DOI:
https://doi.org/10.5753/jbcs.2026.6185Keywords:
Automated Machine Learning, AutoML, Android Malware Detection, Cybersecurity, Machine Learning Pipeline, Explainable Artificial Intelligence (XAI), Transparency in AI, Interpretability , Hyperparameter Optimization, Feature Engineering, Comparative Evaluation, MLOps, Model Selection, Domain-Specific AutoML, Black-Box SystemsAbstract
Malware detection in Android systems requires both cybersecurity domain knowledge and expertise in machine learning (ML). Automated Machine Learning (AutoML) aims to simplify ML development by automating key stages of the pipeline, including pipeline construction, algorithm selection, and hyperparameter optimisation, thereby reducing the need for specialised expertise. However, most general-purpose AutoML frameworks operate largely as black-box systems with limited transparency, interpretability, and experiment traceability. These limitations are particularly problematic in security-sensitive applications, where understanding model behaviour and ensuring reproducibility are essential. To address these challenges, we present MH-AutoML, a meta-heuristic-based AutoML framework designed specifically for Android malware detection. MH-AutoML automates the complete ML pipeline, covering data preprocessing, feature engineering, model selection, and hyperparameter optimisation. In addition, the framework incorporates built-in capabilities for interpretability, debugging, and experiment tracking, which are often absent or only partially supported in existing AutoML solutions. We perform a large-scale empirical evaluation comparing MH-AutoML with seven established AutoML frameworks (Auto-sklearn, AutoGluon, TPOT, HyperGBM, AutoPyTorch, LightAutoML, and MLJAR). The comparison is conducted across nine Android malware benchmark datasets evaluated under both original and class-balanced distributions using 10 random seeds, resulting in a total of 1,440 experimental runs. Using the Matthews Correlation Coefficient (MCC) as the primary evaluation metric, together with Friedman–Nemenyi statistical tests and pairwise Win/Tie/Loss analysis, the results reveal several key findings. First, MLJAR and LightAutoML form a consistently top-performing tier, with MLJAR achieving the highest aggregate dominance score across the analyses. Second, LightAutoML offers the most favourable efficiency–performance trade-off, reaching comparable predictive quality in approximately 9–17 minutes compared with MLJAR’s execution time of about one hour. Third, MH-AutoML achieves competitive recall performance, ranking first on three datasets including Adroit, Androcrawl, and MH-100, although this performance is accompanied by reduced MCC values due to elevated false-positive rates. Fourth, class balancing substantially increases the statistical separation among frameworks, doubling the Friedman test statistic and revealing fragile behaviour in certain tools. AutoPyTorch is the most notable case, collapsing from third position in the Original setting to the last rank under balanced distributions. Finally, no single framework dominates across all datasets, which reinforces the importance of selecting AutoML frameworks according to the specific characteristics of each dataset.
Downloads
References
Amirian, M., Tuggener, L., Chavarriaga, R., Satyawan, Y. P., Schilling, F.-P., Schwenker, F., and Stadelmann, T. (2021). Two to trust: Automl for safe modelling and interpretable deep learning for robustness. In Trustworthy AI-Integrating Learning, Optimization and Reasoning: First International Workshop, TAILOR 2020, Virtual Event, September 4-5, 2020, Revised Selected Papers 1, pages 268-275. Springer. DOI: 10.1007/978-3-030-73959-1_23.
Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., et al. (2020). Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information fusion. DOI: 10.1016/j.inffus.2019.12.012.
Assolin, J., Canto, G., Kreutz, D., and Feitosa, E. (2024). MH-AutoML: Transparência, interpretabilidade e desempenho na detecção de malware android. In Anais Estendidos do XXIV Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 113-120, Porto Alegre, RS, Brasil. SBC. DOI: 10.5753/sbseg_estendido.2024.243362.
Assolin, J., Kreutz, D., Siqueira, G., Rocha, V., Miers, C., Mansilha, R., and Feitosa, E. (2022). Droidautoml: uma ferramenta de automl para o domínio de detecção de malwares android. In Anais Estendidos do XXII Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 135-142, Porto Alegre, RS, Brasil. SBC. DOI: 10.5753/sbseg_estendido.2022.227037.
Bahri, M., Salutari, F., Putina, A., and Sozio, M. (2022). AutoML: state of the art with a focus on anomaly detection, challenges, and research directions. International Journal of Data Science and Analytics. DOI: 10.1007/s41060-022-00309-0.
Baratchi, M., Wang, C., Limmer, S., van Rijn, J. N., Hoos, H., Bäck, T., and Olhofer, M. (2024). Automated machine learning: past, present and future. Artificial intelligence review, 57(5):122. DOI: 10.1007/s10462-024-10726-1.
Barbudo, R., Ventura, S., and Romero, J. R. (2023). Eight years of automl: categorisation, review and trends. Knowledge and Information Systems, 65(12):5097-5149. DOI: 10.1007/s10115-023-01935-1.
Bifarin, O. O. and Fernández, F. M. (2024). Automated machine learning and explainable ai (automl-xai) for metabolomics: improving cancer diagnostics. Journal of the American Society for Mass Spectrometry, 35(6):1089-1100. DOI: 10.1021/jasms.3c00403.
Bilal, M., Ali, G., Iqbal, M. W., Anwar, M., Malik, M. S. A., and Kadir, R. A. (2022). Auto-prep: Efficient and automated data preprocessing pipeline. IEEE Access. DOI: 10.1109/ACCESS.2022.3198662.
Bragança, H., Rocha, V., Barcellos, L., Souto, E., Kreutz, D., and Feitosa, E. (2023). Capturing the behavior of android malware with mh-100k: A novel and multidimensional dataset. In Anais do XXIII Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 510-515, Porto Alegre, RS, Brasil. SBC. DOI: 10.5753/sbseg.2023.233596.
Brännström, M. (2023). Transparency of complex systems: The semantic transparency framework. Available at: [link].
Carvalho, D. V., Pereira, E. M., and Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics. DOI: 10.3390/electronics8080832.
Colaco, C. W., Bagwe, M. D., Bose, S. A., and Jain, K. (2021). DefenseDroid: A Modern Approach to Android Malware Detection. Strad Research, 8(5):271-282. DOI: 10.37896/sr8.5/027.
da Silva, M. C., Licari, B., Tavares, G. M., and Junior, S. B. (2024). Benchmarking automl clustering frameworks. In AutoML Conference 2024 (ABCD Track). DOI: 10.3390/info15010063.
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine learning research, 7(Jan):1-30. Available at: [link].
Doke, A. and Gaikwad, M. (2021). Survey on automated machine learning (automl) and meta learning. In 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT), pages 1-5. IEEE. DOI: 10.1109/ICCCNT51525.2021.9579526.
Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., and Smola, A. (2020). AutoGluon-Tabular: Robust and accurate AutoML for structured data. [link].
Ferreira, L., Pilastri, A., Martins, C. M., Pires, P. M., and Cortez, P. (2021). A comparison of AutoML tools for machine learning, deep learning and xgboost. In IJCNN, pages 1-8. DOI: 10.1109/IJCNN52387.2021.9534091.
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., and Hutter, F. (2015). Efficient and robust automated machine learning. Advances in neural information processing systems, 28. DOI: 10.1007/978-3-030-05318-5_6.
Garouani, M. and Bouneffa, M. (2023). Unlocking the black box: Towards interactive explainable automated machine learning. In International Conference on Intelligent Data Engineering and Automated Learning, pages 458-469. Springer. DOI: 10.1007/978$-3$-031$-48232$-8_42.
Gijsbers, P., Bueno, M. L., Coors, S., LeDell, E., Poirier, S., Thomas, J., Bischl, B., and Vanschoren, J. (2024). Amlb: an automl benchmark. Journal of Machine Learning Research. DOI: 10.48550/arXiv.2207.12560.
Giovanelli, J., Bilalli, B., and Abelló, A. (2022). Data pre-processing pipeline generation for autoetl. Information Systems. DOI: 10.1016/j.is.2021.101957.
Guerra-Manzanares, A., Bahsi, H., and Nõmm, S. (2021). KronoDroid: Time-based Hybrid-featured Dataset for Effective Android Malware Detection and Characterization. Computers & Security, 110:102399. DOI: 10.1016/j.cose.2021.102399.
Hariri-Ardebili, M. A., Mahdavi, P., and Pourkamali-Anaraki, F. (2024). Benchmarking automl solutions for concrete strength prediction: Reliability, uncertainty, and dilemma. Construction and Building Materials, 423:135782. DOI: 10.1016/j.conbuildmat.2024.135782.
Hasan, R., Dattana, V., Mahmood, S., and Hussain, S. (2024). Towards transparent diabetes prediction: Combining automl and explainable ai for improved clinical insights. Information, 16(1):7. DOI: 10.3390/info16010007.
Hutter, F., Kotthoff, L., and Vanschoren, J. (2019). Automated machine learning: methods, systems, challenges. Springer Nature. DOI: 10.1007/978-3-030-05318-5.
Jian Yang, Xuefeng Li, H. W. (2020). HyperGBM: A Full Pipeline AutoML Tool Integrated With Various GBM Models version 0.2.x. Available at:[link].
Karmaker S. et. al. (2021). Automl to date and beyond: Challenges and opportunities. ACM Computing Surveys, 54(8). DOI: 10.1145/3470918.
Kovalevsky, V., Delhibabu, R., and Zhukova, N. (2024). Automl framework for physical activities recognition. In 2024 3rd International Conference on Artificial Intelligence For Internet of Things (AIIoT), pages 1-5. IEEE. DOI: 10.1109/AIIoT58432.2024.10574646.
Krzywanski, J., Sztekler, K., Skrobek, D., Grabowska, K., Ashraf, W. M., Sosnowski, M., Ishfaq, K., Nowak, W., and Mika, L. (2024). AutoML-based predictive framework for predictive analysis in adsorption cooling and desalination systems. Energy Science & Engineering, 12(5):1969-1986. DOI: 10.1002/ese3.1725.
Kundu, P. P., Anatharaman, L., and Truong-Huu, T. (2021). An empirical evaluation of automated machine learning techniques for malware detection. In Proceedings of the 2021 ACM Workshop on Security and Privacy Analytics. DOI: 10.1145/3445970.3451155.
Le, T. T., Fu, W., and Moore, J. H. (2020). Scaling tree-based automated machine learning to biomedical big data with a feature set selector. Bioinformatics, 36(1):250-256. DOI: 10.1093/bioinformatics/btz470.
LeDell, E. and Poirier, S. (2020). H2O AutoML: Scalable automatic machine learning. In Proceedings of the AutoML Workshop at ICML, volume 2020. Available at: [link].
Love, P. E., Fang, W., Matthews, J., Porter, S., Luo, H., and Ding, L. (2023). Explainable artificial intelligence (xai): Precepts, models, and opportunities for research in construction. Advanced Engineering Informatics. DOI: 10.1016/j.aei.2023.102024.
Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems. DOI: 10.48550/arXiv.1705.07874.
Mahindru, A. (2018). Android permission dataset. Mendeley Data. Available at: [link].
Mahmood, S., Hasan, R., Hussain, S., and Adhikari, R. (2025). An interpretable and generalizable machine learning model for predicting asthma outcomes: Integrating automl and explainable ai techniques. World, 6(1):15. DOI: 10.3390/world6010015.
Martín, A., Calleja, A., Menéndez, H. D., Tapiador, J., and Camacho, D. (2016). Adroit: Android malware detection using meta-information. In 2016 IEEE Symposium Series on Computational Intelligence (SSCI), pages 1-8. IEEE. DOI: 10.1109/SSCI.2016.7849904.
Mi, J.-X., Li, A.-D., and Zhou, L.-F. (2020). Review study of interpretation methods for future interpretable machine learning. IEEE Access. DOI: 10.1109/ACCESS.2020.3032756.
Moayeri, M., Rabbat, M., Ibrahim, M., and Bouchacourt, D. (2024). Embracing diversity: Interpretable zero-shot classification beyond one vector per class. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2302-2321. DOI: 10.1145/3630106.3659039.
Molino, P., Dudin, Y., and Miryala, S. S. (2019). Ludwig: a type-based declarative deep learning toolbox. arXiv preprint arXiv:1909.07930. DOI: 10.48550/arXiv.1909.07930.
Narkar, S., Zhang, Y., Liao, Q. V., Wang, D., and Weisz, J. D. (2021). Model lineupper: Supporting interactive model comparison at multiple levels for automl. In Proceedings of the 26th International Conference on Intelligent User Interfaces, pages 170-174. DOI: 10.1145/3397481.3450658.
Naser, M. (2023). Machine learning for all! benchmarking automated, explainable, and coding-free platforms on civil and environmental engineering problems. Journal of Infrastructure Intelligence and Resilience. DOI: 10.1016/j.iintel.2023.100028.
Nasimian, A. et. al. (2024). Alphaml: A clear, legible, explainable, transparent, and elucidative binary classification platform for tabular data. Patterns, 5(1). DOI: 10.1101/2023.06.20.545752.
Neto, H. A., Alves, R. C., and Campos, S. V. (2020). Nasirt: Automl based learning with instance-level complexity information. arXiv preprint arXiv:2008.11846. DOI: 10.48550/arXiv.2008.11846.
Oakes, B. J., Famelis, M., and Sahraoui, H. (2024). Building domain-specific machine learning workflows: A conceptual framework for the state of the practice. ACM Transactions on Software Engineering and Methodology, 33(4):1-50. DOI: 10.1145/3638243.
Oliveira, S. d., Topsakal, O., and Toker, O. (2024). Benchmarking automated machine learning (automl) frameworks for object detection. Information, 15(1):63. DOI: 10.3390/info15010063.
Płońska, A. and Płoński, P. (2021). Mljar: State-of-the-art automated machine learning framework for tabular data. Version 0.10, 3. Available at: [link].
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). "why should i trust you?" explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. DOI: 10.1145/2939672.2939778.
Rocha, V., Bragança, H., Kreutz, D., and Feitosa, E. (2024). Mh-fsf: um framework para reprodução, experimentação e avaliação de métodos de seleção de características. In Anais Estendidos do XXIV Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 121-128. SBC. DOI: 10.5753/sbseg_estendido.2024.243363.
Salehin, I., Islam, M. S., Saha, P., Noman, S., Tuni, A., Hasan, M. M., and Baten, M. A. (2024). AutoML: A systematic review on automated machine learning with neural architecture search. Journal of Information and Intelligence, 2(1):52-81. DOI: 10.1016/j.jiixd.2023.10.002.
Santu, S. K. K., Hassan, M. M., Smith, M. J., Xu, L., Zhai, C., and Veeramachaneni, K. (2021). AutoML to date and beyond: Challenges and opportunities. ACM Computing Surveys. DOI: 10.1145/3470918.
Schaller, M., Kruse, M., Ortega, A., Lindauer, M., and Rosenhahn, B. (2025). AutoML for multi-class anomaly compensation of sensor drift. Measurement, 250:117097. DOI: 10.1016/j.measurement.2025.117097.
Singh, A., Patel, S., Bhadani, V., Kumar, V., and Gaurav, K. (2024). Automl-gwl: automated machine learning model for the prediction of groundwater level. Engineering Applications of Artificial Intelligence, 127:107405. DOI: 10.1016/j.engappai.2023.107405.
Sırt, M. and Eyüpoğlu, C. (2025). Comprehensive benchmarking analysis for evaluating effectiveness of transfer learning-based feature engineering in automl. The Journal of Cognitive Systems, 9(2):30-37. DOI: 10.52876/jcs.1604889.
SISTO, A. (2012). Androcrawl: studying alternative android marketplaces. Available at: [link].
Tian, J. and Che, C. (2024). Automated machine learning: A survey of tools and techniques. Journal of Industrial Engineering and Applied Science, 2(6):71-76. DOI: 10.70393/6a69656173.323336.
Trirat, P., Jeong, W., and Hwang, S. J. (2024). Automl-agent: A multi-agent llm framework for full-pipeline automl. arXiv preprint arXiv:2410.02958. DOI: 10.48550/arXiv.2410.02958.
Truong, A., Walters, A., Goodsitt, J., Hines, K., Bruss, C. B., and Farivar, R. (2019a). Towards automated machine learning: Evaluation and comparison of automl approaches and tools. arXiv preprint arXiv:1908.05557. DOI: 10.1109/ICTAI.2019.00209.
Truong, A., Walters, A., Goodsitt, J., Hines, K., Bruss, C. B., and Farivar, R. (2019b). Towards automated machine learning: Evaluation and comparison of AutoML approaches and tools. In IEEE 31st ICTAI. DOI: 10.1109/ICTAI.2019.00209.
Vakhrushev, A., Ryzhkov, A., Savchenko, M., Simakov, D., Damdinov, R., and Tuzhilin, A. (2021). Lightautoml: Automl solution for a large financial services ecosystem. arXiv preprint arXiv:2109.01528. DOI: 10.48550/arXiv.2109.01528.
Vilone, G. and Longo, L. (2020). Explainable artificial intelligence: a systematic review. arXiv preprint arXiv:2006.00093. DOI: 10.48550/arXiv.2006.00093 .
Weerts, H., Pfisterer, F., Feurer, M., Eggensperger, K., Bergman, E., Awad, N., Vanschoren, J., Pechenizkiy, M., Bischl, B., and Hutter, F. (2023). Can fairness be automated? guidelines and opportunities for fairness-aware AutoML. Journal of Artificial Intelligence Research. DOI: 10.1613/jair.1.14747.
Wever, M., Tornede, A., Mohr, F., and Hüllermeier, E. (2021). Automl for multi-label classification: Overview and empirical evaluation. IEEE transactions on pattern analysis and machine intelligence, 43(9):3037-3054. DOI: 10.1109/TPAMI.2021.3051276.
Wu, J., Chen, X.-Y., Zhang, H., Xiong, L.-D., Lei, H., and Deng, S.-H. (2019). Hyperparameter optimization for machine learning models based on bayesian optimization. Journal of Electronic Science and Technology. DOI: 10.11989/JEST.1674-862X.80904120.
Wu, J., Wang, H., Ni, C., Zhang, C., and Lu, W. (2024). Data pipeline training: Integrating automl to optimize the data flow of machine learning models. In 2024 7th International Conference on Advanced Algorithms and Control Engineering (ICAACE), pages 730-734. IEEE. DOI: 10.1109/ICAACE61206.2024.10549260.
Yang, L. and Shami, A. (2020). On hyperparameter optimization of machine learning algorithms: Theory and practice. Neurocomputing, 415:295-316. DOI: 10.1016/j.neucom.2020.07.061.
Ye, T., Meng, J., Xiao, Y., Lu, Y., Zheng, A., and Liang, B. (2025). Integrated automl-based framework for optimizing shale gas production: A case study of the fuling shale gas field. Energy Geoscience, 6(1):100365. DOI: 10.1016/j.engeos.2024.100365.
Yerima, S. Y. and Sezer, S. (2018). Droidfusion: A novel multilevel classifier fusion approach for android malware detection. IEEE transactions on cybernetics, 49(2):453-466. DOI: 10.1109/TCYB.2017.2777960.
Yuan, H., Yu, K., Xie, F., Liu, M., and Sun, S. (2024). Automated machine learning with interpretation: a systematic review of methodologies and applications in healthcare. Medicine Advances, 2(3):205-237. DOI: 10.1002/med4.75.
Zhang, X., Zhai, J., Ma, S., Guan, X., and Shen, C. (2025). Dream: Debugging and repairing automl pipelines. ACM Transactions on Software Engineering and Methodology, 34(4):1-29. DOI: 10.1145/3702992.
Zheng, R., Qu, L., Cui, B., Shi, Y., and Yin, H. (2023). Automl for deep recommender systems: A survey. ACM Transactions on Information Systems, 41(4):1-38. DOI: 10.1145/3579355.
Zimmer, L., Lindauer, M., and Hutter, F. (2020). Auto-pytorch tabular: Multi-fidelity metalearning for efficient and robust autodll. arxiv 2020. arXiv preprint arXiv:2006.13799. DOI: 10.48550/arXiv.2006.13799.
Zimmer, L., Lindauer, M., and Hutter, F. (2021). Auto-pytorch: multi-fidelity metalearning for efficient and robust autodl. IEEE Trans. on Pattern Analysis and MI, 43(9):3079-3090. DOI: 10.1109/TPAMI.2021.3067763.
Zöller, M.-A. and Huber, M. F. (2021). Benchmark and survey of automated machine learning frameworks. Journal of artificial intelligence research, 70:409-472. DOI: 10.48550/arXiv.1904.12054 .
Zöller, M.-A., Titov, W., Schlegel, T., and Huber, M. F. (2023). Xautoml: a visual analytics tool for understanding and validating automated machine learning. ACM Transactions on Interactive Intelligent Systems, 13(4):1-39. DOI: 10.1145/3625240.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Joner Assolin, Gabriel Canto, Diego Kreutz, Hendrio Bragança, Eduardo Feitosa , Angelo Nogueira, Vanderson Rocha

This work is licensed under a Creative Commons Attribution 4.0 International License.

