RISK-AWARE CYBERATTACK DETECTION IN WIRELESS NETWORKS CONSIDERING THE ROBUSTNESS OF MACHINE LEARNING MODELS
DOI:
https://doi.org/10.28925/2663-4023.2026.34.1333Keywords:
cybersecurity, wireless networks, intrusion detection system, machine learning, adversarial attack, SHAP, selective classification, risk assessmentAbstract
Intrusion detectors for wireless networks are still commonly assessed on unchanged test sets. Such evaluation does not cover a deliberate modification of controllable traffic features or the use of model explanations to identify a vulnerable direction. A high test score therefore says little about detector behaviour under an active attack. This paper examines a risk-aware pipeline that combines a gradient-boosted base classifier, SHAP explanations, a logistic risk meta-classifier, and selective referral to a human analyst. Risk is described by three signals: the base attack probability, cosine stability of a local SHAP explanation relative to a class-specific training centroid, and Mahalanobis distance from the training distribution. This combination is intended to capture not only the predicted class, but also whether the reasoning pattern and the record itself are familiar to the model. A fixed five percent of records is referred for review according to explanation instability, predictive uncertainty, or predictive entropy. The experiments use a controlled synthetic dataset and NSL-KDD. Two attacks are evaluated: a one-step, gradient-free SHAP-sign perturbation and a transferred projected-gradient attack. The results are not uniform across datasets. On the synthetic benchmark, uncertainty and entropy provide the best selective results; under transferred PGD, macro-F1 on the retained 95 percent rises from 0.882 ± 0.008 for the base detector to 0.908 ± 0.008. NSL-KDD gives a different ordering. Explanation stability is more useful on clean and SHAP-sign-perturbed records, whereas under transferred PGD neither selective configuration improves on C3: macro-F1 is 0.699 for C3, 0.696 for C4, and 0.694 for C4-conf/ent. The base detector also deteriorates markedly, from 0.791 ± 0.003 on clean NSL-KDD to 0.579 ± 0.004 under SHAP-sign and 0.637 ± 0.002 under transferred PGD. These results do not support choosing a review rule from clean validation data alone. Selective referral can manage the accuracy–coverage trade-off, but it is not an adversarial defence. Before deployment, the pipeline should be tested on current Wi-Fi, IoT, 5G, and 6G traffic, with semantically valid perturbations, adaptive attacks, and review rates matched to analyst capacity.
Downloads
References
Adesina, D., Hsieh, C.-C., Sagduyu, Y. E., & Qian, L. (2023). Adversarial machine learning in wireless communications using RF data: A review. IEEE Communications Surveys & Tutorials, 25(1), 77–100. https://doi.org/10.1109/COMST.2022.3205184
He, K., Kim, D. D., & Asghar, M. R. (2023). Adversarial machine learning for network intrusion detection systems: A comprehensive survey. IEEE Communications Surveys & Tutorials, 25(1), 538–566. https://doi.org/10.1109/COMST.2022.3233793
Senevirathna, T., Siniarski, B., Liyanage, M., & Wang, S. (2024). Deceiving post-hoc explainable AI (XAI) methods in network intrusion detection. In 2024 IEEE 21st Consumer Communications & Networking Conference (CCNC) (pp. 107-112). IEEE. https://doi.org/10.1109/CCNC51664.2024.10454633
Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41-46. https://doi.org/10.1109/TIT.1970.1054406
Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4878-4887). NeurIPS proceedings
Mohale, V. Z., & Obagbuwa, I. C. (2025). Evaluating machine learning-based intrusion detection systems with explainable AI: Enhancing transparency and interpretability. Frontiers in Computer Science, 7, Article 1520741. https://doi.org/10.3389/fcomp.2025.1520741
Arreche, O., Guntur, T. R., Roberts, J. W., & Abdallah, M. (2024). E-XAI: Evaluating black-box explainable AI frameworks for network intrusion detection. IEEE Access, 12, 23954-23988. https://doi.org/10.1109/ACCESS.2024.3365140
Birahim, S. A., Paul, A., Rahman, F., Islam, Y., Roy, T., Hasan, M. A., Haque, F., & Chowdhury, M. E. H. (2025). Intrusion detection for wireless sensor network using particle swarm optimization based explainable ensemble machine learning approach. IEEE Access, 13, 13711-13730. https://doi.org/10.1109/ACCESS.2025.3528341
Papernot, N., McDaniel, P., & Goodfellow, I. (2016). Transferability in machine learning: From phenomena to black-box attacks using adversarial samples [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1605.07277
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2018). Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations (ICLR).
Cohen, J. M., Rosenfeld, E., & Kolter, J. Z. (2019). Certified adversarial robustness via randomized smoothing. In Proceedings of the 36th International Conference on Machine Learning (Vol. 97, pp. 1310-1320). PMLR. PMLR proceedings
Tavallaee, M., Bagheri, E., Lu, W., & Ghorbani, A. A. (2009). A detailed analysis of the KDD CUP 99 data set. In 2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications (pp. 1-6). IEEE. https://doi.org/10.1109/CISDA.2009.5356528
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765-4774). NeurIPS proceedings
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135-1144). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778
Athalye, A., Carlini, N., & Wagner, D. (2018). Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In Proceedings of the 35th International Conference on Machine Learning (Vol. 80, pp. 274-283). PMLR. PMLR proceedings
Biggio, B., & Roli, F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognition, 84, 317-331. https://doi.org/10.1016/j.patcog.2018.07.023
Kim, B., Sagduyu, Y. E., Davaslioglu, K., Erpek, T., & Ulukus, S. (2022). Channel-aware adversarial attacks against deep learning-based wireless signal classifiers. IEEE Transactions on Wireless Communications, 21(6), 3868-3880. https://doi.org/10.1109/TWC.2021.3124855
Zhang, W., Krunz, M., & Ditzler, G. (2024). Stealthy adversarial attacks on machine learning-based classifiers of wireless signals. IEEE Transactions on Machine Learning in Communications and Networking, 2, 261-279. https://doi.org/10.1109/TMLCN.2024.3366161
Kolias, C., Kambourakis, G., Stavrou, A., & Gritzalis, S. (2016). Intrusion detection in 802.11 networks: Empirical evaluation of threats and a public dataset. IEEE Communications Surveys & Tutorials, 18(1), 184-208. https://doi.org/10.1109/COMST.2015.2402161
Tabassi, E. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ольга Партика

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.