Key Takeaways & Executive Findings
- •• Proposes CRGT-SA, a hybrid deep learning model integrating CNN, LSTM, gated TCN, and self-attention for network intrusion detection. • Achieves state-of-the-art performance with 91.5% binary and 90.5% multi-class accuracy on the UNSW-NB15 dataset, outperforming traditional and deep learning baselines. • Demonstrates strong generalization ability through additional validation on the NSL-KDD dataset. • Overcomes limitations of shallow machine learning by automatically extracting spatiotemporal features and selecting significant attributes via self-attention.
Abstract
To address the challenge of cyberattacks, intrusion detection systems (IDSs) are introduced to recognize intrusions and protect computer networks. Among all these IDSs, conventional machine learning methods rely on shallow learning and have unsatisfactory performance. Unlike machine learning methods, deep learning methods are the mainstream methods because of their capability to handle mass data without prior knowledge of specific domain expertise. Concerning deep learning, long short-term memory (LSTM) and temporal convolutional networks (TCNs) can be used to extract temporal features from different angles, while convolutional neural networks (CNNs) are valuable for learning spatial properties. Based on the above, this paper proposes a novel interlaced and spatiotemporal deep learning model called CRGT-SA, which combines CNN with gated TCN and recurrent neural network (RNN) modules to learn spatiotemporal properties, and imports the self-attention mechanism to select significant features. More specifically, our proposed model splits the feature extraction into multiple steps with a gradually increasing granularity, and executes each step with a combined CNN, LSTM, and gated TCN module. Our proposed CRGT-SA model is validated using the UNSW-NB15 dataset and is compared with other compelling techniques, including traditional machine learning and deep learning models as well as state-of-the-art deep learning models. According to the simulation results, our proposed model exhibits the highest accuracy and F1-score among all the compared methods. More specifically, our proposed model achieves 91.5% and 90.5% accuracy for binary and multi-class classifications respectively, and demonstrates its ability to protect the Internet from complicated cyberattacks. Moreover, we conduct another series of simulations on the NSL-KDD dataset; the simulation results of comparison with other models further prove the generalization ability of our proposed model.
1. Introduction
The latest report from Juniper Research states that there will be 84 billion network connections in 2024 (Oseni et al., 2023). Meanwhile, cyberattacks are constantly evolving and pose a significant threat to a wide variety of cutting-edge technologies, such as smart hospitals, power, medical, and campuses (Lv et al., 2021). For example, attacks against critical infrastructures used for power generation can lead to loss of property or even personal safety (Cook et al., 2020). Intrusion detection systems (IDSs) are responsible for identifying intrusions that evade security technique, and providing a vital second level of resistance to protect computer networks (Fang et al., 2021).
Multiple recent IDS studies adopt machine learning and deep learning methods to solve security issues (Al-Garadi et al., 2020; Chen et al., 2022). Conventional machine learning methods include logistic regression (LR), Gaussian naive Bayes (GNB), K-nearest neighbors (KNN), adaptive boosting (AdaB), and random forest (RF), and almost all the methods mentioned above use shallow learning which relies on manual feature engineering to extract features (Vasilomanolakis et al., 2015). Because of the huge amount of data, shallow learning is unable to solve real-time problems (Laghrissi et al., 2021). As a result, these methods achieve unsatisfactory performance in identifying different types of cyberattacks (Wang K et al., 2023).
Unlike machine learning methods, deep learning methods have become the most dominant roles in the field of intrusion detection, because of their ability to deal with mass data without prior knowledge of specific domain expertise. For instance, convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and temporal convolutional networks (TCNs), as standard deep learning technologies, can deal with network intrusions with different degrees of difficulty, complexity, and distributivity (Wang XF et al., 2020). To achieve high accuracy in detecting and classifying different types of cyberattacks, this study proposes a new network intrusion detection approach with a hybrid deep learning model.
Loading authentic research manuscript (Pages 1–5)...
Jue CHEN, Wanxiao LIU, Xihe QIU, Wenjing LV, Yujie XIONG (2025). CRGT-SA: an interlaced and spatiotemporal deep learning model for network intrusion detection. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400459
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is CRGT-SA?
CRGT-SA is a novel interlaced and spatiotemporal deep learning model for network intrusion detection that combines CNN, LSTM, gated TCN, and self-attention mechanisms to extract spatiotemporal features and select significant attributes.
What datasets were used to validate the model?
The model was validated on the UNSW-NB15 dataset and the NSL-KDD dataset to demonstrate its performance and generalization ability.
What accuracy does CRGT-SA achieve?
On UNSW-NB15, it achieves 91.5% accuracy for binary classification and 90.5% for multi-class classification, outperforming traditional machine learning and state-of-the-art deep learning models.
Why is deep learning preferred over machine learning for intrusion detection?
Deep learning methods can handle massive data without prior domain expertise, automatically learning hierarchical features, unlike shallow machine learning methods which rely on manual feature engineering and often show unsatisfactory performance.
How does CRGT-SA differ from other hybrid deep learning models?
It uniquely interlaced CNN, LSTM, and gated TCN modules in multiple steps with gradually increasing granularity and integrates self-attention for feature selection, enhancing detection accuracy and generalization.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena