Key Takeaways & Executive Findings
- •• E-CGL introduces a combined replay sampling strategy based on node importance and diversity to mitigate catastrophic forgetting in continual graph learning. • By sharing weights between a GNN and a lightweight MLP, E-CGL bypasses expensive message-passing, yielding average training and inference speedups of 15.83× and 4.89×. • The method achieves state-of-the-art results on four CGL datasets, reducing average catastrophic forgetting to −1.1%. • E-CGL effectively handles topological interdependencies between sequential graph snapshots while maintaining efficiency at scale.
Abstract
Continual learning (CL) has emerged as a crucial paradigm for learning from sequential data while retaining previous knowledge. Continual graph learning (CGL), characterized by dynamically evolving graphs from streaming data, presents distinct challenges that demand efficient algorithms to prevent catastrophic forgetting. The first challenge stems from the interdependencies between different graph data, in which previous graphs influence new data distributions. The second challenge is handling large graphs in an efficient manner. To address these challenges, we propose an efficient continual graph learner (E-CGL) in this paper. We address the interdependence issue by demonstrating the effectiveness of replay strategies and introducing a combined sampling approach that considers both node importance and diversity. To improve efficiency, E-CGL leverages a simple yet effective multi-layer perceptron (MLP) model that shares weights with a graph neural network (GNN) during training, thereby accelerating computation by circumventing the expensive message-passing process. Our method achieves state-of-the-art results on four CGL datasets under two settings, while significantly lowering the catastrophic forgetting value to an average of −1.1%. Additionally, E-CGL achieves the training and inference speedup by an average of 15.83× and 4.89×, respectively, across four datasets. These results indicate that E-CGL not only effectively manages correlations between different graph data during continual training but also enhances efficiency in large-scale CGL.
1. Introduction
Graphs have garnered significant research attention due to their ubiquity in modeling relational data (Kipf and Welling, 2017; Veličković et al., 2018; Xu et al., 2019). In real-world applications, graph data continuously expand, and new patterns emerge over time. For instance, recommendation networks may introduce new product categories, citation networks may see the rise of novel research topics, and chemical design may discover new molecules and drugs. To adapt to these evolving scenarios and provide up-to-date predictions, graph models must constantly learn and update. However, traditional training strategies (Niepert et al., 2016; Kipf and Welling, 2017) experience catastrophic forgetting (Yuan et al., 2023) when adapting to new data, resulting in poor performance on previous tasks.
As a result, methods that can rapidly adapt to new classes while maintaining performance on previous tasks are required. Continual learning (CL, also known as incremental learning or lifelong learning) aims to do this by learning from new data while retaining prior knowledge (Thrun, 1994). While CL has been extensively studied in areas such as computer vision (Rebuffi et al., 2017; Buzzega et al., 2020; Ni et al., 2023) and natural language processing (Srinivasan et al., 2022; Fan et al., 2022), its application to graph-structured data presents unique challenges for a diverse set of downstream tasks, including supervised node classification and unsupervised dynamic network embedding (Choi et al., 2024; Gao et al., 2024; Wang ZZ et al., 2025).
The first challenge arises from the topological interdependence of graph snapshots, in which new tasks inherit structural and attributive dependencies from previous tasks. Topological interdependence occurs when sequential graph snapshots {G1, G2, ..., GT} share nodes/edges or have attribute continuity, resulting in conditional dependencies between their distributions. Specifically, in continual graph learning (CGL), the conditional likelihood of observing Gt depends on previous graphs: p(θ|G1:t) ∝ p(Gt|G1:t−1, θ)p(θ|G1:t−1)/p(Gt), where p(θ|G1:t) refers to the posterior distribution of the model parameter θ after observing all snapshots from round 1 to t, and p(θ|G1:t−1) is an estimate of θ when only the first t−1 snapshots are observed.
Loading authentic research manuscript (Pages 1–5)...
Jianhao Guo, Zixuan Ni, Yun Zhu, Siliang Tang (2025). E-CGL: an efficient continual graph learner. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2500162
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is continual graph learning (CGL)?
Continual graph learning (CGL) is a paradigm that enables graph models to learn from dynamically evolving graph snapshots in streaming data while retaining previously acquired knowledge. It addresses challenges such as catastrophic forgetting and the interdependencies that exist among sequential graph data.
How does E-CGL mitigate catastrophic forgetting?
E-CGL uses replay strategies combined with a sampling approach that considers both node importance and diversity. This allows the model to effectively manage correlations between different graph data during continual training, reducing average catastrophic forgetting to −1.1%.
What makes E-CGL efficient for large graphs?
E-CGL shares weights between a graph neural network (GNN) and a simple multi-layer perceptron (MLP) during training, thus circumventing the expensive message-passing process. This leads to average training and inference speedups of 15.83× and 4.89× across four datasets.
What are the main challenges addressed by E-CGL?
The two main challenges are the topological interdependence of sequential graph snapshots—where new tasks inherit structural and attributive dependencies from previous tasks—and the need to handle large graphs efficiently without sacrificing performance.
In what settings was E-CGL evaluated?
E-CGL was evaluated on four continual graph learning datasets under two settings. It achieved state-of-the-art results while significantly lowering catastrophic forgetting and improving both training and inference efficiency.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena