Key Takeaways & Executive Findings
- •• Identifies the three primary GPU memory consumers in large-scale model training: model parameters, model states, and model activations. • Provides a systematic overview of memory optimization techniques, including parallelism, offloading, and activation checkpointing, tailored for limited GPU memory. • Addresses the GPU memory wall problem, highlighting the gap between exponential parameter growth and linear memory capacity increase. • Concludes with future research directions, advocating for continued innovation in memory-efficient training methods for large-scale language models.
Abstract
Large-scale models have gained significant attention in a wide range of fields, such as computer vision and natural language processing, due to their effectiveness across various applications. However, a notable hurdle in training these large-scale models is the limited memory capacity of graphics processing units (GPUs). In this paper, we present a comprehensive survey focused on training large-scale models with limited GPU memory. The exploration commences by scrutinizing the factors that contribute to the consumption of GPU memory during the training process, namely model parameters, model states, and model activations. Following this analysis, we present an in-depth overview of the relevant research work that addresses these aspects individually. Finally, the paper concludes by presenting an outlook on the future of memory optimization in training large-scale language models, emphasizing the necessity for continued research and innovation in this area. This survey serves as a valuable resource for researchers and practitioners keen on comprehending the challenges and advancements in training large-scale language models with limited GPU memory.
1. Introduction
With the advent of deep learning (LeCun et al., 2015), neural networks have witnessed significant advancements across diverse domains, such as speech recognition (Povey et al., 2011; Ze et al., 2013; Cho et al., 2014), computer vision (Krizhevsky et al., 2012; Ji et al., 2013; Ren SQ et al., 2015), and natural language processing (Chowdhury, 2003; Sutskever et al., 2014; Dong et al., 2019). The landscape changed dramatically in 2017 with the introduction of Transformer (Vaswani et al., 2017), which surpassed conventional neural network models in language translation tasks, capturing widespread attention from various fields (Kitaev et al., 2020; Liu Z et al., 2021; Han K et al., 2023). Large-scale models, such as BERT (Devlin et al., 2019), ERNIE (Sun Y et al., 2019), T5 (Raffel et al., 2020), and GPT-3 (Brown et al., 2020; Zhou J et al., 2024), which are generally built on the Transformer model, have demonstrated state-of-the-art performance in a wide range of applications, including reading comprehension, question answering (Rajpurkar et al., 2016), and adversarial generations (Zellers et al., 2018).
Recent studies have consistently shown that larger models with increased parameter sizes exhibit improved capacity and representation capabilities (Qiu et al., 2020). Consequently, there has been a rapid escalation in the size of these models, necessitating the need for memory optimization techniques during their training (Sun Y et al., 2021; Zeng et al., 2021). The exponential growth of the number of model parameters places high demands for graphics processing unit (GPU) memory. Unfortunately, the linear increase of GPU memory capacity occurring recently cannot meet the requirements of the exponential growth of the number of model parameters. Therefore, the further advancement of large-scale models is substantially hindered by the limited GPU memory, i.e., the GPU memory wall problem (Gholami et al., 2024; Rajbhandari et al., 2021), as shown in Fig. 1, posing a significant impediment to progress in this domain. The demand for GPU memory remains a crucial factor when training large-scale models. For instance, training a model with more than one trillion parameters necessitates the usage of more than 180 A100 GPUs, each equipped with 80 GB of memory. Therefore, training large-scale models with limited GPU memory is a pressing research challenge that demands further advancements.
Training large-scale models with limited GPU memory has been an urgent issue to handle. Liang et al. (2022) presented a survey that introduced auto parallelism in detail. They gave a comprehensive understanding and analysis of parallel and distributed training of deep neural networks (DNNs). Gusak et al. (2022) introduced the techniques of large-scale model training, including memory optimization methods. They summarized the main strategies contributing to training large-scale models. However, their comparison is not enough. Different from them, we present a comprehensive analysis aimed at identifying the key factors that consume GPU memory during the training of DNNs and providing a survey about how to train large-scale models with limited GPU memory. Drawing upon our analysis, a thorough study of the prevalent techniques that are employed in training large-scale models with limited GPU memory is presented. This study delves into the significant memory consumption caused by model parameters, model states, and model activations.
Loading authentic research manuscript (Pages 1–5)...
Yu Tang, Linbo Qiao, Lujia Yin, Peng Liang, Ao Shen, Zhilin Yang, Lizhi Zhang, Dongsheng Li (2025). Training large-scale language models with limited GPU memory: a survey. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2300710
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the GPU memory wall problem in large-scale model training?
The GPU memory wall problem refers to the fundamental mismatch between the exponential growth of model parameters and the relatively linear increase in GPU memory capacity. This makes it increasingly difficult to fit large-scale models into available GPU memory, hindering further advancements in the field.
What are the main factors that consume GPU memory during training?
The primary factors are model parameters, model states (including optimizer states and gradients), and model activations. These components collectively contribute to the high memory demand during training of large-scale models.
What techniques are surveyed for training large-scale models with limited GPU memory?
The survey provides an in-depth overview of memory optimization techniques addressing the three main factors. It covers methods such as parameter offloading, gradient checkpointing, mixed-precision training, and various parallelism strategies (data, model, pipeline, and auto parallelism).
How can large-scale language models be trained with limited GPU memory?
By adopting a combination of memory optimization techniques, including efficiently managing model parameters, states, and activations, as well as leveraging parallel and distributed training strategies. The survey offers a comprehensive analysis of these approaches and their trade-offs.
What are the future research directions for memory optimization in training large-scale models?
The paper emphasizes the need for continued research and innovation, particularly in developing new algorithms and architectures that reduce memory footprint, improving memory management at the system level, and exploring novel offloading and compression techniques to bridge the gap between model scale and hardware capacity.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena