SinoTechIntel Academic Portal
Official PDF TranslationFrontiers of Information Technology & Electronic Engineering

Training large-scale language models with limited GPU memory: a survey

Authors: Yu Tang; Linbo Qiao; Lujia Yin; Peng Liang; Ao Shen; Zhilin Yang; Lizhi Zhang; Dongsheng Li

DOI: 10.1631/FITEE_2300710Status: Verified Translated Edition
Sponsored AdvertisementAd Placement Area
reCAPTCHA Bot Shield Active

Preparing Secure Academic Download

Verifying human reader & generating high-resolution document...

Verifying Document Integrity15s remaining
← Back to Article
Protected by Google reCAPTCHA v3.PrivacyTerms
Sponsored ContentAdSense In-Feed Ad Slot

Key Findings in This Report

• Identifies the three primary GPU memory consumers in large-scale model training: model parameters, model states, and model activations. • Provides a systematic overview of memory optimization techniques, including parallelism, offloading, and activation checkpointing, tailored for limited GPU memory. • Addresses the GPU memory wall problem, highlighting the gap between exponential parameter growth and linear memory capacity increase. • Concludes with future research directions, advocating for continued innovation in memory-efficient training methods for large-scale language models.