SinoTechIntel Academic Portal
Official PDF TranslationFrontiers of Information Technology & Electronic Engineering

Memory-efficient tensor parallelism for long-sequence Transformer training

Authors: Peng LIANG; Linbo QIAO; Yanqi SHI; Hao ZHENG; Yu TANG; Dongsheng LI

DOI: 10.1631/FITEE_2400602Status: Verified Translated Edition
Sponsored AdvertisementAd Placement Area
reCAPTCHA Bot Shield Active

Preparing Secure Academic Download

Verifying human reader & generating high-resolution document...

Verifying Document Integrity15s remaining
← Back to Article
Protected by Google reCAPTCHA v3.PrivacyTerms
Sponsored ContentAdSense In-Feed Ad Slot

Key Findings in This Report

• METP avoids duplicated tensors and uses send/recv communication instead of collective communication, improving memory efficiency. • With double buffering, METP achieves effective overlap between computation and communication, with a theoretical condition for full overlap. • METP provides O(1/p^3) memory overhead without FlashAttention and saves at least 41.7% memory compared to TP when using FlashAttention. • On eight A100 GPUs, METP increases the maximum sequence length by 2.38–2.99 times versus other parallelism methods.