Key Takeaways & Executive Findings
- •• Comprehensive analysis decomposes transformer inference overhead to identify primary bottlenecks across operator, runtime, and control layers. • Three-tier scheduling framework on Shenwei AI accelerator MPE cuts host-device launches to approximately 1/10,000 of the original PyTorch-GPU setup. • Zero-copy memory management with segment-page fusion substantially reduces memory access latency and improves overall inference efficiency. • Fast model loading eliminates redundant verification and initialization computations, slashing large-model loading time from 22,128.31 ms to 1,041.72 ms.
Abstract
Transformer models have become a cornerstone of various natural language processing (NLP) tasks. However, the substantial computational overhead during the inference remains a significant challenge, limiting their deployment in practical applications. In this study, we address this challenge by minimizing the inference overhead in transformer models using the controlling element on artificial intelligence (AI) accelerators. Our work is anchored by four key contributions. First, we conduct a comprehensive analysis of the overhead composition within the transformer inference process, identifying the primary bottlenecks. Second, we leverage the management processing element (MPE) of the Shenwei AI (SWAI) accelerator, implementing a three-tier scheduling framework that significantly reduces the number of host-device launches to approximately 1/10 000 of the original PyTorch-GPU setup. Third, we introduce a zero-copy memory management technique using segment-page fusion, which significantly reduces memory access latency and improves overall inference efficiency. Finally, we develop a fast model loading method that eliminates redundant computations during model verification and initialization, reducing the total loading time for large models from 22 128.31 ms to 1041.72 ms. Our contributions significantly enhance the optimization of transformer models, enabling more efficient and expedited inference processes on AI accelerators.
1. Introduction
Over the past decade, pre-trained language models based on the transformer architecture (Vaswani et al., 2017) have become a leading paradigm in the domain of natural language processing (NLP). Notable examples of such models include BERT (Devlin et al., 2019), GPT-2 (Radford et al., 2019), LLaMA (Touvron et al., 2023), wav2vec2.0 (Baevski et al., 2020), and whisper (Radford et al., 2023). These transformer models have significantly advanced state-of-the-art models in terms of accuracy, surpassing traditional models. Nonetheless, the computational intensity during the inference phase poses a significant barrier to their integration into real-world applications, which require strict criteria such as low latency, rapid inference capabilities, and cost-effective operational overhead.
There are currently two strategies for enhancing the inference efficiency of transformer models. The first pertains to the optimization of specific operators, such as Softermax (Stevens et al., 2021) and FLASHATTENTION (Dao et al., 2022), and involves reducing accuracy and minimizing input/output (I/O) costs to enhance the operators’ computational speed. The second focuses on the optimization of the model’s runtime during inference, using various techniques such as structured pruning (Kim YJ et al., 2020), knowledge distillation (Fang et al., 2021), operator fusion, memory management (Chen SY et al., 2021), and parallel scheduling strategies (Du et al., 2022). These approaches collectively aim to improve inference efficacy by reducing model size and enhancing overall system performance. However, the intrinsic computational overhead (Ma et al., 2021) associated with transformer inference has received relatively little scholarly attention, despite its critical role and the need for comprehensive research in this area.
To minimize the overhead during the inference process, we propose an innovative approach that uses the hardware control mechanisms of artificial intelligence (AI) accelerators. We begin by conducting a detailed analysis of the transformer model’s inference control process, exploring the components of the runtime overhead through carefully designed experiments. Based on this analysis, we introduce a fast loading method that significantly reduces the overhead caused by loading models with PyTorch-GPU. We also develop a three-tier scheduling framework for interacting with the host and the controlling element on the accelerator, leveraging the accelerator’s scheduling capabilities. To minimize device’s launch expenditures, we embrace a holistic full-graph optimization strategy. Additionally, we implement a zero-copy memory management protocol based on segment-page fusion, which eliminates data transmission-related costs. These optimizations target the reduction of overhead throughout the inference process.
Loading authentic research manuscript (Pages 1–5)...
Yulong ZHAO, Chunzhi WU, Yizhuo WANG, Lufei ZHANG, Yaguang ZHANG, Wenyuan SHEN, Hao FAN, Hankang FANG, Yi QIN, Xin LIU (2025). Minimizing transformer inference overhead using controlling element on Shenwei AI accelerator. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400453
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the main contribution of this paper?
The paper minimizes transformer inference overhead on Shenwei AI accelerators through four key contributions: a comprehensive overhead composition analysis, a three-tier scheduling framework using the management processing element, a zero-copy memory management technique based on segment-page fusion, and a fast model loading method that eliminates redundant computations.
How does the three-tier scheduling framework reduce inference overhead?
The three-tier scheduling framework leverages the management processing element (MPE) of the Shenwei AI accelerator to significantly reduce the number of host-device launches to approximately 1/10,000 of the original PyTorch-GPU setup, thereby minimizing launch expenditures.
What is zero-copy memory management and how does it improve efficiency?
Zero-copy memory management is a technique using segment-page fusion that eliminates data transmission-related costs and reduces memory access latency, thereby improving the overall inference efficiency of transformer models.
How much faster is the proposed model loading method compared to PyTorch-GPU?
The proposed fast model loading method reduces the total loading time for large models from 22,128.31 ms to 1,041.72 ms, representing a substantial improvement by eliminating redundant computations during model verification and initialization.
Why is minimizing transformer inference overhead important?
Transformer models have high computational intensity during inference, which poses a significant barrier to real-world deployment. Minimizing overhead reduces latency, supports rapid inference, and lowers operational costs, making these models more practical for production applications.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena