SinoTechIntel Academic Portal
Official PDF TranslationFrontiers of Information Technology & Electronic Engineering

Minimizing transformer inference overhead using controlling element on Shenwei AI accelerator

Authors: Yulong ZHAO; Chunzhi WU; Yizhuo WANG; Lufei ZHANG; Yaguang ZHANG; Wenyuan SHEN; Hao FAN; Hankang FANG; Yi QIN; Xin LIU

DOI: 10.1631/FITEE_2400453Status: Verified Translated Edition
Sponsored AdvertisementAd Placement Area
reCAPTCHA Bot Shield Active

Preparing Secure Academic Download

Verifying human reader & generating high-resolution document...

Verifying Document Integrity15s remaining
← Back to Article
Protected by Google reCAPTCHA v3.PrivacyTerms
Sponsored ContentAdSense In-Feed Ad Slot

Key Findings in This Report

• Comprehensive analysis decomposes transformer inference overhead to identify primary bottlenecks across operator, runtime, and control layers. • Three-tier scheduling framework on Shenwei AI accelerator MPE cuts host-device launches to approximately 1/10,000 of the original PyTorch-GPU setup. • Zero-copy memory management with segment-page fusion substantially reduces memory access latency and improves overall inference efficiency. • Fast model loading eliminates redundant verification and initialization computations, slashing large-model loading time from 22,128.31 ms to 1,041.72 ms.