Key Takeaways & Executive Findings
- •• Introduces AOCQ, a three-level quantization method (operator, framework, loss) that adaptively corrects channel and token outliers in vision Transformers, reducing quantization error. • Achieves 81.57% accuracy on DeiT-Base with 8-bit post-training quantization, only a 0.28 percentage point drop, and 4× faster runtime. • Enables ultra-low 4-bit weight quantization for Swin and DeiT across classification and object detection tasks, with a minimal accuracy loss of about 2% and nearly 8× less memory. • Demonstrates that AOCQ effectively mitigates the uneven activation distributions that limit standard PTQ methods, supporting efficient edge deployment.
Abstract
Transformers have demonstrated considerable success across various domains but are constrained by their significant computational and memory requirements. This poses challenges for deployment on resource-constrained devices. Quantization, as an effective model compression method, can significantly reduce the operational time of Transformers on edge devices. Notably, Transformers display more substantial outliers than convolutional neural networks, leading to uneven feature distribution among different channels and tokens. To address this issue, we propose an adaptive outlier correction quantization (AOCQ) method for Transformers, which significantly alleviates the adverse effects of these outliers. AOCQ adjusts the notable discrepancies in channels and tokens across three levels: operator level, framework level, and loss level. We introduce a new operator that equivalently balances the activations across different channels and insert an extra stage to optimize the activation quantization step on the framework level. Additionally, we transfer the imbalanced activations across tokens and channels to the optimization of model weights on the loss level. Based on the theoretical study, our method can reduce the quantization error. The effectiveness of the proposed method is verified on various benchmark models and tasks. Surprisingly, DeiT-Base with 8-bit post-training quantization (PTQ) can achieve 81.57% accuracy with a 0.28 percentage point drop while enjoying 4× faster runtime. Furthermore, the weights of Swin and DeiT on several tasks, including classification and object detection, can be post-quantized to ultra-low 4 bits, with a minimal accuracy loss of 2%, while requiring nearly 8× less memory.
1. Introduction
Transformer-based architectures (Vaswani et al., 2017) have shown great power in natural language processing (NLP) tasks (Choi et al., 2018; Devlin et al., 2019). Increasingly, vision Transformers (ViTs) have also achieved competitive performance on many computer vision (CV) tasks including image classification, object detection, object segmentation, and other vision tasks recently (Carion et al., 2020; Chen ZS et al., 2021; Dosovitskiy et al., 2021; Graham et al., 2021; Yuan L et al., 2021; Touvron et al., 2021a; Dong et al., 2022; Liu Z et al., 2022b; Yu et al., 2022). Transformers consist of a number of blocks containing multi-head self-attention (MHSA) and feed-forward networks (FFNs), which enables the extraction of highly discriminative features. However, these Transformer-based models are notable for their substantial computational intensity and extensive memory requirements, posing significant challenges for deployment on resource-constrained devices (Alam et al., 2023; Chitty-Venkata et al., 2023). Consequently, there is an urgent industry requirement to compress and accelerate these Transformer-based models to facilitate broader application.
Much effort has been invested in facilitating the deployment of Transformers, including pruning (Han S et al., 2015; Zheng et al., 2022), distillation (Touvron et al., 2021b), quantization (Yao et al., 2022), and the direct design of more efficient Transformers (Choromanski et al., 2021; Yang et al., 2022). Among these methods, quantization employs low-bit precision for weight and activation values without altering the model architecture, making it particularly suitable for carefully designed efficient Transformers. There are primarily two types of quantization methods: post-training quantization (PTQ) and quantization-aware training (QAT) (Chitty-Venkata et al., 2023). Unlike the QAT method, which necessitates the entire training dataset, PTQ only requires unlabeled calibration images, thereby enabling rapid quantization and deployment. Consequently, our focus is on PTQ methods to compress Transformers. Besides, quantization and other techniques (Han S et al., 2015; Touvron et al., 2021b; Yang et al., 2022) are complementary for model acceleration.
Previous studies, such as ranking-aware (Liu ZH et al., 2021), adopt a ranking loss to make the order of the self-attention results after quantization as consistent as possible. Fully quantized vision Transformer (FQ-ViT) (Lin et al., 2022) observes serious inter-channel variation in LayerNorm inputs and extreme non-uniform distributions in attention maps, and thus uses power-of-two factor (PTF) and log-int-softmax (LIS) to reduce the performance degradation. Scale reparameterization for post-training quantization of vision Transformers (RepQ-ViT) (Li ZK et al., 2023) uses channel-wise quantization to deal with the imbalance between channels and uses a log√2 quantizer to compress the power-law features of the latter.
Loading authentic research manuscript (Pages 1–5)...
Zheyang LI, Chaoxiang LAN, Kai ZHANG, Wenming TAN, Ye REN, Jun XIAO (2025). An adaptive outlier correction quantization method for vision Transformers. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400994
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is AOCQ?
AOCQ (Adaptive Outlier Correction Quantization) is a post-training quantization method for vision Transformers that mitigates activation outliers by adjusting channel and token discrepancies at operator, framework, and loss levels, reducing quantization error while preserving accuracy.
How does AOCQ improve Transformer quantization?
AOCQ introduces a channel-balancing operator, optimizes the activation quantization step at the framework level, and transfers imbalanced activations to weight optimization at the loss level, collectively reducing quantization error and enabling fast, low-bit deployment.
What accuracy does DeiT-Base achieve with 8-bit PTQ?
DeiT-Base achieves 81.57% accuracy with 8-bit post-training quantization, only a 0.28 percentage point drop, with 4× faster runtime.
Can AOCQ quantize weights to 4 bits?
Yes, Swin and DeiT weights can be post-quantized to ultra-low 4 bits on tasks including classification and object detection, with a minimal accuracy loss of about 2% and nearly 8× less memory.
Why are outliers a problem in vision Transformers?
Vision Transformers exhibit larger outliers than CNNs, causing uneven feature distributions among channels and tokens, which can significantly degrade the performance of standard quantization methods; AOCQ specifically corrects these outliers.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena