SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2400867Original Research

SAPER-AI accelerator: a systolic array-based power-efficient reconfigurable AI accelerator

Fahad Bin Muslim¹,Kashif Inayat¹,Muhammad Zain Siddiqi¹,Safiullah Khan¹,Tayyeb Mahmood¹,Ihtesham ul Islam¹

Faculty of Computer Science and Engineering, GIK Institute, Topi 23460, Pakistan

Read Executive PreviewQuick FAQ
SAPER-AI accelerator: a systolic array-based power-efficient reconfigurable AI accelerator
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:January 11, 2025Edition:Vol. 32, Issue 1 • pp. 540-552Citation:Fahad Bin Muslim et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:low-power designdeep learning

Key Takeaways & Executive Findings

  • • SAPER-AI achieves 10–25% power efficiency improvement for 32×32 and 64×64 systolic array configurations through coarse-grained row/column PE deactivation. • The accelerator utilizes Unified Power Format (UPF) for simplified power intent specification, enabling rapid design with negligible microarchitectural optimization effort. • Power-delay product (PDP) improves by approximately 6% for larger systolic array sizes, demonstrating the scalability of the proposed power management technique. • ResNet50 workloads consistently outperform MobileNet on systolic arrays due to more regular convolution patterns, with the gap widening as array size increases.
Sponsored Research Highlight

Abstract

Deep learning (DL) accelerators are critical for handling the growing computational demands of modern neural networks. Systolic array (SA)-based accelerators consist of a 2D mesh of processing elements (PEs) working cooperatively to accelerate matrix multiplication. The power efficiency of such accelerators is of primary importance, especially considering the edge AI regime. This work presents the SAPER-AI accelerator, an SA accelerator with power intent specified via a unified power format representation in a simplified manner with negligible microarchitectural optimization effort. Our proposed accelerator switches off rows and columns of PEs in a coarse-grained manner, thus leading to SA microarchitecture complying with the varying computational requirements of modern DL workloads. Our analysis demonstrates enhanced power efficiency ranging between 10% and 25% for the best case 32×32 and 64×64 SA designs, respectively. Additionally, the power delay product (PDP) exhibits a progressive improvement of around 6% for larger SA sizes. Moreover, a performance comparison between the MobileNet and ResNet50 models indicates generally better SA performance for the ResNet50 workload. This is due to the more regular convolutions portrayed by ResNet50 that are more favored by SAs, with the performance gap widening as the SA size increases.

1. Introduction

The meteoric rise in deep learning (DL)-based solutions has revolutionized a variety of domains such as image processing, pattern recognition, and transportation. To make full use of the benefits that such DL models can accord, huge operational costs need to be incurred due to the computational complexities accompanying such models (Yüzügüler et al., 2023). Therefore, specialized DL accelerators are necessary to offer enhanced computational prowess while still being energy efficient. Several such accelerators have been proposed on both sides of the DL continuum, i.e., the cloud (Bobda et al., 2022; Li et al., 2023) and the edge (Seshadri et al., 2022; Loh et al., 2025).

Deep neural networks (DNNs) incorporate matrix multiplication as the primary primitive, which in turn, fortunately offers a lot of parallelism, and its acceleration is imperative to achieve tremendous processing demands of the DL workloads (Muslim et al., 2024). Several general matrix multiplication (GEMM) accelerators found in the literature are based on systolic arrays (SAs) (Jouppi et al., 2017; Song et al., 2019; Lai and Zhang, 2024). SA is a two-dimensional mesh of processing elements (PEs) with each PE being fed inputs from the left and top sides. Each PE performs a multiplication and accumulation (MAC) operation every clock cycle, and the partial result and the input are fed to the neighboring PE via pipeline registers. In this way, the PEs collaborate to offer enhanced parallelism in processing data efficiently and accelerate the DNN computation. A typical SA architecture is depicted in Fig. 1.

One of the main factors necessitating the usage of customized architectures for DL acceleration is the resulting improvement in the energy efficiency offered by the accelerators. Such customized hardware must comply with the tight power budget, which is important both at the cloud and the edge. However, the power efficiency of such designs at the edge due to the constrained power availability is of even greater importance (Kim et al., 2020). Thus, a lot of research has been done in recent years to improve the power efficiency of such hardware accelerators. The research found in the literature mainly targets complex microarchitectural optimizations, e.g., by reducing expensive memory access optimizations or by quantization of the networks leading to approximate computing, thus leading to enhanced energy savings (Chen YJ et al., 2016; Moons et al., 2016). Moreover, the modern DNNs are accompanied by increased sparsity, much like their biological counterparts and can generalize well when compared to their denser versions (Hoefler et al., 2021; Guo et al., 2024).

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Fahad Bin Muslim, Kashif Inayat, Muhammad Zain Siddiqi, Safiullah Khan, Tayyeb Mahmood, Ihtesham ul Islam (2025). SAPER-AI accelerator: a systolic array-based power-efficient reconfigurable AI accelerator. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400867
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is the SAPER-AI accelerator?

SAPER-AI is a systolic array-based reconfigurable AI accelerator that improves power efficiency by selectively powering off rows and columns of processing elements, targeting edge AI applications.

How does SAPER-AI achieve power savings?

It employs unified power format (UPF) to express power intent, allowing coarse-grained deactivation of PE rows and columns based on computational requirements, leading to 10-25% power efficiency gains.

What are systolic arrays in AI accelerators?

Systolic arrays are 2D meshes of processing elements that perform multiplication-accumulation operations in a coordinated pipeline, enabling efficient matrix multiplication for deep learning workloads.

Why does ResNet50 show better performance than MobileNet on SAPER-AI?

ResNet50 exhibits more regular convolution patterns, which map more efficiently onto systolic array architectures, and this performance advantage grows with larger array sizes.

What is the significance of this research?

It demonstrates a low-effort design approach for energy-efficient AI accelerators suitable for edge devices, balancing power efficiency and computational performance.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF