SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2400547Original Research

An end-to-end automatic methodology to accelerate the accuracy evaluation of deep neural networks under hardware transient faults

Jiajia JIAO¹,Ran WEN¹,Hong YANG¹

College of Information Engineering, Shanghai Maritime University, Shanghai 201306, China

Read Executive PreviewQuick FAQ
An end-to-end automatic methodology to accelerate the accuracy evaluation of deep neural networks under hardware transient faults
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:May 15, 2025Edition:Vol. 32, Issue 5 • pp. 796-808Citation:Jiajia JIAO et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:deep neural networkshardware transient faultsfault injectionsilent data corruptionsafety-critical misclassificationA-Meanfast evaluationautomatic evaluation tool

Key Takeaways & Executive Findings

  • • Introduces A-Mean, a unified and end-to-end automatic methodology for rapid evaluation of hardware transient faults on DNNs, achieving up to 922.80× speedup over TensorFI+. • Estimates both general classification accuracy and application-specific safety-critical misclassification (SCM) using silent data corruption rates of basic operations and a two-level mean calculation mechanism. • Employs a max-policy to handle non-sequential structures and a worst-case scheme to compute enlarged SCM and halved accuracy, enabling conservative reliability assessment. • Provides an easy-to-use automatic tool, publicly available, for prompt fault evaluation across diverse DNN models and datasets with minimal accuracy loss (0.77% on average).
Sponsored Research Highlight

Abstract

Hardware transient faults are proven to have a significant impact on deep neural networks (DNNs), whose safety-critical misclassification (SCM) in autonomous vehicles, healthcare, and space applications is increased up to four times. However, the inaccuracy evaluation using accurate fault injection is time-consuming and requires several hours and even a couple of days on a complete simulation platform. To accelerate the evaluation of hardware transient faults on DNNs, we design a unified and end-to-end automatic methodology, A-Mean, using the silent data corruption (SDC) rate of basic operations (such as convolution, addition, multiply, ReLU, and max-pooling) and a static two-level mean calculation mechanism to rapidly compute the overall SDC rate, for estimating the general classification metric accuracy and application-specific metric SCM. More importantly, a max-policy is used to determine the SDC boundary of non-sequential structures in DNNs. Then, the worst-case scheme is used to further calculate the enlarged SCM and halved accuracy under transient faults, via merging the static results of SDC with the original data from one-time dynamic fault-free execution. Furthermore, all of the steps mentioned above have been implemented automatically, so that this easy-to-use automatic tool can be employed for prompt evaluation of transient faults on diverse DNNs. Meanwhile, a novel metric "fault sensitivity" is defined to characterize the variation of transient fault-induced higher SCM and lower accuracy. The comparative results with a state-of-the-art fault injection method TensorFI+ on five DNN models and four datasets show that our proposed estimation method A-Mean achieves up to 922.80 times speedup, with just 4.20% SCM loss and 0.77% accuracy loss on average. The artifact of A-Mean is publicly available at https://github.com/breatrice321/A-Mean.

1. Introduction

Transient faults primarily result from inherent circuit malfunctions, such as external radiation and internal electrical interference. These transient faults often occur randomly and cause performance loss or area/power overhead for detection and mitigation, so that some ignored faults have the potential to corrupt the program or lead to incorrect data. However, in safety-sensitive applications, such as autonomous vehicles, healthcare, and space applications, transient faults could lead to immeasurable loss of life and property. In recent years, deep neural networks (DNNs) have been adopted in safety-critical applications due to their remarkable problem-solving capabilities. However, these safety-critical applications necessitate dependable, robust, and efficient support from popular DNNs. Consequently, it is important to assess the accuracy of various DNN models under transient faults to guarantee their acceptable prediction quality.

Two categories of mainstream approaches are usually employed to evaluate the impact of faults on DNNs: (1) Fault injection involves injecting faults into DNNs intentionally many times so that the impacts can be quantified by the ratio of the number of injections with observed wrong results to the total number of injections, such as CAFI, saca-FI, TensorFI, and TensorFI+. (2) Fault-free analysis uses a few fault-free simulations and simple analytical models for fast evaluation of fault impacts on DNNs, such as SERN, APPRAISER, DeepVigor, and saca-AVF. The former is accurate but time-consuming, while the latter is fast but inaccurate. The DNN-driven safety-critical applications require not only high reliability for guaranteed safety but also high performance for real-time decisions. Therefore, fast fault-free analysis is preferred for safety-critical applications. However, how to improve the evaluation accuracy of the impacts of transient faults on DNNs with the advantage of high speed is a challenge.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Jiajia JIAO, Ran WEN, Hong YANG (2025). An end-to-end automatic methodology to accelerate the accuracy evaluation of deep neural networks under hardware transient faults. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400547
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is A-Mean and how does it improve fault evaluation?

A-Mean is an end-to-end automatic methodology that accelerates the accuracy evaluation of deep neural networks under hardware transient faults. It uses the silent data corruption (SDC) rate of basic operations and a static two-level mean calculation to rapidly estimate classification accuracy and safety-critical misclassification (SCM), achieving up to 922.80 times speedup over traditional fault injection methods like TensorFI+.

How does A-Mean handle non-sequential structures in DNNs?

A-Mean employs a max-policy to determine the SDC boundary of non-sequential structures, such as residual connections or parallel branches. This ensures worst-case reliability estimation for complex DNN architectures.

What metrics does A-Mean evaluate?

A-Mean evaluates two primary metrics: general classification accuracy and application-specific safety-critical misclassification (SCM). It also introduces a novel metric called 'fault sensitivity' to characterize the variation of transient fault-induced higher SCM and lower accuracy.

What are the known limitations of existing fault evaluation methods?

Traditional fault injection methods are accurate but time-consuming, often requiring hours to days for complete simulations. Fault-free analytical methods are fast but less accurate. A-Mean bridges this gap by providing high speed while maintaining low accuracy loss (0.77% on average compared to TensorFI+).

How can researchers access the A-Mean tool?

The artifact of A-Mean is publicly available at https://github.com/breatrice321/A-Mean, allowing researchers and practitioners to evaluate transient faults on diverse DNN models and datasets.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF