Key Takeaways & Executive Findings
- •• Interaction theory offers an axiomatic framework that translates DNN decision logic into symbolic interaction concepts, moving beyond empirical explanation methods. • It provides a unified mathematical account of diverse deep learning phenomena, including generalization, adversarial sensitivity, representation bottleneck, and learning dynamics. • The theory reveals that learning dynamics follow a two-phase evolution of interaction complexity, explaining why generalization and adversarial sensitivity change during training. • By unifying empirical attribution and adversarial-transferability-boosting methods, interaction theory advances toward a first-principles explanation for explainable AI.
Abstract
Most explanation methods are designed in an empirical manner, so exploring whether there exists a first-principles explanation of a deep neural network (DNN) becomes the next core scientific problem in explainable artificial intelligence (XAI). Although it is still an open problem, in this paper, we discuss whether the interaction-based explanation can serve as the first-principles explanation of a DNN. The strong explanatory power of interaction theory comes from the following aspects: (1) it establishes a new axiomatic system to quantify the decision-making logic of a DNN into a set of symbolic interaction concepts; (2) it simultaneously explains various deep learning phenomena, such as generalization power, adversarial sensitivity, representation bottleneck, and learning dynamics; (3) it provides mathematical tools that uniformly explain the mechanisms of various empirical attribution methods and empirical adversarial-transferability-boosting methods; (4) it explains the extremely complex learning dynamics of a DNN by analyzing the two-phase dynamics of interaction complexity, which further reveals the internal mechanism of why and how the generalization power/adversarial sensitivity of a DNN changes during the learning process.
1. Introduction
Although deep neural networks (DNNs) have exhibited superior performance in various tasks, their decision-making logic is not transparent. The lack of interpretability is especially critical in high-risk tasks, such as medical diagnosis and autonomous driving. Many studies on explainable artificial intelligence (AI) aim to explain the complex learning behaviors of a DNN. However, the explainable AI community has not reached a consensus on the “first-principles explanation” that can faithfully explain both the knowledge and the generalization power of neural networks. How to define the first-principles explanation of a DNN remains an open problem.
As a result, many explanation methods are designed mainly in an empirical manner. To this end, the first-principles explanation is supposed to be a unified theory system with a set of axioms and theorems that can comprehensively explain the internal mathematical mechanisms of various deep learning phenomena, including knowledge representation, generalization power, adversarial robustness, and learning dynamics.
In this paper, we aim to discuss the potential of the theory system of interactions serving as the first-principles explanation of a DNN, including the theoretical achievements of interaction theory and its limitations. Specifically, we will analyze the representation power of interaction theory from various perspectives, e.g., explaining the knowledge representation of a DNN, explaining the mechanism behind the performance, using interactions to unify empirical deep learning methods, and, more importantly, explaining the extremely complex learning dynamics of a DNN. All these perspectives are necessary issues with respect to the first-principles explanation of a DNN.
Loading authentic research manuscript (Pages 1–5)...
Huilin Zhou, Qihan Ren, Junpeng Zhang, Quanshi Zhang (2025). Towards the first principles of explaining DNNs: interactions explain the learning dynamics. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2401025
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the first-principles explanation of a DNN?
It is a unified theory system with a set of axioms and theorems that can comprehensively explain the internal mathematical mechanisms of deep learning phenomena, including knowledge representation, generalization power, adversarial robustness, and learning dynamics.
Why are interaction-based explanations considered powerful for understanding DNNs?
Interaction theory establishes an axiomatic system that quantifies decision logic into symbolic concepts, explains multiple learning phenomena, unifies empirical attribution methods, and reveals two-phase dynamics of interaction complexity underlying learning.
How does interaction theory explain adversarial sensitivity?
It provides mathematical tools that uniformly explain empirical attribution and adversarial-transferability-boosting methods, and shows how changes in interaction complexity during learning are linked to adversarial sensitivity.
What are the two-phase dynamics of interaction complexity?
The learning dynamics of a DNN exhibit two phases of interaction complexity, which reveal internal mechanisms of why and how generalization power and adversarial sensitivity change during the learning process.
Why do some scholars criticize post-hoc explanation methods?
Researchers such as Rudin, Ghassemi, and Adebayo argue that post-hoc explanations can be unfaithful, misleading, or independent of the model, making them unreliable for high-stakes applications like medical diagnosis.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena