Key Takeaways & Executive Findings
- •• AnDi architecture integrates analogue and digital computing cores to overcome limitations of analogue CIM in handling complex regression tasks requiring precise floating-point calculations. • The dual-domain floating-point (DDFP) processor enables FP compatibility and decouples NN algorithms from hardware, allowing training without considering specific architectures. • Fine-grained dual-domain mapping optimizes weight allocation between digital and analogue cores, improving energy efficiency and reducing memory management overhead. • NN feature-enhancing technique mitigates noise accumulation in analogue CIM by deploying lightweight enhancing layers on digital cores.
Abstract
Developing efficient neural network (NN) computing systems is crucial in the era of artificial intelligence (AI). Traditional von Neumann architectures have both the issues of "memory wall" and "power wall", limiting the data transfer between memory and processing units. Compute-in-memory (CIM) technologies, particularly analogue CIM with memristor crossbars, are promising because of their high energy efficiency, computational parallelism, and integration density for NN computations. In practical applications, analogue CIM excels in tasks like speech recognition and image classification, revealing its unique advantages. For instance, it efficiently processes vast amounts of audio data in speech recognition, achieving high accuracy with minimal power consumption. In image classification, the high parallelism of analogue CIM significantly speeds up feature extraction and reduces processing time. With the boosting development of AI applications, the demands for computational accuracy and task complexity are rising continually. However, analogue CIM systems are limited in handling complex regression tasks with needs of precise floating-point (FP) calculations. They are primarily suited for the classification tasks with low data precision and a limited dynamic range. A novel analogue-digital unified CIM architecture (named as AnDi) has been developed by integrating the analogue and digital computing cores, which aims to address the above challenges of analogue CIM. The key component is the dual-domain floating-point (DDFP) processor, which serves as a universal data type to represent both FP and integer (INT) numbers. This processor enables FP compatibility regardless of the native capabilities of analogue CIM or digital cores. The DDFP data structure consists of an INT tensor for the feature map and an FP scale. The DDFP processor manages data flow between the analogue and digital domains, performing quantization or dequantization operation. This effectively decouples the NN algorithm from the underlying hardware, enabling more general NN computation. The training process for the top-level algorithm can proceed without considering specific hardware architectures, as all data-scaling operations are handled at the hardware level. To maximize the potential of AnDi architecture, several strategies have been proposed. The fine-grained dual-domain mapping divides weight matrices larger than analogue CIM arrays into smaller ones. Weights are allocated to either digital or analogue cores based on their size. Low-parallel weight blocks are assigned to digital cores, thus maximizing the utilization rates of digital multiply-accumulate (MAC) cores. This approach enhances system-level energy efficiency by avoiding weight splits into additional blocks within the same layer, thereby reducing memory management overhead. The NN feature-enhancing technique addresses the noise accumulation issue in analogue CIM. Lightweight enhancing layers are integrated into the network and deployed on noise-free digital cores for inference.
1. Introduction
Developing efficient neural network (NN) computing systems is crucial in the era of artificial intelligence (AI). Traditional von Neumann architectures have both the issues of "memory wall" and "power wall", limiting the data transfer between memory and processing units. Compute-in-memory (CIM) technologies, particularly analogue CIM with memristor crossbars, are promising because of their high energy efficiency, computational parallelism, and integration density for NN computations.
In practical applications, analogue CIM excels in tasks like speech recognition and image classification, revealing its unique advantages. For instance, it efficiently processes vast amounts of audio data in speech recognition, achieving high accuracy with minimal power consumption. In image classification, the high parallelism of analogue CIM significantly speeds up feature extraction and reduces processing time. With the boosting development of AI applications, the demands for computational accuracy and task complexity are rising continually. However, analogue CIM systems are limited in handling complex regression tasks with needs of precise floating-point (FP) calculations. They are primarily suited for the classification tasks with low data precision and a limited dynamic range.
Loading authentic research manuscript (Pages 1–5)...
Liang Chu, Wenjun Li (2025). A leap forward in compute-in-memory system for neural network inference. SinoTechIntel Verified Research. https://doi.org/10.1088/1674-4926/25020028
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the AnDi architecture?
AnDi is a novel analogue-digital unified compute-in-memory (CIM) architecture that integrates analogue and digital computing cores to address the limitations of analogue CIM in handling complex regression tasks requiring precise floating-point calculations.
How does the dual-domain floating-point (DDFP) processor work?
The DDFP processor serves as a universal data type representing both floating-point (FP) and integer (INT) numbers. It consists of an INT tensor for the feature map and an FP scale, managing data flow between analogue and digital domains through quantization or dequantization operations.
What are the key strategies to maximize AnDi's potential?
Key strategies include fine-grained dual-domain mapping to allocate weights optimally between digital and analogue cores, and NN feature-enhancing technique to mitigate noise accumulation in analogue CIM by deploying lightweight enhancing layers on digital cores.
What are the benefits of AnDi over traditional analogue CIM?
AnDi enables FP compatibility, decouples NN algorithms from hardware, improves energy efficiency through optimized weight mapping, and reduces noise accumulation, thus supporting more general NN computation including complex regression tasks.
What is the significance of decoupling NN algorithms from hardware?
Decoupling allows training of top-level algorithms without considering specific hardware architectures, as all data-scaling operations are handled at the hardware level, simplifying the development process and enhancing portability.
Related Technical Papers & Translations
A Novel Approach for Enhanced Brain Tumor Segmentation Using Multimodal MRI and Deep Learning
Brain tumor segmentation from multimodal MRI is crucial for diagnosis and treatment planning. In this study, we propose a novel deep learning framework that integrates structural and functional imaging modalities to improve segmentation accuracy. Our method employs a multi-scale attention mechanism and a hybrid loss function to handle class imbalance and boundary ambiguity. Evaluated on the BraTS benchmark, our approach achieves state-of-the-art performance, with Dice scores of 0.91, 0.87, and 0.84 for whole tumor, core, and enhancing tumor, respectively. Furthermore, we demonstrate the generalizability of our model across different scanners and protocols. Our findings suggest that the proposed method can significantly aid clinical decision-making and surgical planning.
Investigation of coupled acoustic and electrical responses and early warning approaches during re-loading of damaged coal
Initial damage from engineering disturbances in deep coal mining degrades mechanical properties and heightens dynamic-hazard risks, challenging conventional monitoring. This study probes the coupled acoustic-electrical responses of initially damaged coal under reloading and develops a multi-parameter, multi-level dynamic integrated early-warning model. Using a true-triaxial Split Hopkinson Pressure Bar (SHPB) system, we prepared specimens with graded damage by varying static deviatoric stresses and dynamic impacts. Uniaxial compression reloading was conducted with synchronous acoustic emission (AE) and resistivity monitoring. Joint time-domain responses of force, acoustics, and electricity delineated distinct loading stages. Time-frequency features were extracted via Fourier and wavelet transforms; crack architecture was quantified by 3D AE localization and fractal-dimension analysis. Initial damage markedly reduced load-bearing capacity. Resistivity decreased sharply with increasing deviatoric stress, while cumulative AE counts increased strongly. The AE spectrum evolved from bimodal to broadband with low- and high-frequency enhancement. The resistivity spectrum showed progressive bandwidth broadening, energy amplification, and high-frequency advancement. The AE spatial fractal dimension rose significantly during compaction. An integrated warning system combining multiscale entropy fusion, Temporal Convolutional Network (TCN)-Transformer forecasting, recurrence-network analysis, and a Bayesian framework yielded a 28.4 s lead time, offering a theoretical basis and technical pathway for intelligent prevention of dynamic hazards.
Influence of aggregate particle size on fracture behavior and energy evolution of cemented rockfill in the post-peak stage
Cemented rockfill (CRF) combines structural support with sustainable reuse of coal-derived solid waste. This study integrates digital image correlation, acoustic emission monitoring, and finite–discrete element simulations to investigate mechanical behavior, fracture development, and energy evolution of CRF containing 54% aggregate content with three grain-size distributions (5–10, 10–20, and 20–30 mm). Results indicate finer aggregates raise compressive strength and elastic modulus, and increase post-peak softening and residual stiffness. Fracture patterns transition from dominantly unidirectional failure in coarse specimens to pronounced X-shaped conjugate shear in fine specimens, with cracks initiating at boundaries and propagating inward. The proportion of failed joints at comparable strains decreases markedly with finer gradation, reflecting a more homogeneous crack network that enhances post-peak load retention and produces frequent minor stress fluctuations. Energy analyses reveal a coarse > medium > fine ordering in cumulative dissipation; however, finer aggregates delay rapid kinetic and dissipative energy release, promoting slower energy redistribution and improved load resistance. These findings quantify how aggregate gradation controls deformational mechanisms, crack topology, and energy partitioning, and provide design guidance for optimizing aggregate size and cementitious composition to enhance ductility, energy absorption, and structural reliability of CRF in underground engineering.