Key Takeaways & Executive Findings
- •• The composite aerogel pressure sensors exhibited low hysteresis (13.69%), wide detection range (6.25 Pa-1200 kPa), and cyclic stability to acquire stable and accurate pronunciation signals. • Over 6888 and 4158 pronunciation signals were collected by the pressure sensor and utilized for training the convolutional neural network model, allowing for accurate recognition of six dialects (96.2% accuracy) and seven words (96.6% accuracy). • The work emphasizes innovation in both material design and methodology, bridging sensing performance and mechanical properties. • This research advances silent speech recognition for human–machine interaction and physiological signal monitoring, particularly for non-standard language users.
Abstract
Wearable pressure sensors capable of adhering comfortably to the skin hold great promise in sound detection. However, current intelligent speech assistants based on pressure sensors can only recognize standard languages, which hampers effective communication for non-standard language people. Here, we prepare an ultralight Ti3C2Tx MXene/chitosan/polyvinylidene difluoride composite aerogel with a detection range of 6.25 Pa-1200 kPa, rapid response/recovery time, and low hysteresis (13.69%). The wearable aerogel pressure sensor can detect speech information through the throat muscle vibrations without any interference, allowing for accurate recognition of six dialects (96.2% accuracy) and seven different words (96.6% accuracy) with the assistance of convolutional neural networks. This work represents a significant step forward in silent speech recognition for human–machine interaction and physiological signal monitoring.
1. Introduction
Spoken recognition as a branch of speech recognition can assist people with language barriers as well as human–computer interactions to express ideas and give instructions. The present spoken recognition involves the detection of sound waves directly, including spectral analysis, extraction and comparison of acoustic features, and acoustic texture analysis [1–3]. However, the direct detection approach is susceptible to interference by the transmission media, ambient noise, and the physiological state of the speakers. Speech recognition through mechanical sensors can avoid these defects by detecting the vibration of throat muscles based on the anatomical foundation of the throat during vocalization [4–7].
Wearable pressure sensors that can convert throat vibrations into visualized electrical signals have received widespread attention in detecting speech information [8–11]. Initially, speech recognition was mainly implemented by comparing the waveforms of electrical signals of throat vibrations captured using pressure sensors or tactile sensors [12, 13]. In addition, the pressure sensor can detect vibrations within the throat muscles and distinguish different pronunciations by simple signal processing, such as calculating the slope of signal peaks and comparing peak widths [14]. With the advancement of artificial intelligence (AI) technology, machine learning was introduced to build models for training and recognition of different pronunciations, particularly the combination of pressure sensors and machine learning [15–20]. Convolutional neural network (CNN) and support vector machine have frequently been introduced to identify the collected pronunciation signal for speech recognition [8, 21]. However, pressure sensors for speech recognition are currently restricted to identifying standard languages, which hampers effective communication for dialect speakers [5, 22]. For tone languages, the differences between dialect pronunciations are tone and pitch, which are generated by the throat muscles controlling the movement of the hyoid bone and cartilage. The elevation or depression of voice pitch is closely related to the contraction and relaxation of the throat muscles. The primary challenge in dialect recognition through pressure sensors with narrow-detection range and hysteresis lies in the difficulty in capturing the subtle and rapid vibrations of throat muscles during the vocalization process [23]. These factors place stringent demands on the pressure-sensing performance, such as low detection limit, high stability, and hysteresis characteristics.
To fulfill the requirements for speech recognition, Ti3C2Tx MXene has emerged as a promising candidate for wearable pressure sensors due to its adjustable layer spacing and superior conductivity [24–26]. However, pure Ti3C2Tx typically suffers from mechanical brittleness and oxidization, rendering it susceptible to collapse during repeated cycles [27]. To prevent sensitivity degradation under mechanical stimuli, compositing Ti3C2Tx layers with a nanostructured polymer matrix offer
Loading authentic research manuscript (Pages 1–5)...
Yanan Xiao, He Li, Tianyi Gu, Xiaoteng Jia, Shixiang Sun, Yong Liu, Bin Wang, He Tian, Peng Sun, Fangmeng Liu, Geyu Lu (2024). Ti3C2Tx Composite Aerogels Enable Pressure Sensors for Dialect Speech Recognition Assisted by Deep Learning. Nano-Micro Letters. https://doi.org/10.1007/s40820-024-01605-z
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the main innovation of this research?
The main innovation lies in the design of a Ti3C2Tx MXene/chitosan/polyvinylidene difluoride composite aerogel pressure sensor that combines low hysteresis, wide detection range, and high stability, enabling accurate dialect speech recognition through deep learning.
How does the pressure sensor detect speech?
The wearable aerogel pressure sensor detects throat muscle vibrations during vocalization, converting them into electrical signals without interference from ambient noise or transmission media.
What are the key performance metrics of the sensor?
The sensor exhibits a detection range of 6.25 Pa to 1200 kPa, rapid response/recovery time, and low hysteresis of 13.69%, ensuring stable and accurate pronunciation signal acquisition.
What accuracy was achieved for dialect and word recognition?
With the assistance of convolutional neural networks, the system achieved 96.2% accuracy for six dialects and 96.6% accuracy for seven different words.
What is the significance of this work?
This work advances silent speech recognition for human–machine interaction and physiological signal monitoring, particularly benefiting non-standard language speakers by enabling effective communication.
Related Technical Papers & Translations
Direct Repair of the Crystal Structure and Coating Surface of Spent LiFePO4 Materials Enables Superfast Li-Ion Migration
The rapid accumulation of spent LiFePO4 (LFP) cathodes from retired lithium-ion batteries necessitates the development of effective and environmental-friendly recycling strategies. In this context, direct regeneration has emerged as a promising approach for reclaiming LFP cathode materials, offering a streamlined pathway to restore their electrochemical functionality. We report an integrated regeneration protocol that simultaneously repairs the degraded crystal structure and reconstructs the damaged carbon coating in spent LFP. The regenerated cathode material had superfast lithium-ion diffusion kinetics and a stable cathode–electrolyte interface, giving a remarkable rate capability with specific capacities of 122 mAh g−1 at 5C and 106 mAh g−1 at 10C (1C = 170 mA g−1). It also maintained capacities of 110.7 mAh g−1 (5C) and 84.1 mAh g−1 (10C) after 400 cycles. It could be used in harsh environments and could be stably cycled at subzero temperatures (−10 and −20 °C) and in solid-state electrolyte batteries. Life cycle assessment combined with economic evaluation using the EverBatt model reveals that this direct regeneration approach has high economic and environmental benefits.
Oxide Semiconductor for Advanced Memory Architectures: Atomic Layer Deposition, Key Requirement and Challenges
Oxide semiconductors (OSs), introduced by the Hosono group in the early 2000s, have evolved from display backplane materials to promising candidates for advanced memory and logic devices. The exceptionally low leakage current of OSs and compatibility with three-dimensional (3D) architectures have recently sparked renewed interest in their use in semiconductor applications. This review begins by exploring the unique material properties of OSs, which fundamentally originate from their distinct electronic band structure. Subsequently, we focus on atomic layer deposition (ALD), a core technique for growing excellent OS films, covering both basic and advanced processes compatible with 3D scaling. The basic surface reaction mechanisms—adsorption and reaction—and their roles in film growth are introduced. Furthermore, material design strategies, such as cation selection, crystallinity control, anion doping, and heterostructure engineering, are discussed. We also highlight challenges in memory applications, including contact resistance, hydrogen instability, and lack of p-type materials, and discuss the feasibility of ALD-grown OSs as potential solutions. Lastly, we provide an outlook on the role of ALD-grown OSs in memory technologies. This review bridges material fundamentals and device-level requirements, offering a comprehensive perspective on the potential of ALD-driven OSs for next-generation semiconductor memory devices.
Laser powder bed fusion of biodegradable Zn-4Cu alloy: Processing, microstructure and properties
Zn's natural degradability and biocompatibility make it a promising candidate for implants, however, its mechanical properties remain insufficient for bone applications. In this study, the performance of Zn was enhanced by developing Zn-Cu alloys via laser powder bed fusion (LPBF). Optimal LPBF parameters for forming stable tracks were achieved by adjusting laser power and scanning speed. Under optimized conditions of 100 W and 100 mm/s, high-density (99.58%) Zn-Cu alloys with improved hardness (68.2HV) and yield strength (160 MPa) were achieved. These improvements are attributed to solid solution strengthening, segregation strengthening, and grain refinement. The Zn-Cu alloys also demonstrated favorable degradation behavior, with a rate of 0.16 mm/year. This degradation is primarily driven by micro-galvanic corrosion between the CuZn5 phase and Zn matrix, along with refined grains and increased grain boundary density. This work demonstrates a viable strategy for fabricating Zn-based implants with enhanced structural integrity and mechanical performance via LPBF.