SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2400061Original Research

Dynamic prompting class distribution optimization for semi-supervised sound event detection

Lijian Gao¹,Qing Zhu¹,Yaxin Shen¹,Qirong Mao¹,Yongzhao Zhan¹

School of Computer Science and Communication Engineering, Jiangsu University, Zhenjiang 212016, China

Read Executive PreviewQuick FAQ
Dynamic prompting class distribution optimization for semi-supervised sound event detection
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:April 15, 2025Edition:Vol. 32, Issue 4 • pp. 371-383Citation:Lijian Gao et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:semi-supervised learningsound event detectionprompt tuningclass distribution learningpseudo-labelingteacher-student frameworkDCASE challengeaudio tagging

Key Takeaways & Executive Findings

  • • PADO introduces dynamic prompt tuning to optimize class distribution learning in semi-supervised sound event detection, effectively mitigating noisy interference from pseudo-labels and domain gaps. • The method achieves significant performance improvements over state-of-the-art approaches on DCASE 2019, 2020, and 2021 challenge datasets. • PADO maintains model generalization while improving the efficiency of class distribution learning, addressing the trade-off between pseudo-label quality and quantity. • The framework is readily extendable to other benchmark models, demonstrating versatility beyond specific SSED architectures.
Sponsored Research Highlight

Abstract

Semi-supervised sound event detection (SSED) tasks typically leverage a large amount of unlabeled and synthetic data to facilitate model generalization during training, reducing overfitting on a limited set of labeled data. However, the generalization training process often encounters challenges from noisy interference introduced by pseudo-labels or domain knowledge gaps. To alleviate noisy interference in class distribution learning, we propose an efficient semi-supervised class distribution learning method through dynamic prompt tuning, named prompting class distribution optimization (PADO). Specifically, when modeling real labeled data, PADO dynamically incorporates independent learnable prompt tokens to explore prior knowledge about the true distribution. Then, the prior knowledge serves as prompt information, dynamically interacting with the posterior noisy-class distribution information. In this case, PADO achieves class distribution optimization while maintaining model generalization, leading to a significant improvement in the efficiency of class distribution learning. Compared with state-of-the-art methods on the SSED datasets from DCASE 2019, 2020, and 2021 challenges, PADO achieves significant performance improvements. Furthermore, it is readily extendable to other benchmark models.

1. Introduction

Sound event detection (SED) has gained significant attention due to its practical relevance in various real-world applications such as audio surveillance (Crocco et al., 2016; Park and Kim, 2020), acoustic scene understanding (Imoto et al., 2020), and human–machine interaction (Fu et al., 2019). SED tasks require recognizing the categories of events and marking the onset and offset time for each event in a mixed audio recording. This generally involves two separate sub-tasks: audio tagging and audio localization (Mesaros et al., 2021). Specifically, when given an input spectrogram, the SED model will output two-level predictions: clip-level probabilities for audio tagging and frame-level probabilities for event localization.

Traditionally, training a well-performing SED model requires an ample amount of manually labeled training data (Gao LJ et al., 2019, 2022, 2024; Li et al., 2020; Serizel et al., 2020). However, the fine-grained manual annotation at the frame level is extremely time-consuming and results in a severe shortage of annotated samples, presenting a major hurdle for SED research in the big data era. To address this issue, researchers have shifted their focus towards semi-supervised sound event detection (SSED) tasks, leveraging large-scale unlabeled and synthetic data for generalization learning to effectively mitigate overfitting on the limited labeled data. Recently, a teacher–student framework (Yan et al., 2020; Koh et al., 2021; Zheng et al., 2021b; Gao LJ et al., 2023) was established by the most popular methods in SSED commonly based on the consistency regularization assumption (Tarvainen and Valpola, 2017). In this framework, the teacher model provides pseudo-labels to guide the student model for generalization training on the unlabeled and synthetic data, while the student model exploits real labeled data concurrently for supervised training.

Despite the success of the teacher–student framework in generalization learning, noisy interference introduced by pseudo-labels or domain knowledge bias of synthetic data in class distribution learning is often under-considered, resulting in suboptimal performance in SED tasks. Therefore, one of the most challenging tasks in SSED is alleviating the noisy interference in class distribution learning. To this end, recent works have achieved performance gains through effective pseudo-labeling (PL) strategies (Chan and Chin, 2021; Koh et al., 2021), or try to transfer the domain knowledge from synthetic data domain to real domain (Zheng et al., 2021a). However, there is a trade-off between the quality and quantity of pseudo-labels, leading to a reduction in the utilization of unlabeled data. Additionally, effective...

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Lijian Gao, Qing Zhu, Yaxin Shen, Qirong Mao, Yongzhao Zhan (2025). Dynamic prompting class distribution optimization for semi-supervised sound event detection. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400061
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is PADO in sound event detection?

PADO (Prompting class distribution optimization) is a method for semi-supervised sound event detection that uses dynamic prompt tuning to optimize class distribution learning, mitigating noisy interference from pseudo-labels and domain gaps.

How does PADO improve the performance of SSED models?

PADO dynamically incorporates learnable prompt tokens to explore prior knowledge about the true distribution, which then interacts with posterior noisy-class distribution information. This optimizes class distribution learning while maintaining generalization, leading to significant performance improvements on DCASE challenge datasets.

Which datasets were used to evaluate PADO?

PADO was evaluated on semi-supervised sound event detection datasets from the DCASE 2019, 2020, and 2021 challenges.

Can PADO be extended to other models beyond the proposed framework?

Yes, PADO is readily extendable to other benchmark models, making it a versatile approach for improving class distribution learning in various semi-supervised learning settings.

What is the main challenge addressed by PADO in SSED?

PADO addresses the challenge of noisy interference in class distribution learning, which arises from pseudo-labels and domain knowledge gaps in synthetic data, leading to more robust and accurate sound event detection.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF