SinoTechIntel Academic Portal
Official PDF TranslationFrontiers of Information Technology & Electronic Engineering

Multi-talker audio–visual speech recognition towards diverse scenarios

Authors: Yuxiao LIN; Tao JIN; Xize CHENG; Zhou ZHAO; Fei WU

DOI: 10.1631/FITEE_2500411Status: Verified Translated Edition
Sponsored AdvertisementAd Placement Area
reCAPTCHA Bot Shield Active

Preparing Secure Academic Download

Verifying human reader & generating high-resolution document...

Verifying Document Integrity15s remaining
← Back to Article
Protected by Google reCAPTCHA v3.PrivacyTerms
Sponsored ContentAdSense In-Feed Ad Slot

Key Findings in This Report

• A novel end-to-end AVSR framework is proposed for realistic multi-talker scenarios, addressing both unknown speaker counts and modality misalignment. • The speaker-number-aware mixture-of-experts (SA-MoE) mechanism adaptively fuses audio and visual information based on the number of overlapping speakers, using speaker counting as an auxiliary task. • A cross-modal realignment (CMR) module robustly handles asynchronous audio-video inputs, overcoming temporal misalignment in real-world recordings. • The challenge-based curriculum learning (CBCL) strategy prioritizes difficult samples, improving training efficiency and overall performance on complex multi-talker speech.
Download Full PDF: Multi-talker audio–visual speech recognition towards diverse scenarios | SinoTechIntel | SinoTechIntel