• The proposed attention-based multi-domain fusion approach effectively mitigates information redundancy inherent in simple concatenation, significantly enhancing active sonar target recognition performance.
• Combined 1DCNN-LSTM and 2DCNN with channel attention extract complementary deep features from time-domain and spectral-domain representations.
• Multi-domain cross-attention fusion strengthens inter-domain information interaction, improving feature representation and generalization under low signal-to-clutter ratios.
• Experiments demonstrate superiority over single-domain and existing fusion methods, with robust performance in challenging underwater acoustic environments.