• Proposes DRMSpell, a multimodal pretrained language model that dynamically reweights phonological and visual modalities to enhance Chinese spelling correction (CSC) performance.
• Introduces a dynamically reweighting multimodality (DRM) module that adaptively determines the contribution of each modality per character, improving the model's ability to target different error types.
• Develops an independent-modality masking strategy (IMS) during pretraining that strengthens multimodal interaction and robustness against incorrect modal information.
• Achieves state-of-the-art results on widely used CSC benchmarks, demonstrating effective modeling of cross-modal interactions and resilience to noisy modal inputs.