Key Takeaways & Executive Findings
- •• Proposes DRMSpell, a multimodal pretrained language model that dynamically reweights phonological and visual modalities to enhance Chinese spelling correction (CSC) performance. • Introduces a dynamically reweighting multimodality (DRM) module that adaptively determines the contribution of each modality per character, improving the model's ability to target different error types. • Develops an independent-modality masking strategy (IMS) during pretraining that strengthens multimodal interaction and robustness against incorrect modal information. • Achieves state-of-the-art results on widely used CSC benchmarks, demonstrating effective modeling of cross-modal interactions and resilience to noisy modal inputs.
Abstract
Chinese spelling correction (CSC) is a task that aims to detect and correct the spelling errors that may occur in Chinese texts. However, the Chinese language exhibits a high degree of complexity, characterized by the presence of multiple phonetic representations known as pinyin, which possess distinct tonal variations that can correspond to various characters. Given the complexity inherent in the Chinese language, the CSC task becomes imperative for ensuring the accuracy and clarity of written communication. Recent research has included external knowledge into the model using phonological and visual modalities. However, these methods do not effectively target the utilization of modality information to address the different types of errors. In this paper, we propose a multimodal pretrained language model called DRMSpell for CSC, which takes into consideration the interaction between the modalities. A dynamically reweighting multimodality (DRM) module is introduced to reweight various modalities for obtaining more multimodal information. To fully use the multimodal information obtained and to further strengthen the model, an independent-modality masking strategy (IMS) is proposed to independently mask three modalities of a token in the pretraining stage. Our method achieves state-of-the-art performance on most metrics constituting widely used benchmarks. The findings of the experiments demonstrate that our method is capable of modeling the interactive information between modalities and is also robust to incorrect modal information.
1. Introduction
Chinese spelling correction (CSC) is a crucial task in the field of natural language processing, as it plays a vital role in ensuring the accuracy and clarity of written communication in Chinese (Cheng et al., 2020; Xu et al., 2021). CSC aims to detect and correct spelling errors that may occur in Chinese texts, which can be particularly challenging due to the nature of Chinese characters represented as pictographs (Liu et al., 2021). For instance, Chinese characters may possess multiple phonetic representations known as pinyin, where each pinyin comprises distinct tonal variations. These tonal variations can also correspond to different characters (Sun et al., 2021). Given these complexities, texts in the Chinese language frequently encounter spelling errors that are both phonetically and visually similar, as illustrated in Fig. 1. Additionally, the meaning of a sentence is modified dramatically when some characters are incorrect. This is a scenario wherein the given context is also impacted to a great extent (Huang et al., 2021). Given these problems arising in Chinese texts, CSC is challenging and important for some downstream tasks such as optical character recognition (OCR) and automatic speech recognition (ASR) (Bhardwaj et al., 2022; Kim et al., 2022).
Some works developed some confusion sets containing phonologically or visually similar character pairs (Wang et al., 2018; Cheng et al., 2020; Ma et al., 2023), aiming to model the relationship between these characters. However, these confusion sets possess characters that are limited in scope and used by heuristic rules, which results in poor use of phonological and visual information. Recent works have introduced external knowledge into the pretrained language model (Guo et al., 2021; Zhang RQ et al., 2021; Liang et al., 2023) to overcome the shortcomings of the confusion set. However, due to Chinese spelling errors often involving homophones or similarly-looking characters, the appropriateness of the modal information used would vary across the different error types. Therefore, it is essential to differentiate the contributions made by different modal features to error correction. The existing approaches have not adequately modeled the semantic connections between the different modalities. They engage in merely introducing multimodal information through simple fusion operations, such as summation, thereby reducing the contribution of the different modalities.
In this paper, we propose DRMSpell, a multimodal pretrained language model for CSC. A dynamically reweighting multimodality (DRM) component is proposed for the reweighting of different modal inputs for each character in the sequence. DRM can dynamically determine which modality to rely on more according to th
Loading authentic research manuscript (Pages 1–5)...
Yinghao LI, Heyan HUANG, Baojun WANG, Yang GAO (2025). DRMSpell: dynamically reweighting multimodality for Chinese spelling correction. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2300816
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is Chinese spelling correction (CSC)?
Chinese spelling correction (CSC) is a natural language processing task that aims to detect and correct spelling errors in Chinese texts. These errors often involve phonologically or visually similar characters, making the task challenging due to the complexity of the Chinese language, including its multiple phonetic representations (pinyin) and tonal variations.
What is DRMSpell?
DRMSpell is a multimodal pretrained language model proposed for Chinese spelling correction. It dynamically reweights phonological and visual modalities to better utilize external knowledge and model interactions between modalities, achieving state-of-the-art performance on widely used benchmarks.
How does the dynamically reweighting multimodality (DRM) module work?
The DRM module dynamically reweights different modal inputs (e.g., phonological and visual features) for each character in the sequence. It determines which modality to rely on more based on the specific error type, thereby improving the model's ability to address different kinds of spelling errors.
What is the independent-modality masking strategy (IMS)?
IMS is a pretraining strategy that independently masks the three modalities (character, pinyin, and shape) of a token. This encourages the model to learn robust multimodal representations and enhances its ability to handle incorrect or missing modal information during inference.
What are the key contributions of the DRMSpell paper?
The key contributions are threefold: (1) proposing a novel multimodal pretrained language model DRMSpell for CSC, (2) introducing a dynamically reweighting multimodality module to adaptively fuse modalities, and (3) developing an independent-modality masking strategy to strengthen multimodal interaction and robustness, leading to state-of-the-art results on benchmark datasets.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena