Key Takeaways & Executive Findings
- •• DDiNER integrates a hierarchical industrial domain dictionary with BERT, BiLSTM, and CRF for multilevel feature fusion, effectively addressing ambiguous entity boundaries and semantic overlaps in complex industrial Chinese text. • The framework achieves superior performance with average precision, recall, and F1-scores of 95.75%, 95.73%, and 95.74%, respectively, outperforming state-of-the-art models. • Validation on an independent dataset demonstrates strong robustness and capability in recognizing unseen and long-tail entities. • DDiNER provides an effective and scalable solution for industrial Chinese NER, with significant potential for information extraction, knowledge graph construction, and intelligent decision-making.
Abstract
Accurate Chinese named entity recognition (NER) in the process industry is crucial for applications such as information extraction, knowledge graph construction, and intelligent decision-making. However, challenges, including ambiguous entity boundaries, semantic overlaps, and limited annotated data, significantly hinder performance. To address these issues, this study proposes DDiNER, a domain dictionary-guided Chinese NER framework that integrates a hierarchical industrial domain dictionary with bidirectional encoder representations from Transformers (BERT) via a hierarchical lexicon adapter (HLA), combined with bidirectional long short-term memory (BiLSTM) and conditional random field (CRF) layers for multilevel feature fusion. Experimental results show that DDiNER achieves superior performance, with average precision, recall, and F1-scores of 95.75%, 95.73%, and 95.74%, respectively, outperforming state-of-the-art models. Validation on an independent dataset confirms its robustness and strong capability in recognizing unseen and long-tail entities. This study provides an effective and scalable solution for industrial Chinese NER, with significant potential for downstream intelligent applications.
1. Introduction
Named entity recognition (NER) aims to automatically extract specific entities from massive unstructured text and identify their corresponding categories (Gao et al., 2021; Liu P et al., 2022; Ehrmann et al., 2024), serving as a fundamental task for knowledge extraction and an essential foundation for downstream applications such as knowledge graph construction (Liu C and Yang, 2022; Zhong et al., 2024), information retrieval (Kumar and Starly, 2022), and question-answering systems (Hu Z and Ma, 2023; Prasanna et al., 2024). In the process industry domain, NER is a cornerstone of natural language processing (NLP), enabling the transformation of massive amounts of unstructured industrial text into structured machine-interpretable knowledge, which supports applications such as production optimization, safety monitoring, and intelligent decision-making.
However, most of the current research is focused more on domains such as general purpose (Geng et al., 2023; Yang et al., 2024), finance (Zhang et al., 2023, 2024), agriculture (G et al., 2023; De et al., 2025), and medicine (Hu Z and Ma, 2023; Hu Y et al., 2024), compared with the process industry domain, especially the Chinese process industry domain. Compared with general-domain Chinese NER, process-industry Chinese NER faces significantly greater challenges due to the complexity of domain knowledge, diversity of terminology, and heterogeneity of data sources. Texts in process industrial contexts, such as production records, equipment maintenance logs, and operational guidelines, often contain highly specialized terminologies, ambiguous abbreviations, compound expressions, and mixed data formats, making accurate entity recognition considerably difficult.
There are three obvious challenges. First, process industrial texts often contain multi-word compound entities that describe specific equipment or phenomena, such as “continuous casting crystallizer,” which combines “continuous casting” and “crystallizer.” Accurate identification of such entities requires precise handling of compound terms and correct Chinese word segmentation; otherwise, segmentation errors can propagate and lead to recognition inaccuracies. Second, many abbreviations and terminologies in process industry texts exhibit polysemy, meaning that their interpretations vary across operational contexts. For example, “CNC” may refer to “computer numerical control” in mechanical manufacturing but represent another meaning in other automation contexts, necessitating NER systems to incorporate domain knowledge and context-aware disambiguation strategies. Third, unlike general-domain NER, process industry NER suffers from a shortage of large-scale and expert-annotated corpora due to the complexity and confidentiality of industrial data. This data scarcity makes it difficult to train deep learning models effectively and limits the performance of generic pre-trained language models on domain-specific tasks.
Loading authentic research manuscript (Pages 1–5)...
Ronghui LIU, Wei CUI, Xiaojun LIANG, Weihua GUI (2025). DDiNER: domain dictionary-guided Chinese named entity recognition for complex industrial contexts. Engineering Information Technology & Electronic Engineering. https://doi.org/10.1631/ENG_ITEE_2025_0047
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is DDiNER?
DDiNER is a domain dictionary-guided Chinese named entity recognition framework designed for complex industrial contexts. It integrates a hierarchical industrial domain dictionary with BERT via a hierarchical lexicon adapter, combined with BiLSTM and CRF layers for multilevel feature fusion.
What challenges does process industry Chinese NER face?
Process industry Chinese NER faces challenges such as ambiguous entity boundaries, semantic overlaps, polysemous abbreviations, compound expressions, and limited expert-annotated corpora, making accurate recognition difficult.
How does the domain dictionary improve NER performance?
The hierarchical domain dictionary is integrated via a hierarchical lexicon adapter (HLA) with BERT, injecting domain-specific lexical knowledge and enabling better handling of compound entities and context-aware disambiguation.
What were the experimental results of DDiNER?
DDiNER achieved average precision, recall, and F1-scores of 95.75%, 95.73%, and 95.74%, respectively, outperforming state-of-the-art models and showing strong robustness on unseen and long-tail entities.
What are potential downstream applications of DDiNER?
The framework can support information extraction, knowledge graph construction, intelligent decision-making, production optimization, and safety monitoring in the process industry.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena