Key Takeaways & Executive Findings
- •• Random forest achieved the highest predictive accuracy (R² = 0.8541) for Zn recovery from carbonate ores in NaOH leaching. • SHAP analysis identified NaOH concentration, leaching time, and solid-to-liquid ratio as the most positive influencers, while Ca, Fe, and Pb inhibited recovery. • The study compiled 422 experimental observations and compared four regression models, offering a scalable framework for process optimization. • Machine learning combined with explainable AI provides actionable insights for reagent optimization in hydrometallurgical zinc production.
Abstract
This study addresses the challenge of predicting zinc (Zn) recovery from carbonate ores via sodium hydroxide (NaOH) leaching. This complex process influenced by variable ore composition, surface passivation effects, and nonlinear reaction dynamics, which complicate reagent optimization and process control in hydrometallurgical operations. To tackle this, a dataset containing 422 experimental observations was compiled from previous studies, incorporating ore composition and process parameters, such as NaOH concentration, leaching time, temperature, and solid-to-liquid ratio. Four regression models (decision tree, neural network, generalized additive model, and random forest) were trained and evaluated using performance metrics, such as coefficient of determination (R²), root mean squared error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and symmetrical mean absolute percentage error (SMAPE). Among these, the random forest model achieved the best predictive accuracy, with R² value of 0.8541 on the test set and the lowest error rates, demonstrating its effectiveness in capturing the complex relationships between input variables and Zn recovery. Explainable artificial intelligence, particularly SHapley additive exPlanations (SHAP) analysis, revealed that NaOH concentration, leaching time, and solid-to-liquid ratio had the most positive influence on Zn recovery, whereas elements such as Ca, Fe, and Pb had inhibitory effects. These findings align with known geochemical behavior and provide valuable insights for reagent optimization and process efficiency in leaching processes. This study demonstrates the practical potential of machine learning in mineral processing, offering a scalable framework for optimizing Zn recovery from non-sulfide ores and a data-driven approach to enhance decision-making in hydrometallurgical applications.
1. Introduction
In recent years, the metal industry and technological advancements have progressed rapidly, leading to heightened demand for a variety of metals. Mining companies, beneficiation facilities, and metal producers aim to meet this demand with greater efficiency and profitability. Since the primary sources of metals are getting depleted and secondary sources (such as process tailings and discarded end-user products) become valuable alternatives, the major concern in metal production is identifying the optimal production methods and conditions to get more efficiency and profitability. Among the commonly used metals, zinc (Zn) is particularly significant due to its unique physical and chemical properties, especially in galvanization and alloying processes. Zn is primarily extracted from sphalerite (ZnS) minerals using the well-known conventional flotation method. However, due to high demand, sulfide deposits are insufficient, leading to a shift in Zn production from primary sources to non-sulfide ores, with smithsonite (ZnCO3) being the most commonly utilized mineral for Zn extraction.
Unlike Zn extracted from sulfide ores, Zn derived from the mineral smithsonite has been obtained by leaching procedures. Numerous researchers have employed both acidic and alkaline leaching systems, utilizing a range of acids, including sulfuric, hydrochloric, and organic acids, as well as alkaline agents such as sodium hydroxide (NaOH) and ammoniacal solutions. However, studies showed that alkaline leaching produced much better dissolution recoveries because it was more selective and prevented the dissolution of undesirable compounds into the final solution, which would have contaminated the solution. Thus, a lot of research has been done in the literature on the leaching of Zn from various smithsonite ores using NaOH solutions [1–6]. The majority of these studies have concentrated on enhancing recovery value and identifying the influencing parameters. To evaluate the effectiveness of these parameters on the leaching recovery of Zn, researchers have predominantly focused on several factors, including the concentration of NaOH, leaching duration, leaching temperature, solid-to-liquid ratio, and stirring speed. Given the extensive body of published research, it is essential to determine which leaching factors are considered most critical and their effects on Zn recovery.
Consequently, a machine learning approach is warranted to evaluate the leaching parameters in relation to their importance for Zn recovery through leaching. Machine learning (ML), a subset of artificial intelligence, serves as a powerful technology for optimizing sophisticated and labor-intensive operational investigations. By leveraging data from previous studies, ML systems employ statistical and mathematical techniques to establish an appropriate model for experimental inquiries. Through the application of nonlinear functions, these techniques are capable of revealing intricate relationships between input and output data, thereby enabling the extrapolation of conclusions for new scenarios.
Loading authentic research manuscript (Pages 1–5)...
Ilker Erkan, Mehmet Akif Günen (2025). Comparison of Zn recovery prediction from carbonate ores with machine-learning methods. Journal of Mineral Metallurgy and Materials Science. https://doi.org/10.1007/s12613-025-3286-4
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
Which machine learning model performed best for predicting zinc recovery?
The random forest model achieved the best predictive accuracy with an R² value of 0.8541 on the test set, outperforming decision tree, neural network, and generalized additive models.
What factors most influence zinc recovery during NaOH leaching?
SHAP analysis revealed that NaOH concentration, leaching time, and solid-to-liquid ratio have the most positive influence on zinc recovery, while elements such as Ca, Fe, and Pb exhibit inhibitory effects.
Why is alkaline leaching with NaOH preferred for smithsonite ores?
Alkaline leaching using NaOH is more selective and prevents the dissolution of undesirable contaminants into the final solution, resulting in better dissolution recoveries compared to acidic leaching.
How many experimental observations were used in the study?
The dataset comprised 422 experimental observations compiled from previous studies, including ore composition and process parameters.
What is the significance of using SHAP analysis in this study?
SHAP analysis provides explainable AI insights, identifying the direction and magnitude of each input variable's influence on zinc recovery, which helps in optimizing reagent use and process conditions.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena