SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2500287Original Research

Jiu fusion artificial intelligence (JFA): a two-stage reinforcement learning model with hierarchical neural networks and human knowledge for Tibetan Jiu chess

Xiali LI¹,Xiaoyu FAN¹,Junzhi YU¹,Zhicheng DONG¹,Xianmu CAIRANG¹,Ping LAN¹

Minzu University of China, Beijing 100081, China

Read Executive PreviewQuick FAQ
Jiu fusion artificial intelligence (JFA): a two-stage reinforcement learning model with hierarchical neural networks and human knowledge for Tibetan Jiu chess
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:February 25, 2025Edition:Vol. 32, Issue 2 • pp. 761-773Citation:Xiali LI et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:Reinforcement learningDeep reinforcement learning

Key Takeaways & Executive Findings

  • • JFA is a two-stage DRL model for Tibetan Jiu chess with separate strategic layout and hierarchical battle sub-models, enabling phase-specific learning. • Knowledge-guided pruning and auxiliary agents reduce layout decision time to approximately 1/147 of AlphaZero while achieving a 74% win rate. • Combined SLM and HBM achieve an 81% win rate against other models, comparable to a human amateur 4-dan player, with HBM alone at 70%. • JFA won first place at the 2024 China National Computer Game Tournament, demonstrating robust performance under limited hardware resources.
Sponsored Research Highlight

Abstract

Tibetan Jiu chess, recognized as a national intangible cultural heritage, is a complex game comprising two distinct phases: the layout phase and the battle phase. Improving the performance of deep reinforcement learning (DRL) models for Tibetan Jiu chess is challenging, especially given the constraints of hardware resources. To address this, we propose a two-stage model called JFA, which incorporates hierarchical neural networks and knowledge-guided techniques. The model includes sub-models: strategic layout model (SLM) for the layout phase and hierarchical battle model (HBM) for the battle phase. Both sub-models use similar network structures and employ parallel Monte Carlo tree search (MCTS) methods for independent self-play training. HBM is structured as a hierarchical neural network, with the upper network selecting movement and jump capturing actions and the lower network handling square capturing actions. Human knowledge-based auxiliary agents are introduced to assist SLM and HBM, simulating the entire game and providing reward signals based on square capturing or victory outcomes. Additionally, within the HBM, we propose two human knowledge-based pruning methods that prune parallel MCTS and capture actions in the lower network. In the experiments against a layout model using the AlphaZero method, SLM achieves a 74% win rate, with the decision-making time being reduced to approximately 1/147 of the time required by the AlphaZero model. SLM also won the first place at the 2024 China National Computer Game Tournament. HBM achieves a 70% win rate when playing against other Tibetan Jiu chess models. When used together, SLM and HBM in JFA achieve an 81% win rate, comparable to the level of a human amateur 4-dan player. These results demonstrate that JFA effectively enhances artificial intelligence (AI) performance in Tibetan Jiu chess.

1. Introduction

Tibetan Jiu chess is a complete information game, encompassing both layout and battle phases (see Section 1 in the supplementary materials). It is recognized as a national intangible cultural heritage and an exemplary traditional cultural symbol of the Chinese nation. The research on its game algorithms not only provides an ideal experimental platform for advancing artificial intelligence (AI) algorithms, but also fosters innovative approaches for the preservation and transmission of traditional culture.

Deep reinforcement learning (DRL) algorithms have been widely applied to board games, with the most achieving performance that surpasses top human players, such as AlphaGo (Silver et al., 2016), AlphaZero (Silver et al., 2018), Libratus (Brown and Sandholm, 2018), and Suphx (Li JJ et al., 2020). The outstanding performance of these agents relies on substantial hardware resources for model training. DRL algorithms for Tibetan Jiu chess include temporal-difference reinforcement learning algorithms, hybrid DRL models, and phased game strategies. However, the aforementioned algorithms or models have not achieved breakthrough progress in enhancing the playing strength of the agents, and the research on Tibetan Jiu chess game strategies faces the following three challenges:

  1. For most complete information games, DRL algorithms combined with Monte Carlo tree search (MCTS) are commonly employed, where the evaluation results of the game tree guide the training of the neural network. In Tibetan Jiu chess, the game consists of layout and battle phases, with simulations extending from leaf nodes to the end of the game. This leads to excessively long search paths in the game tree. Furthermore, using a single model throughout the entire gameplay process makes it challenging for the agent to learn accurately, thereby hindering improvements in playing strength.
  2. Tibetan Jiu chess lacks large-scale, high-quality training datasets. Employing self-play methods to generate training data is time-consuming under the limited hardware resources typically available in laboratory environments.
SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Xiali LI, Xiaoyu FAN, Junzhi YU, Zhicheng DONG, Xianmu CAIRANG, Ping LAN (2025). Jiu fusion artificial intelligence (JFA): a two-stage reinforcement learning model with hierarchical neural networks and human knowledge for Tibetan Jiu chess. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2500287
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is Tibetan Jiu chess?

Tibetan Jiu chess is a complete information game recognized as a national intangible cultural heritage. It consists of two distinct phases: the layout phase and the battle phase, making it a complex board game for AI research.

What is JFA in the context of this paper?

JFA (Jiu fusion artificial intelligence) is a two-stage reinforcement learning model designed for Tibetan Jiu chess. It integrates hierarchical neural networks and human knowledge, comprising a strategic layout model (SLM) for the layout phase and a hierarchical battle model (HBM) for the battle phase.

What were the key results of JFA?

SLM achieved a 74% win rate against an AlphaZero-based layout model, reducing decision time to about 1/147. HBM achieved a 70% win rate against other models. When combined, JFA reached an 81% win rate, comparable to a human amateur 4-dan player.

How does JFA address hardware constraints?

JFA employs parallel Monte Carlo tree search, hierarchical network structures, and human knowledge-based pruning methods. These techniques reduce computational overhead and training time, enabling effective performance under limited hardware resources typical of laboratory environments.

What is the significance of this research?

This work demonstrates that AI can be significantly enhanced for a cultural board game without massive hardware infrastructure. It won first place at the 2024 China National Computer Game Tournament and contributes to the preservation and transmission of traditional culture through advanced AI algorithms.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF