Key Takeaways & Executive Findings
- •• Proposes a training–synthesizing framework that integrates learned gait-conditioned policies into a single multiskill locomotion policy. • Achieves low-cost, smooth gait switching and controllable gaits via reinforcement learning without manual tuning. • Demonstrates seamless gait transitions while maintaining energy optimality across all velocity commands. • Addresses energy efficiency, robustness, and hardware safety in learned multigait control for quadruped robots.
Abstract
Quadruped robots are able to exhibit a range of gaits, each with its own traversability and energy efficiency characteristics. By actively coordinating between gaits in different scenarios, energy-efficient and adaptive locomotion can be achieved. This study investigates the performances of learned energy-efficient policies for quadrupedal gaits under different commands. We propose a training–synthesizing framework that integrates learned gait-conditioned locomotion policies into an efficient multiskill locomotion policy. The resulting control policy achieves low-cost smooth switching and controllable gaits. Our results of the learned multiskill policy demonstrate seamless gait transitions while maintaining energy optimality across all commands.
1. Introduction
Quadruped robots have successfully navigated complex environments using various control approaches, but their adaptability and efficiency still fall short compared to biological animals. In addition to differences in body structures, a significant reason is that animals can easily adopt the most suitable gait pattern and frequency and switch between gaits rapidly, smoothly, and robustly. Different gaits enable animals to effectively handle diverse terrain conditions at different speeds while maintaining good energy efficiency (Hildebrand, 1965; Hoyt and Taylor, 1981). As quadruped robots continue to be deployed in complex and unstructured environments, they will inevitably encounter new challenges that require emulating the strategies used by their natural counterparts.
Controllable gaits and active switching capabilities offer significant advantages for quadruped robot control (Haynes and Rizzi, 2006; Hsiao-Wecksler et al., 2010). By incorporating higher-level decision inputs, robots can activate different gaits in real time (Xi et al., 2016). Moreover, these capabilities enable robots to go beyond locomotion and perform tasks such as dancing and leaping using specially designed gait controllers (Margolis and Agrawal, 2022).
However, the integration of multiple gaits to improve adaptability has been largely unexplored in existing research. Most existing frameworks either restrict locomotion to a single predefined gait type or disregard gait controllability altogether by leaving gait selection to the policy. Although this approach prioritizes simplicity and ease of tuning, it sacrifices optimality and controllability. As manual fine-tuning is typically required for multigait integration, many researchers opt not to explore this avenue.
Loading authentic research manuscript (Pages 1–5)...
Zhicheng WANG, Xin ZHAO, Meng Yee (Michael) CHUAH, Zhibin LI, Jun WU, Qiuguo ZHU (2025). Efficient learning of robust multigait quadruped locomotion for minimizing the cost of transport. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2401070
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the main contribution of this paper?
The paper proposes a training–synthesizing framework that integrates learned gait-conditioned locomotion policies into a single multiskill policy for quadruped robots. This enables efficient, smooth, and controllable gait switching while minimizing the cost of transport.
How does the proposed framework achieve efficient multigait locomotion?
The framework first trains single-gait policies conditioned on gait phase and velocity commands, then uses reinforcement learning to train a gait selector module. This integrates the skills into a multiskill policy that transitions seamlessly between gaits while maintaining energy optimality.
What are the advantages of learned gait-conditioned policies over traditional methods?
Learned policies eliminate the need for time-consuming manual tuning of heuristics and oscillators used in central pattern generators. They automatically discover optimal gait coordination strategies, providing smoother transitions, better controllability, and improved energy efficiency.
How does the policy maintain energy optimality across different commands?
The multiskill policy is trained to minimize the cost of transport (energy per distance) while responding to velocity commands. It learned to select and switch between gaits optimally, ensuring that the robot always uses the most energy-efficient gait for the given speed.
What are the practical applications of this research?
This research enhances the adaptability and efficiency of quadruped robots in unstructured environments. It enables real-time gait switching for tasks like search and rescue, inspection, and exploration, where energy efficiency and robust locomotion are critical.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena