SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2500100Original Research

Mind the Gap: towards generalizable autonomous penetration testing via domain randomization and meta-reinforcement learning

Shicheng Zhou¹,Jingju Liu¹,Yuliang Lu¹,Jiahai Yang¹,Yue Zhang¹,Jie Chen¹

National University of Defense Technology, Hefei, China

Read Executive PreviewQuick FAQ
Mind the Gap: towards generalizable autonomous penetration testing via domain randomization and meta-reinforcement learning
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:October 10, 2025Edition:Vol. 32, Issue 10 • pp. 301-313Citation:Shicheng Zhou et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:autonomous penetration testingreinforcement learningdomain randomizationmeta-reinforcement learninglarge language modelcybersecuritygeneralizationnetwork security

Key Takeaways & Executive Findings

  • • Proposes GAP, a generalizable autonomous penetration testing framework combining a real-to-sim-to-real pipeline with domain randomization and meta-reinforcement learning. • Addresses the training environment dilemma by enabling efficient policy learning in realistic environments through synthetic environment generation. • Introduces a large language model-powered domain randomization method for creating diverse training environments to improve generalization. • Demonstrates zero-shot policy transfer in similar environments and rapid policy adaptation in dissimilar environments across various vulnerable virtual machines.
Sponsored Research Highlight

Abstract

With the increasing number of vulnerabilities exposed on the Internet, autonomous penetration testing (pentesting) has emerged as a promising research area. Reinforcement learning (RL) is a natural fit for studying this topic. However, two key challenges limit the applicability of RL-based autonomous pentesting in real-world scenarios: the training environment dilemma—training agents in simulated environments is sample-efficient while ensuring that their realism remains challenging; poor generalization ability—agents’ policies often perform poorly when transferred to unseen scenarios, with even slight changes potentially causing a significant generalization gap. To address both challenges, we propose a generalizable autonomous pentesting framework termed GAP, which aims to achieve efficient policy training in realistic environments and train generalizable agents capable of drawing inferences about other cases from one instance. GAP introduces a real-to-sim-to-real pipeline that enables end-to-end policy learning in unknown real environments while constructing realistic simulations and improves agents’ generalization ability by leveraging domain randomization and meta-RL learning. We are among the first to apply domain randomization in autonomous pentesting and propose a large language model-powered domain randomization method for synthetic environment generation. We further apply meta-RL to improve agents’ generalization ability in unseen environments by leveraging synthetic environments. Combining the two methods effectively bridges the generalization gap and improves agents’ policy adaptation performance. Simulations are conducted on various vulnerable virtual machines, with results showing that GAP can enable policy learning in various realistic environments, achieve zero-shot policy transfer in similar environments, and achieve rapid policy adaptation in dissimilar environments.

1. Introduction

Penetration testing (pentesting) is an authorized simulated cyberattack methodology for identifying security vulnerabilities, allowing organizations to proactively enhance their defenses. However, as more vulnerabilities are exposed on the Internet, traditional manual-based pentesting becomes costly, time-consuming, and personnel-constrained. Autonomous pentesting has emerged as a promising research area, and reinforcement learning (RL) is a suitable method for optimizing the sequential decision-making process involved. RL trains agents through trial and error without needing predefined environmental models, making it a natural fit.

RL-based autonomous pentesting aims to train agents to explore and exploit vulnerabilities in target hosts. Yet, RL algorithms typically require large numbers of training samples, which is impractical in real-world environments where interactions are time-consuming and risky. A common solution is to train agents in simulated or emulated environments and then transfer the learned policies to real scenarios. This approach introduces two key challenges: the training environment dilemma (balancing realism and efficiency) and poor generalization ability.

Challenge 1—the training environment dilemma—represents a conflict between the realism of the training environment and the efficiency of the training process. Prior work has used simulated environments such as network simulators (e.g., NASim), which are sample-efficient but lack realism. Alternatively, training in realistic environments is more faithful but can be inefficient and risky. GAP aims to overcome these limitations by integrating domain randomization and meta-RL to train generalizable policies in realistic environments.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Shicheng Zhou, Jingju Liu, Yuliang Lu, Jiahai Yang, Yue Zhang, Jie Chen (2025). Mind the Gap: towards generalizable autonomous penetration testing via domain randomization and meta-reinforcement learning. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2500100
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is GAP in autonomous penetration testing?

GAP (Generalizable Autonomous Pentesting) is a framework proposed to achieve efficient policy training in realistic environments and improve agents' generalization ability. It combines domain randomization and meta-reinforcement learning to bridge the gap between simulated training and real-world deployment.

How does GAP address the training environment dilemma?

GAP introduces a real-to-sim-to-real pipeline that enables end-to-end policy learning in unknown real environments while constructing realistic simulations. This allows agents to train efficiently in simulated environments while maintaining realism, thus resolving the realism-efficiency trade-off.

What is domain randomization in this context?

Domain randomization is a technique that generates diverse synthetic environments to improve the generalization of policies. In GAP, a large language model powers the domain randomization method to create varied training environments, helping agents adapt to unseen scenarios.

How does meta-reinforcement learning improve policy adaptation?

Meta-RL enables agents to learn how to learn, allowing them to quickly adapt their policies to new environments with minimal fine-tuning. GAP leverages meta-RL on synthetic environments to achieve rapid policy adaptation in dissimilar real-world scenarios.

What are the main challenges in RL-based autonomous pentesting?

The main challenges are the training environment dilemma (balancing sample efficiency and realism) and poor generalization ability (policies failing to transfer to unseen scenarios). GAP addresses both by combining domain randomization and meta-RL, demonstrating zero-shot transfer in similar environments and rapid adaptation in dissimilar ones.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF