• Proposes GAP, a generalizable autonomous penetration testing framework combining a real-to-sim-to-real pipeline with domain randomization and meta-reinforcement learning.
• Addresses the training environment dilemma by enabling efficient policy learning in realistic environments through synthetic environment generation.
• Introduces a large language model-powered domain randomization method for creating diverse training environments to improve generalization.
• Demonstrates zero-shot policy transfer in similar environments and rapid policy adaptation in dissimilar environments across various vulnerable virtual machines.