SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2400487Original Research

Active inference of protocol state machines from incomplete message domains

Maohua GUO¹,Yuefei ZHU¹,Jinlong FEI¹

Key Laboratory of Cyberspace Security, Ministry of Education, Zhengzhou 450001, China

Read Executive PreviewQuick FAQ
Active inference of protocol state machines from incomplete message domains
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:January 19, 2025Edition:Vol. 32, Issue 1 • pp. 352-364Citation:Maohua GUO et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:protocol reverse engineeringstate machine inferenceactive learningincomplete message domainsL* algorithmminimally adequate teachernetwork securityRTSPSMTP

Key Takeaways & Executive Findings

  • • Introduces a novel active inference method for protocol state machines using the minimally adequate teacher (MAT) framework, addressing incomplete message domains. • Achieves a comprehensive protocol state machine by broadening the input space via session completion and deterministic mutation techniques. • Optimizes the L+ M algorithm with traffic deduplication, expanded prefix tree acceptor construction, response-based query optimization, and random counterexample generation, reducing execution time by ~40.7%. • Demonstrates significant efficiency gains on RTSP and SMTP protocols, cutting connections by ~28.6% and interactions by ~46.6% compared to AALpy.
Sponsored Research Highlight

Abstract

Inferring protocol state machines from observable information presents a significant challenge in protocol reverse engineering (PRE), especially when passively collected traffic suffers from message loss, resulting in an incomplete protocol state space. This paper introduces an innovative method for actively inferring protocol state machines using the minimally adequate teacher (MAT) framework. By incorporating session completion and deterministic mutation techniques, this method broadens the range of protocol messages, thereby constructing a more comprehensive input space for the protocol state machine from an incomplete message domain. Additionally, the efficiency of active inference is improved through several optimizations for the L+ M algorithm, including traffic deduplication, the construction of an expanded prefix tree acceptor (EPTA), query optimization based on responses, and random counterexample generation. Experiments on the real-time streaming protocol (RTSP) and simple mail transfer protocol (SMTP), which use Live555 and Exim implementations across multiple versions, demonstrate that this method yields more comprehensive protocol state machines with enhanced execution efficiency. Compared to the L+ M algorithm implemented by AALpy, Act_Infer achieves an average reduction of approximately 40.7% in execution time and significantly reduces the number of connections and interactions by approximately 28.6% and 46.6%, respectively.

1. Introduction

In complex network environments, ensuring precise and stable data transmission between communication entities relies on standardizing their interactive behaviors through network protocols. The request for comments (RFC, https://www.rfc-editor.org/), sponsored and published by the Internet Society (ISOC), serves as the comprehensive repository for nearly all public Internet standards, encompassing 9565 protocol specification documents as of March 2024. However, the plethora of botnets, Trojans, and malware employing proprietary protocols for communication (Chandler, 2023) poses significant challenges to network security oversight, driving an increased focus on protocol reverse engineering (PRE) in fields such as intrusion detection (Abdulganiyu et al., 2023; Saied et al., 2024), vulnerability mining (Pham et al., 2020; Natella, 2022; Yu et al., 2024), and malware analysis (Antonakakis et al., 2017; de Carli et al., 2017).

PRE (Huang et al., 2022) can be broadly classified into two implementation approaches. The first approach involves analyzing program instructions (Ma et al., 2022), using dynamic taint analysis or symbolic execution techniques for high-precision reverse engineering. However, obstacles such as protection mechanisms, obfuscated binaries, or packaged files often hinder dynamic debugging and the analysis of protocol entities. The second approach centers on analyzing message sequences (Li et al., 2023), typically spanning preprocessing, protocol format inference, and protocol state machine inference stages. Preprocessing lays the groundwork for reverse engineering, including tasks like data clustering (Le et al., 2024). Common clustering algorithms include the nearest neighbor clustering (K-means), unweighted pair-group method with arithmetic means (UPGMA), and density-based spatial clustering of applications with noise (DBSCAN). Protocol format inference is accomplished through static traffic analysis to segment protocol messages (Wang YP et al., 2012; Kleber et al., 2018, 2022; Sun et al., 2019; Chandler et al., 2023), infer protocol keywords (Ye et al., 2021; Chandler, 2023; Tang et al., 2023), and perform semantic analysis (Bermudez et al., 2016; Wang XW et al., 2020; Wang YP et al., 2022). At this stage, the protocol is typically categorized as either a text-based protocol or a binary-based protocol based on whether the messages can be converted into JSON, XML, or other types of textual files. The detailed analysis of the protocol format lays a foundation for the inference stage of the protocol state machine. Protocol state machine inference (Székely et al., 2021; Sun et al., 2022) consolidates all known information and analysis results, employing state machines to formalize and regularize protocol interactive behavior. The research presented in this paper is based on static execution traces.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Maohua GUO, Yuefei ZHU, Jinlong FEI (2025). Active inference of protocol state machines from incomplete message domains. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400487
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is the main contribution of this paper?

The paper proposes an active inference method using the minimally adequate teacher (MAT) framework to construct protocol state machines from incomplete message domains, significantly improving completeness and efficiency.

How does the method handle incomplete protocol state spaces?

It broadens the message domain through session completion and deterministic mutation, allowing the construction of a more comprehensive input space for the state machine.

What optimizations are made to the L+ M algorithm?

The optimizations include traffic deduplication, construction of an expanded prefix tree acceptor (EPTA), response-based query optimization, and random counterexample generation.

How does Act_Infer perform compared to AALpy?

Act_Infer reduces execution time by about 40.7% and reduces connections and interactions by approximately 28.6% and 46.6%, respectively.

On which protocols was the method tested?

The method was validated on the real-time streaming protocol (RTSP) and simple mail transfer protocol (SMTP) using Live555 and Exim implementations across multiple versions.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF