Key Takeaways & Executive Findings
- •• A de-blocking adaptive feedback control (AFC) design is proposed for shared-buffer CIOQ switching, achieving theoretical 100% non-blocking via a credit timeout detection mechanism (CTDM) that eliminates head-of-line blocking. • The combination of VOQ dynamic regulation (VDRA) and threshold dynamic adaptive (TDAA) algorithms effectively mitigates congestion spreading from buffer overflow, enhancing system stability. • Experimental results show a maximum throughput of 1499.66 Gb/s, minimum latency of 83 ns, effective throughput ratio of 96.94%, and packet loss rate as low as 0.6% under typical traffic. • Compared to traditional CIOQ and IQ architectures, throughput improves by 15.12% and 20.55%, while forwarding latency is reduced by 26.9% and 54.7%, respectively, demonstrating superior performance for complex data exchange demands.
Abstract
To address the issues of head-of-line (HOL) blocking at the virtual output queue (VOQ) level, packet loss, and congestion spreading caused by buffer overflow in the shared-buffer-based combined input and output queued (CIOQ) switching architecture, while enhancing its performance and stability, we propose a de-blocking adaptive feedback control (AFC) design in this study. The introduction of the credit timeout detection mechanism (CTDM) enables the CIOQ to achieve theoretical 100% non-blocking state, effectively eliminating the impact of HOL blocking. With the combined effect of the proposed VOQ dynamic regulation algorithm (VDRA) and threshold dynamic adaptive algorithm (TDAA), it can reduce the risk of congestion spreading caused by buffer overflow and consequently improve the overall performance of the system. Both theoretical analysis and experimental results demonstrate that, under typical traffic conditions, the proposed design achieves a maximum throughput of 1499.66 Gb/s and a minimum latency of 83 ns. Additionally, the effective throughput ratio reaches 96.94%, with a data link layer packet (DLLP) loss ratio of merely 0.61% and a packet loss rate as low as 0.6%. In comparison with traditional CIOQ and input queued (IQ) switch architectures, the proposed design demonstrates improvements in throughput by 15.12% and 20.55%, and forwarding latency is reduced by 26.9% and 54.7%, respectively, and the system stability is stronger, which can fully satisfy the demand for data exchange in complex situations.
1. Introduction
As the third-generation high-speed serial expansion bus standard succeeding instruction set architecture (ISA) and peripheral component interconnect (PCI), peripheral component interconnect express (PCIe) delivers high bandwidth, low latency, and reliability. These advantages enable its ubiquitous adoption in core infrastructure domains including hyperscale data centers (HDC), hyper-converged infrastructure (HCI), high-performance computing (HPC), and artificial intelligence (AI). In practice, however, there is a significant mismatch between the physical link speed and the forwarding performance of the switching architecture.
On the one hand, with the iterative evolution of the PCIe technology standard, its physical layer link speed continues to break through, and the bidirectional theoretical throughput can be as high as 512 GB/s in the PCIe 7.0 protocol. On the other hand, the existing switching equipment packet processing capability exhibits obvious technical lag, and its actual forwarding rate and throughput cannot effectively match the high-speed transmission capacity provided by the underlying physical link. This technological development imbalance leads to the transmission pressure of the entire network gradually shifting to the switching architecture of the core node, which handles the packet forwarding performance and load stability issues, becoming a key bottleneck restricting the performance of the entire network.
The main task of the switching architecture is to provide a way to route packets from the input port to the output port quickly and efficiently. As a null-separated non-blocking switching architecture, the crossbar switching architecture has been widely used in various routers and switches. It is functionally equivalent to multiple parallel buses; it can simultaneously complete the matching between multiple input and output ports in parallel, and compared to the bus-based switching architecture, it has stronger switching capability. In the crossbar switching architecture implementation, the combined input and output queued (CIOQ)-type architecture with shared-buffer fully integrates the advantages of traditional input queued (IQ) and output queued (OQ) switching architecture. By setting up dual buffers at the input and output ends, it not only provides throughput close to IQ and OQ in a diverse input traffic environment, but also effectively balances the traffic and mitigates...
Loading authentic research manuscript (Pages 1–5)...
Rui Zheng, Jianliang Shen, Fan Zhang, Ping Lv, Peijie Li, Yu Shao, Zhengbin Zhu (2025). De-blocking adaptive feedback control design for shared-buffer CIOQ switching architecture. Engineering Information Technology & Electronic Engineering. https://doi.org/10.1631/ENG_ITEE_2025_0180
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is the main contribution of the paper?
The paper proposes a de-blocking adaptive feedback control (AFC) design for shared-buffer-based combined input and output queued (CIOQ) switching architecture, incorporating a credit timeout detection mechanism (CTDM) and dynamic regulation algorithms to achieve theoretical 100% non-blocking and improve performance.
How does the proposed design improve throughput and latency?
Under typical traffic conditions, the design achieves a maximum throughput of 1499.66 Gb/s, minimum latency of 83 ns, effective throughput ratio of 96.94%, and low packet loss rates. Compared with traditional CIOQ and IQ architectures, throughput is improved by 15.12% and 20.55%, while latency is reduced by 26.9% and 54.7%, respectively.
What is the role of the credit timeout detection mechanism (CTDM)?
CTDM enables the CIOQ to achieve a theoretical 100% non-blocking state by effectively eliminating the impact of head-of-line (HOL) blocking at the virtual output queue (VOQ) level.
How does the design address congestion spreading?
The combination of the VOQ dynamic regulation algorithm (VDRA) and threshold dynamic adaptive algorithm (TDAA) reduces the risk of congestion spreading caused by buffer overflow, thereby enhancing system stability and overall performance.
What is the significance of this work for PCIe-based networks?
The proposed switching architecture addresses the mismatch between physical link speed and forwarding performance in PCIe-based systems, enabling high-speed data exchange in complex scenarios such as data centers, HPC, and AI infrastructure.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena