Key Takeaways & Executive Findings
- •• The proposed query-selection encoder (QSE) significantly accelerates training convergence and improves detection accuracy for end-to-end object detectors. • The hierarchical feature-aware attention (HFA) mechanism suppresses similar feature representations and highlights discriminative ones, expediting feature selection. • QSE is versatile and can be seamlessly integrated into both CNN- and Transformer-based detection architectures. • Extensive experiments on MS COCO, CrowdHuman, and PASCAL VOC demonstrate that QSE enhances end-to-end performance with fewer training epochs.
Abstract
End-to-end object detection methods have attracted extensive interest recently since they alleviate the need for complicated human-designed components and simplify the detection pipeline. However, these methods suffer from slower training convergence and inferior detection performance compared to conventional detectors, as their feature fusion and selection processes are constrained by insufficient positive supervision. To address this issue, we introduce a novel query-selection encoder (QSE) designed for end-to-end object detectors to improve the training convergence speed and detection accuracy. QSE is composed of multiple encoder layers stacked on top of the backbone. A lightweight head network is added after each encoder layer to continuously optimize features in a cascading manner, providing more positive supervision for efficient training. Additionally, a hierarchical feature-aware attention (HFA) mechanism is incorporated in each encoder layer, including in- and cross-level feature attention, to enhance the interaction between features from different levels. HFA can effectively suppress similar feature representations and highlight discriminative ones, thereby accelerating the feature selection process. Our method is highly versatile in accommodating both CNN- and Transformer-based detectors. Extensive experiments were conducted on the popular benchmark datasets MS COCO, CrowdHuman, and PASCAL VOC to demonstrate the effectiveness of our method. The results showed that CNN- and Transformer-based detectors using QSE can achieve better end-to-end performance within fewer training epochs.
1. Introduction
Object detection is a crucial task in computer vision, aiming to find targets of interest in images by circling bounding boxes and predicting categories (Pu et al., 2021; Qin et al., 2023; Wang CY et al., 2023). Traditional object detectors built by convolutional neural networks (CNNs) adopt a dense prediction paradigm, which perform classification and localization tasks based on pre-defined densely tiled bounding boxes (Girshick, 2015; Ren et al., 2015; Lin TY et al., 2017) or grid points in the two-dimensional (2D) image plane (Tian et al., 2019; Zhou et al., 2019). One-to-many label assignments are the core scheme of these methods, in which each ground-truth box is assigned to multiple predictions of detectors as the supervised target. Despite their excellent performance, these detectors rely heavily on hand-designed components, i.e., non-maximum suppression (NMS), to remove duplicated predictions during inference, which introduces additional hyper-parameters to tune and thus causes sub-optimal performance in dense scenes (Li S et al., 2023; Zhang SL et al., 2023).
To achieve a more flexible end-to-end detection, DEtection TRansformer (DETR) (Carion et al., 2020) was proposed, viewing object detection as a set prediction problem and introducing a Transformer encoder–decoder architecture. Adopting a sparse prediction paradigm, DETR reasons about the global image context and outputs the final predictions by using a small set of learnable object queries. The one-to-one label assignment plays a crucial role in DETR for conducting end-to-end detection, where each ground-truth box is assigned only one prediction. Hence, DETR outputs only a single prediction for each object during inference and NMS is no longer necessary. This approach has encouraged many subsequent improvements (Yao et al., 2021; Wang YN et al., 2022; Li F et al., 2023). In addition, POTO (Wang JF et al., 2021) and OneNet (Sun PZ et al., 2021b) attempt to adopt one-to-one label assignment in CNN-based detectors to realize end-to-end detection. However, these methods suffer from extremely slow training convergence and relatively low performance on small objects. The core reason for this problem is a conflict between one-to-one label assignment and sufficient positive supervision (Jia et al., 2023; Hou et al., 2024). During the training process, the one-to-one matching scheme assigns a single positive prediction to each ground-truth box, which leads to negative predictions dominating most of the loss function, causing insufficient positive supervision. Therefore, more training iterations are required for convergence. To alleviate this issue, previous studies introduced additional training-only architectures, such as query denoising (Li F et al., 2022), mul...
Loading authentic research manuscript (Pages 1–5)...
Zuyi WANG, Zhimeng ZHENG, Jun MENG, Li XU (2025). End-to-end object detection using a query-selection encoder with hierarchical feature-aware attention. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400960
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What problem does the query-selection encoder (QSE) solve?
QSE addresses the slow training convergence and inferior detection performance of existing end-to-end detectors by providing more positive supervision through cascading lightweight head networks after each encoder layer.
How does hierarchical feature-aware attention (HFA) improve feature selection?
HFA incorporates in- and cross-level feature attention to suppress similar feature representations and highlight discriminative ones, which accelerates the feature selection process and enhances interaction between features from different levels.
Which benchmark datasets were used to evaluate QSE?
Extensive experiments were conducted on MS COCO, CrowdHuman, and PASCAL VOC, demonstrating that QSE improves end-to-end detection performance within fewer training epochs.
Is QSE compatible with CNN-based detectors?
Yes, QSE is designed to be versatile and can be integrated into both CNN-based detectors (e.g., POTO, OneNet) and Transformer-based detectors (e.g., DETR), improving their end-to-end performance.
What is the main advantage of QSE over traditional end-to-end detectors?
QSE accelerates training convergence and boosts detection accuracy by providing richer positive supervision and a more effective feature selection mechanism, while remaining compatible with diverse backbone architectures.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena