Key Takeaways & Executive Findings
- •• Proposes ProC-KD, a novel cross-task knowledge distillation method that removes the label-space constraint between teacher and student networks. • Introduces a prototype learning module to capture invariant intrinsic local object representations from the teacher network. • Develops a task-adaptive feature augmentation module that enhances student features with generalized prototypes to improve generalization. • Demonstrates effectiveness across various visual tasks, confirming the practicality of cross-task knowledge distillation.
Abstract
Recently, large-scale pretrained models have revealed their benefits in various tasks. However, due to the enormous computation complexity and storage demands, it is challenging to apply large-scale models to real scenarios. Existing knowledge distillation methods require mainly the teacher model and the student model to share the same label space, which restricts their application in real scenarios. To alleviate the constraint of different label spaces, we propose a prototype-guided cross-task knowledge distillation (ProC-KD) method to migrate the intrinsic local-level object knowledge of the teacher network to various task scenarios. First, to better learn the generalized knowledge in cross-task scenarios, we present a prototype learning module to learn the invariant intrinsic local representation of objects from the teacher network. Second, for diverse downstream tasks, a task-adaptive feature augmentation module is proposed to enhance the student network features with the learned generalization prototype representations and guide the learning of the student network to improve its generalization ability. Experimental results on various visual tasks demonstrate the effectiveness of our approach for cross-task knowledge distillation scenarios.
1. Introduction
Recently, the Transformer network (Vaswani et al., 2017) has achieved great advances in some visual tasks, for example, image classification (Dosovitskiy et al., 2021; Liu Z et al., 2021; Touvron et al., 2021), object detection (Carion et al., 2020; Zhu XZ et al., 2021; Zhou et al., 2023), image segmentation (Ye LW et al., 2019; Jain et al., 2023), and visual language joint learning (Chen YC et al., 2020; Li LJ et al., 2020; Fu et al., 2023). Based on the self-attention mechanism, Transformer networks can process complete input sequences and possess the advantage of parallelization. Therefore, these networks are usually used to obtain the pretrained model from large-scale datasets (Deng J et al., 2009). Currently, fine-tuning is the common strategy for the utilization of pretrained models in cross-task learning scenarios. After learning the generalized feature representation from large-scale datasets, fine-tuning is performed on the downstream task with the small dataset to boost the performance of the downstream task model. However, applying these large-scale models to practical application scenarios with limited resources (e.g., mobile devices) has become a big challenge due to their enormous computation complexity and huge storage needs.
To solve the above model application issue, some model compression and acceleration technologies have been proposed, for example, parameter pruning (Molchanov et al., 2017; Zhu MH and Gupta, 2018), model quantization (Wu JX et al., 2016), and knowledge distillation (KD) (Hinton et al., 2015). Particularly, KD is a valid approach for model compression, which transfers the knowledge from a large deep neural network into a small one (Hinton et al., 2015). Different from other model compression approaches, KD can decrease the number of parameters and boost the performance of small models on downstream tasks, irrespective of the architectural differences between the teacher network and the student network. It has been successful in a variety of tasks, such as computer vision (Hinton et al., 2015; Romero et al., 2015; Yim et al., 2017; Müller et al., 2019; Park et al., 2019; Gou et al., 2021), natural language processing (Sanh et al., 2019; Sun et al., 2019; Jiao et al., 2020), and speech recognition (Chebotar and Waters, 2016; Kurata and Saon, 2020; Yoon et al., 2021).
However, these KD approaches require mainly the teacher network and the student network to perform the same task; for example, the teacher network and the student network share the same label space, which limits their application in real scenarios, such as downstream tasks in different label spaces as shown in Fig. 1a. By transferring the knowledge from the teacher network to downstream tasks with different label spaces, the cross-task KD method expands the application of the teacher model to a variety of downstream tasks. The existing same-task KD method works mainly to transfer the final prediction logit or the hidden layer knowledge, which is the global-level knowledge alignment and cannot be applied to cross-task KD directly. An earlier work on cross-task KD (Ye HJ et al., 2020) aligns the high-order comparison relationship between models in a local manner; however, this method lags in the representation power of the invariant intrinsic object and is a two-stage distillation method.
Loading authentic research manuscript (Pages 1–5)...
Deng LI, Peng LI, Aming WU, Yahong HAN (2025). Prototype-guided cross-task knowledge distillation. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400383
Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.
Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.
Frequently Asked Questions
What is cross-task knowledge distillation?
Cross-task knowledge distillation transfers knowledge from a teacher model trained on one task to a student model solving a different task, without requiring the same label space. It expands the applicability of pretrained models to diverse downstream tasks.
How does ProC-KD overcome label space constraints?
ProC-KD uses a prototype learning module to extract invariant intrinsic local object representations from the teacher network, which are then transferred via a task-adaptive feature augmentation module to enhance the student's features. This avoids direct alignment of logits or hidden layers that require identical label spaces.
What are the key components of the proposed method?
The method comprises two main modules: a prototype learning module that captures generalized prototype representations from the teacher, and a task-adaptive feature augmentation module that injects these prototypes into the student network to improve its generalization across tasks.
What tasks were evaluated in the experiments?
The paper evaluates ProC-KD on various visual tasks, including image classification, object detection, and image segmentation, demonstrating its effectiveness for cross-task knowledge distillation scenarios.
Why is prototype learning significant in knowledge distillation?
Prototype learning helps distill invariant and intrinsic local-level knowledge from the teacher, which is more transferable across tasks with different label spaces than global-level logits or high-order relationships. This improves the student's ability to generalize to new tasks.
Related Technical Papers & Translations
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption
To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.
A cohesion loss model for determining residual strength of deep bedded sandstone
Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s
Federated model with contrastive learning and adaptive control variates for human activity recognition
Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena