• Proposes TFE, a competitive learning paradigm based on DisenIB theory to enhance temporal fidelity in video action recognition without fine-grained supervision.
• Decouples action-relevant semantics from spurious correlations via adversarial feature disentanglement, improving attention alignment.
• Achieves significant accuracy improvements on UCF101, HMDB-51, and Charades benchmarks.
• Addresses the gap between coarse video-level labels and fine-grained temporal dynamics, reducing attention noise in complex scenarios.