• Proposes GMCoT, a novel graph-augmented multimodal chain-of-thought framework for multi-label zero-shot learning.
• Integrates label graphs into LLM reasoning to model complex semantic relationships among labels.
• Mitigates cross-modal semantic gaps by combining multimodal large language models with graph-based structures.
• Outperforms state-of-the-art methods on benchmark datasets for multi-label zero-shot learning.