• Introduces a novel Video to Command framework that integrates multiple data associations and physical constraints to enhance robot learning from demonstrations.
• Proposes an object-level appearance-contrasting multiple data association strategy to robustly track manipulated objects in visually complex environments.
• Develops a multi-task Video to Command model with a hybrid loss function that ensures generated commands are physically feasible and task-appropriate.
• Achieves over 10% improvement in BLEU_N, METEOR, ROUGE_L, and CIDEr metrics compared to state-of-the-art methods, validated on a dual-arm robot prototype.
Download Full PDF: Learning Manipulation from Expert Demonstrations Based on Multiple Data Associations and Physical Constraints | SinoTechIntel | SinoTechIntel