Back
精读 CAIP:用第一视角人类视频中的手部姿态作为动作代理,通过动作-图像对比预训练获得更适合灵巧操作的视觉表征。
paper deep dive
caip
visuomotor control
representation learning
dexterous manipulation