Back
精读 ByteDance Seed 的 GR-RL:用分布式 critic 学习任务进度、过滤次优示范,再通过形态对称增强和 latent-space online RL,让 VLA 学会多眼孔鞋带穿线。
paper deep dive
gr-rl
vision-language-action model
robot reinforcement learning
dexterous manipulation