Back
精读 Probe, Learn, Distill:冻结 VLA 通才策略,用残差 RL 探测失败状态,再把分布对齐的恢复轨迹蒸馏回通才模型。
paper deep dive
vla
residual reinforcement learning
robot data generation