Back
精读 Physical Intelligence 的 RECAP:用分布式 value function 估计 advantage,再以 advantage conditioning 从示教、自治轨迹和人工干预中迭代改进 VLA。
paper deep dive
recap
vla
robot reinforcement learning
π₀.₆