Back
精读 dVLA:把视觉子目标、文本推理和机器人动作统一为离散扩散去噪,并用 Prefix Mask 与 KV Cache 将推理速度提升到约 2 倍。
paper deep dive
dvla
diffusion vla
multimodal chain-of-thought