Fast-ThinkAct:可言语化的 latent 规划,延迟降 89.3%

本页是 VLM 自我改进 演化轴第⑤阶段的收官, 也是与本方向 motivation 最接近的一篇: Fast-ThinkAct(Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning)[1](CVPR 2026)。

1. 问题

reasoning VLA 的显式 CoT 带来泛化,也带来致命的:

explicit CoThigh inference latency \text{explicit CoT} \rightarrow \underline{\text{high inference latency}}

2. 方法

把 teacher 的冗长 reasoning r1:Tr_{1:T} 蒸馏成 verbalizable latent reasoning(可言语化的 latent CoT), 偏好引导目标对齐操作轨迹,同时迁移语言与视觉两种规划能力, action policy 基于 latent plan 执行。

3. 结果

  • 推理延迟最高降低 89.3%(相对 SOTA reasoning VLA);
  • 保持长 horizon 规划、few-shot 适应与 failure recovery

4. 与本方向只差一步

expensive reasoningpersistent compact capability \underline{ \text{expensive reasoning} \rightarrow \text{persistent compact capability} }

区别在 teacher 是谁:

teacher 训练发生地
Fast-ThinkAct 离线 teacher 模型 离线 GPU
本方向 设备自己的 past-self(自己过去的成功 reasoning) 端侧 LoRA

把 teacher 换成 past-self、训练搬上设备 (端侧训练), "第一次想很久、第二次不用想"就闭环了。

参考文献

[1] HUANG C P, MAN Y Z, YU Z D, et al. Fast-ThinkAct: efficient vision-language-action reasoning via verbalizable latent planning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2026. https://arxiv.org/abs/2601.09708


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""