CoT-VLA:未来子目标图像作为视觉 CoT

本页是 VLM 自我改进 演化轴第⑤阶段的一篇: CoT-VLA(Visual Chain-of-Thought Reasoning for Vision-Language-Action Models)[1](CVPR 2025)——证明 reasoning 不一定只能是文本

1. 方法

VLA 直接 observationactionobservation \rightarrow action 缺乏时间规划。 CoT-VLA 让模型先生成:

future subgoal image \underline{\text{future subgoal image}}

未来图像本身当作 visual chain-of-thought, 再基于它生成 action——推理在图像空间中进行。

2. 对本方向的启发

设备 trajectory 里值得固化的 reasoning target 不必是文字 lesson:

 recovery screenshot / state \text{关键 recovery screenshot / state}

本身就可以是需要固化的"推理中间产物"—— 比如"恢复后的正确文件选择器长什么样"。 这与 ThinkAct 的 visual plan latent 一起构成图像/latent 两级非文本 reasoning 形态。

参考文献

[1] ZHAO Q Q, LU Y, KIM M J, et al. CoT-VLA: visual chain-of-thought reasoning for vision-language-action models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2025. https://openaccess.thecvf.com/content/CVPR2025/html/Zhao_CoT-VLA_Visual_Chain-of-Thought_Reasoning_for_Vision-Language-Action_Models_CVPR_2025_paper.html


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""