Vision-R1:冷启动 + GRPO 的视觉推理强化

本页是 VLM 自我改进 演化轴第①阶段的第二个节点: Vision-R1 [1](2025)——复刻 DeepSeek-R1 路线的多模态版本: 先造约 200K 多模态 CoT 冷启动数据做 SFT, 再用 GRPO 强化视觉数学 reasoning。

1. 方法

两阶段管线:

MLLM+DeepSeek-R1200K CoT SFTGRPO RL \text{MLLM} + \text{DeepSeek-R1} \;\rightarrow\; \text{200K CoT 冷启动数据} \;\rightarrow\; \text{SFT} \;\rightarrow\; \text{GRPO RL}

一个专门的工程贡献是 Progressive Thinking Suppression: 推理变长以后答案质量会退化,它在 RL 过程中逐步抑制冗长的 thinking 段落长度,让"先长思考、后收敛"可控。

2. 对本方向的定位

适合当 SFT + RL reasoning internalization 的 baseline: 它证明冷启动数据 + 可验证奖励(数学题有标准答案)能激发视觉推理。 但与 self-improvement 的区别一针见血:

teacher experiencemodels own deployed experience \text{teacher experience} \neq \text{model's own deployed experience}

冷启动 CoT 是老师写的,不是模型在部署中撞出来的。

参考文献

[1] HUANG W X, JIA B H, ZHAI Z J, et al. Vision-R1: incentivizing reasoning capability in multimodal large language models[J/OL]. arXiv preprint arXiv:2503.06749, 2025. https://arxiv.org/abs/2503.06749


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""