VL-Rethinker:用 RL 把自我反思固化进 policy

本页是 VLM 自我改进 演化轴第③阶段的核心: VL-Rethinker [1](NeurIPS 2025)——明确说不做 reasoning distillation,而是用 RL 直接把"反思行为"固化进模型 policy。

1. 方法

GRPO 的 rollout 之后,强制加入一个 rethinking trigger, 迫使模型走完:

initial reasoningreflectioncorrection \text{initial reasoning} \rightarrow \text{reflection} \rightarrow \text{correction}

整条"先想一遍 → 回头检查 → 改正"的行为被奖励信号强化, 最终成为模型自身的输出习惯。

2. 对本方向的关键启发

reflection capabilityθ \underline{ \text{reflection capability} \rightarrow \theta }

它回答的问题与 RSI 的核心追问几乎同构:

"第一次失败后进行了 recovery,下次是不是直接知道要检查什么?"

VL-Rethinker 证明反思可以被 RL 固化—— 但它的反思由外部 trigger 强制触发, 且发生在 benchmark QA 上;RSI 想要的是设备在真实执行中 自己意识到该反思(见 数据版图 的失败定位问题)。

参考文献

[1] WANG H Z, QU C, HUANG Z M, et al. VL-Rethinker: incentivizing self-reflection of vision-language models with reinforcement learning[C]//Advances in Neural Information Processing Systems 38 (NeurIPS). 2025. https://proceedings.neurips.cc/paper_files/paper/2025/hash/2c84844a559e4f962752570bff456ae4-Abstract-Conference.html


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""