Vision-Zero:策略性博弈 self-play 的可扩展自改进

本页是 VLM 自我改进 演化轴第④阶段的一篇: Vision-Zero(Scalable VLM Self-Improvement via Strategic Gamified Self-Play)[1](ICLR 2026)——回答"VLM self-improvement 为什么 一定要依赖人工构造 reasoning 数据集"。

1. 方法

任意图片自动生成视觉 deduction games(类似"谁是卧底"), 让多个 agent role 进行 self-play,用游戏 outcome 产生训练信号:

imagesself-generated taskself-play trajectoryRL update \text{images} \rightarrow \text{self-generated task} \rightarrow \text{self-play trajectory} \rightarrow \text{RL update}

2. 对本方向的定位

已经非常像 self-evolution,但本质上仍属于 artificial training environment——模型自己发明游戏来练自己。

Vision-Zero:artificial training environmentvsRSI:deployment is the environment \text{Vision-Zero} \; : \; \text{artificial training environment} \qquad vs \qquad \underline{\text{RSI} \; : \; \text{deployment is the environment}}

本方向的差异点:不需要模型自己发明游戏—— 用户每天使用设备本身就在持续产生 curriculum (任务的重复性、环境的本地性、成败的可验证性,见 数据版图 §2)。

参考文献

[1] WANG Q S, LIU B, ZHOU T Y, et al. Vision-Zero: scalable VLM self-improvement via strategic gamified self-play[J/OL]. arXiv preprint arXiv:2509.25541, 2025. https://arxiv.org/abs/2509.25541


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""