SDRT:多样 reasoning 轨迹的自蒸馏

本页是 VLM 自我改进 演化轴第②阶段的第二篇: SDRT(Self-Distillation with Diverse Reasoning Traces)[1](2025)—— 名字里的 self-distillation 与"能力固化"的直觉几乎重合: 自己的 reasoning traces → 自蒸馏 → internalize 推理过程

1. 方法

Mt reasoning-guided responsesself-distillationMt+1 M_t \rightarrow \text{生成多样的 reasoning-guided responses} \rightarrow \text{self-distillation} \rightarrow M_{t+1}

对同一视觉问题采样多条不同推理路径的回答,再蒸馏回模型, 让模型把"多样推理中稳定的部分"内化成自身能力。

2. 两个对本方向关键的细节

  1. 不需要 full fine-tuning:用 intervention adapter 做参数高效更新——这正是设备端 LoRA consolidation 的形状;
  2. 蒸馏目标是自己的输出分布而非老师的,天然规避 teacher–student 分布失配。

3. 与端侧固化的对接

SDRT:reasoning tracesadapter \text{SDRT} \; : \quad \text{reasoning traces} \rightarrow \text{adapter}

real execution failure/recovery trajectoryskill traceon-device LoRA \underline{ \text{real execution failure/recovery trajectory} \rightarrow \text{skill trace} \rightarrow \text{on-device LoRA} }

两者只差一个来源:SDRT 的 reasoning 来自 benchmark 问题上的采样, 不来自真实部署 experience。把采样源换成 数据版图 里的真机轨迹(RealMobile 型), 就是端侧闭环的第一步。

参考文献

[1] WU G D, SONG H, WANG Y W, et al. SDRT: enhance vision-language models by self-distillation with diverse reasoning traces[J/OL]. arXiv preprint arXiv:2503.01754, 2025. https://arxiv.org/abs/2503.01754


© 2026 Yang Huan · yanghuan9812@qq.com

results matching ""

    No results matching ""