xrmeng-dev
您好,感谢开源 VLA-Pruner 的代码。 我在阅读实现时注意到一个可能和论文方法描述不一致的地方,想请教一下 temporal decay 的权重方向是否符合预期。 在 `src/openvla/prismatic/extern/hf/modeling_prismatic.py` 中,历史 action-to-vision attention 的加权方式如下: ```python weights = np.array([self.av_decay ** i for i in range(len(self.av_hist)-1, -1, -1)], dtype=np.float32) guided = np.zeros(256, dtype=np.float32) for i in range(len(weights)): guided += weights[i] * self.av_hist[-1 - i] guided = guided / np.sum(guided) ``` 如果我对历史队列的理解正确,`self.av_hist[-1]` 表示最近一次的 action-to-vision attention,`self.av_hist[-2]` 表示上一次,依此类推。 以 `temporal_w = 3`、`gamma = 0.8` 为例,当前代码会得到: ```text weights = [0.64, 0.8, 1.0] self.av_hist[-1] -> 0.64 # 最新 self.av_hist[-2] -> 0.8 self.av_hist[-3] -> 1.0 # 最旧 ``` 也就是说,当前实现会让最旧的历史 attention 获得最大权重,而最新的历史 attention 获得最小权重。 根据我对论文中 temporal decay 设计的理解,似乎更自然的方向应该是最近的历史 attention 权重最大,越早的历史 attention 权重越小,例如: ```text self.av_hist[-1] -> gamma^0 = 1.0 self.av_hist[-2] -> gamma^1 = 0.8 self.av_hist[-3] -> gamma^2 = 0.64 ``` 想请问一下,论文中 temporal-aware pruning 的预期权重方向是哪一种?当前代码让更早的历史 attention 权重更大,是有意这样设计的吗?