Yuye541
## OpenVLA-OFT 上 FastV 激进剪枝后的 LIBERO 成功率异常偏高 您好,首先感谢你们开源这项工作以及 OpenVLA-OFT 相关代码。这个项目对复现和分析 VLA token pruning 方法非常有帮助。 我最近在 OpenVLA-OFT 上复现 FastV / VLA-Pruner 的 LIBERO 结果时,遇到了一个现象:在非常激进的 FastV 剪枝配置下,LIBERO 成功率仍然异 常偏高,和论文中的结果不太一致。 ### 实验设置 - 模型:`openvla-7b-oft-finetuned-libero-spatial` - Benchmark:`libero_spatial` - 每个 task 的 episode 数:`50` - checkpoint 选用:openvla-7b-oft-finetuned-libero-spatial ### 运行命令 CUDA_VISIBLE_DEVICES=0 python experiments/robot/libero/run_libero_eval.py \ --use_vla_cache False \ --use_fastv True \ --use_vla_pruner False \ --fastv_attention_source prefill \ --pretrained_checkpoint /file02/user/liyuye/models/openvla-7b-oft-finetuned-libero-spatial \ --task_suite_name libero_spatial \ --fastv_k 3 \ --fastv_r 0.9275 \ --seed 7 \ --run_id_note fastv_7.25% \ --num_trials_per_task 50 按照我的理解,fastv_r=0.9275 应该表示 pruning ratio 为 92.75%,也就是只保留约 7.25% 的 visual tokens。 ### 观察到的现象 在该配置下,已经完成的 LIBERO-Spatial episodes/tasks 成功率仍然是 100%。这个结果明显高于我的预期,也和论文中该剪枝比例下 FastV 的结 果不一致。 同时我也跑了剪枝比例87.5%,结果依旧是偏高(97.6%),相较于您论文报告的数据。 我开启 pruning report 后,得到了如下输出: FastV/VLA-Pruner full-sequence pruning disables KV cache for correctness. LlamaModel is using LlamaSdpaAttention, but `torch.nn.functional.scaled_dot_product_attention` does not support `output_attentions=True`. Falling back to the manual attention implementation, for this forward pass only. [pruning-report] mode=fastv layer=3 seq_len_before=605 seq_len_after=131 image_len=512 kept_visual=38 vision_tokens_before=512 vision_tokens_kept=38 pruned_total=474 image_spans=[(1, 257, 19), (257, 513, 19)] full_layers=4 pruned_layers=28 dense_tflops=4.013935 pruned_tflops=1.247990 flop_ratio=0.310914 flop_saving=68.91% 从这个 log 看,FastV 路径应该是生效的: - mode=fastv - layer=3 - image_len=512 - vision_tokens_kept=38,约为 7.4% visual tokens - seq_len_before=605 -> seq_len_after=131 这个现象不符合论文报告的数据,如果您能告诉我是哪里出了问题,我将非常感谢