Missing evaluations from paper

#2 · closed · 10 comments

View on GitHub ↗

wang-first

Hi, thanks for the great work! I noticed that the paper includes several evaluations that don’t seem to be available in the current codebase (e.g., quantization, the Qwen model, RULER 32K, and efficiency experiments). Are there any plans to release these in a future update? Thanks again for the great work.

Comments

FFY0

Yes, we will update. We expect the Ruler update to be available soon, likely this weekend. I will upload the 32K Ruler dataset to Hugging Face. In addition, for the efficiency experiments, I will provide a quick test script. As for the Qwen model and quantization, the issue is caused by differences between the Transformers version and the current codebase version. My collaborator is currently working on it, and it will be merged into the branch once it is finished.

FFY0

Hi, the Ruler dataset has been uploaded and is now available for easy download and use. I have also provided a quick efficiency evaluation script that enables one-click assessment of various methods. 😊 Feel free to reach out if you have any questions or feedback!

FFY0

Hello, my collaborator has uploaded a new update: Upgrade Transformers to v4.57. It now supports the Qwen-3 family, including both dense models and MoE models. If you have any questions, feel free to reach out. ^_^

wang-first

Hi, thank you for your quick response and for updating the code. I tried running the script, but I encountered the following error. This occurred when running cake_global on the LongBench dataset (50% subset, 80% compression) using the Qwen3-30B-A3B-Instruct-2507 model. Could you please let me know if this is an actual issue or just a false alarm? Thanks again for the great work. <img width="1525" height="269" alt="Image" src="https://github.com/user-attachments/assets/fbdd911b-53cf-4a0f-a29f-45f563d9afb0" />

adfh917k

> Hi, thank you for your quick response and for updating the code. > > I tried running the script, but I encountered the following error. This occurred when running cake_global on the LongBench dataset (50% subset, 80% compression) using the Qwen3-30B-A3B-Instruct-2507 model. > > Could you please let me know if this is an actual issue or just a false alarm? > > Thanks again for the great work. > > <img alt="Image" width="1525" height="269" src="https://private-user-images.githubusercontent.com/266317604/566268299-fbdd911b-53cf-4a0f-a29f-45f563d9afb0.PNG?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NzM5MzE0MTUsIm5iZiI6MTc3MzkzMTExNSwicGF0aCI6Ii8yNjYzMTc2MDQvNTY2MjY4Mjk5LWZiZGQ5MTFiLTUzY2YtNGEwZi1hMjlmLTQ1ZjU2M2Q5YWZiMC5QTkc_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTYmWC1BbXotQ3JlZGVudGlhbD1BS0lBVkNPRFlMU0E1M1BRSzRaQSUyRjIwMjYwMzE5JTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3QmWC1BbXotRGF0ZT0yMDI2MDMxOVQxNDM4MzVaJlgtQW16LUV4cGlyZXM9MzAwJlgtQW16LVNpZ25hdHVyZT0yMzE5Zjk2NmNlMWIxN2RlYTIwM2NjYmY2YTI4Zjk3YzYxMGRmM2Q2NzhkNWYwYTRmY2EzYjY2ZWQ1ODIzMDRkJlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdCJ9.-T5PHe2Ry3UH9QJHJOFtXZ9d2EQLoHBMZxRG36lyMSY"> Thank you for pointing out this bug. It has been fixed and cake_global can now run successfully.

wang-first

Hi, I’m trying to compress questions into the prompt (similar to the approach used in CAKE or SnapKV) by setting `compress_questions = True`. However, I encounter the following error when I do so. Could you help me understand why this happens and how I might resolve it? Thank you for your time and for your work on this project. <img width="1534" height="60" alt="Image" src="https://github.com/user-attachments/assets/4446840f-171c-40a3-982f-ae88dda85474" />

FFY0

Hello, Our experiments in the paper are evaluated with compression_quesstions=False to better reflect realistic and more complex compression scenarios. This bug may be related to the recent commit that added support for Qwen-MoE. We suggest switching the model to Llama-3.1-8B to check whether the evaluation can run successfully. In addition, could you provide more detailed error messages and your current evaluation script, including the compression method and model being evaluated? We plan to start working on a fix for this bug in the next few days.

wang-first

Hi, This is the full error message ``` /home/miniconda3/envs/myenv/lib/python3.10/site-packages/jieba/_compat.py:18: UserWarning: pkg_resources is deprecated as an API. See https://setuptools.pypa.io/en/latest/pkg_resources.html. The pkg_resources package is slated for removal as early as 2025-11-30. Refrain from using this package or pin to Setuptools<81. import pkg_resources Model will not return attentions in its output to save memory. Set output_attentions=True if attentions are needed in the output. `torch_dtype` is deprecated! Use `dtype` instead! `torch_dtype` is deprecated! Use `dtype` instead! `torch_dtype` is deprecated! Use `dtype` instead! try save to: /home/MyWS/DefensiveKV/evaluation/results/longbench__Qwen3-14B__EfficientLayerDefensiveKV=0.8_win=32_kerl=5__0.8__frac1.00.csv Loading from disk, data_dir: /home/MyWS/DefensiveKV/defensivekv_dataset/longbench/ Replacing vanilla flash attention in Qwen/Qwen3-14B with flash_attn_varlen_func for head-specific compression support. Loading checkpoint shards: 0%| | 0/8 [00:00<?, ?it/s] Loading checkpoint shards: 100%|██████████| 8/8 [00:00<00:00, 218.72it/s] Device set to use cuda:0 model dtype: torch.bfloat16 0%| | 0/3731 [00:00<?, ?it/s]Model <class 'transformers.models.qwen3.modeling_qwen3.Qwen3ForCausalLM'> not tested 0%| | 0/3731 [00:24<?, ?it/s] Traceback (most recent call last): File "/home/MyWS/DefensiveKV/evaluation/evaluate.py", line 328, in <module> Fire(evaluate) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/fire/core.py", line 143, in Fire component_trace = _Fire(component, args, parsed_flag_args, context, name) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/fire/core.py", line 477, in _Fire component, remaining_args = _CallAndUpdateTrace( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/fire/core.py", line 693, in _CallAndUpdateTrace component = fn(*varargs, **kwargs) File "/home/MyWS/DefensiveKV/evaluation/evaluate.py", line 278, in evaluate output = pipe( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/pipelines/base.py", line 1467, in __call__ return self.run_single(inputs, preprocess_params, forward_params, postprocess_params) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/pipelines/base.py", line 1474, in run_single model_outputs = self.forward(model_inputs, **forward_params) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/pipelines/base.py", line 1374, in forward model_outputs = self._forward(model_inputs, **forward_params) File "/home/MyWS/DefensiveKV/kvpress/pipeline.py", line 198, in _forward answer = self.generate_answer( File "/home/MyWS/DefensiveKV/kvpress/pipeline.py", line 257, in generate_answer outputs = self.model( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl return forward_call(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/utils/generic.py", line 918, in wrapper output = func(self, *args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/models/qwen3/modeling_qwen3.py", line 480, in forward outputs: BaseModelOutputWithPast = self.model( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl return forward_call(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/utils/generic.py", line 1064, in wrapper outputs = func(self, *args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/models/qwen3/modeling_qwen3.py", line 410, in forward hidden_states = decoder_layer( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/modeling_layers.py", line 94, in __call__ return super().__call__(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl return forward_call(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/utils/deprecation.py", line 172, in wrapped_func return func(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/transformers/models/qwen3/modeling_qwen3.py", line 260, in forward hidden_states, _ = self.self_attn( File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/home/miniconda3/envs/myenv/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl return forward_call(*args, **kwargs) File "/home/MyWS/DefensiveKV/kvpress/ada_attn.py", line 267, in forward query_states = self.q_norm(self.q_proj(hidden_states).view(hidden_shape)).transpose(1, 2) RuntimeError: cannot reshape tensor of 0 elements into shape [1, 0, -1, 128] because the unspecified dimension size -1 can be any value and is ambiguous ``` In this evaluation, I set compress_questions=True and ran with efficient_layer_defensivekv method on longbench dataset using Qwen3-14B model (before this, I move pipe outside of exception block to capture the error message)

FFY0

@wang-first Hi, I’ve fixed this bug using a tricky workaround. I’d recommend taking a quick look at the summary of this fix to make sure it doesn’t affect your reproduction results. The main change is that, under the compress_questions setting, the input now includes a additional separator (instead of an empty string). This effectively avoids the empty question issue, but it does introduce a minor formatting change (e.g., an extra newline in the input). I believe this is a reasonable trade-off, though models can sometimes be sensitive to such details. Let me know if you notice any differences on your side. ``` @@ -196,7 +196,7 @@ def evaluate( if compress_questions: df["context"] = df["context"] + df["question"] - df["question"] = "" + df["question"] = "\n" ``` ## Bug Fix: Forward Pass Error with `compress_questions` in LongBench ### Problem Description When using the `compress_questions` setting, the behavior differs across benchmarks: | Benchmark | Chat Template | `df["question"]` after `compress_questions` | Forward Pass | |---------------|:-------------:|:-------------------------------------------:|:------------:| | **LongBench** | ❌ Disabled | `""` (empty string) | ❌ Error | | **Ruler** | ✅ Enabled | Non-empty | ✅ Normal | **Root Cause**: With chat template disabled in LongBench, resetting the question column to an empty string causes the model to receive empty input during the forward pass, leading to unexpected errors. --- ### Fix ```diff - df["question"] = "" + df["question"] = "\n" ``` Replacing the empty string with a newline character ensures a valid separator is preserved between the concatenated `context` and subsequent content, preventing empty input from reaching the forward pass. --- ### Background: Why Disable Chat Template for LongBench Evaluation? The official LongBench evaluation and prior works (e.g., **SnapKV**) all adopt the **chat-template-disabled** setting. To maintain alignment with these results, we follow the same convention in this work. --- ### Important Notes Although this change may appear trivial, please be aware of the following: - Certain **models and tasks** can be **highly sensitive** to even minor modifications in chat template handling - Changing the separator (`""` → `"\n"`) may introduce **unexpected performance fluctuations** in model evaluation If you still encounter related issues, consider experimenting with alternative separators: ```python df["question"] = "\n" # newline (current fix) df["question"] = " " # whitespace ```

FFY0

The issue should have been resolved by now. If you still encounter any problems, please feel free to reopen the issues.