SwordFaith/RAVine

โ˜… 9Forks 2PythonGitHub โ†—Compare

README

RAVine

RAVine (a Reality-Aligned eValuation framework for agentic LLMs with search), is a comprehensive evaluation system for agentic search, encompassing the web environment, benchmark datasets, and a novel evaluation method, serving as a full-process, reproducible, and goal-aligned evaluation sandbox.

Features

  • ๐ŸŽฏMore Precise: RAVine provides a more precise and attributable method for nuggets (claim-level ground truth) extraction, the final generated report evaluation is more accurate.
  • โš™๏ธMore Comprehensive: RAVine not only focuses on end-to-end result evaluation, but also designs detailed search process performance indicators.
  • ๐Ÿ’ฐLower Cost: RAVine provides local search APIs and reduces the cost of calling LLM-Judge(Gemini-2.5-Flash) to ~0.01$ per evaluation data.
  • ๐Ÿš€Full Sandbox: We have packaged the search tools and evaluation framework. You only need to provide an LLM to run or evaluate the agent search on RAVine.

Datasets

We have uploaded the running evaluation data, nuggets, corpus index, etc. to hugging face. You can access them and modify the corresponding path variables in the running script:

How to run/evaluate?

First, install the operating environment. This project mainly uses two environments, and we recommend using uv to install the following two environments separately:

  • Env for vllm: the vllm version we use is 0.9.0.1, run pip install vllm==0.9.0.1.
  • Env for /src: the operating environment of our main program has been exported to requirements_agent.txt. Run uv pip install -r requirements_agent.txt.
    • The JDK version we use is 21.0.7, and we recommend that you use this version to support running the BM25 index based on PySerini.

Second, write the configuration file, which is related to the selection and setting of the model, operating environment, index, and file path. You can find examples at configs/ and write your own config here.

Third, run the vllm service and then run the main program. Example instructions are as follows:

bash scripts/server/vllm.sh your_config_file # run the server of agentic llm
bash scripts/evaluation/run.sh your_config_file

If you need to run multiple configurations at once, we provide the corresponding scripts. Please write the configurations to be run into /scripts/evaluation/run_all_eval.sh and then run it.

For more detailed steps, see Evaluation_and_reproduction.

Experimental Results

Table Description:

  • "Rate" denotes the Task Completion Rate.
  • "Comp." refers to the score of Task Completeness.
  • "Rec." and "Prec." represent Recall and Precision, respectively.
  • "URL Err." denotes the URL Error.
  • Latency is measured in seconds, and cost is measured in dollars.
  • Symbols (โ†‘) and (โ†“) indicate that higher or lower values are preferred, respectively.
  • Models not marked with (Thinking) either run without thinking or lack support for the thinking mode.
  • Bold values indicate the best performance for each corresponding metric in the column.

Evaluation results on RAVine, with a maximum context length of 32k and the index built by gte-modernbert-base:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen2.5-7B-Instruct 19.0 6.8 1.9 1.8 7.1 0.01 3.1 18.7 5.7 4.7 8.8 19.8
Qwen2.5-32B-Instruct 71.4 23.0 14.9 16.5 40.3 0.03 4.0 21.1 6.5 4.4 1.4 28.7
Qwen3-4B (Thinking) 92.9 35.4 14.2 11.6 10.2 0.01 2.3 20.9 6.4 5.6 0.0 16.7
Qwen3-4B 0.0 0.0 0.0 0.0 10.1 0.04 6.2 21.3 6.0 4.7 1.3 28.3
Qwen3-8B (Thinking) 86.9 37.8 10.4 12.1 13.9 0.03 6.6 19.7 6.6 5.1 8.9 27.3
Qwen3-8B 28.6 12.4 4.8 6.1 11.2 0.06 9.3 19.3 5.9 5.0 2.4 23.8
Qwen3-32B (Thinking) 98.8 43.5 11.7 15.1 19.6 0.02 2.8 19.2 5.0 4.0 8.9 22.2
Qwen3-32B 85.7 38.0 12.8 12.6 14.6 0.08 8.5 19.1 6.3 5.0 8.1 20.2
Qwen3-30B-A3B (Thinking) 81.0 35.6 10.6 14.2 33.0 0.10 6.6 19.7 6.2 3.6 10.3 29.3
Qwen3-30B-A3B 77.4 30.9 11.3 14.2 15.7 0.07 7.3 16.8 6.2 3.4 0.6 30.4
Qwen3-4B-Instruct-2507 95.2 39.4 7.6 7.0 8.3 0.02 5.9 17.5 6.5 3.8 13.8 27.6
Qwen3-4B-Thinking-2507 98.8 36.6 12.1 15.4 23.0 0.02 2.1 17.8 5.5 4.4 0.0 40.0
Qwen3-30B-A3B-Instruct-2507 83.3 43.1 17.8 17.5 15.8 0.08 8.6 17.6 7.3 4.5 15.6 21.2
Qwen3-30B-A3B-Thinking-2507 96.4 44.1 15.6 19.0 29.7 0.02 3.5 18.1 4.2 3.5 0.0 31.5
LLaMA-3.1-8B-Instruct 96.4 24.0 3.1 3.1 7.3 0.02 2.7 12.1 8.8 6.6 36.8 15.8

Evaluation results on RAVine, with a maximum context length of 128k and the index built by gte-modernbert-base:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen2.5-7B-Instruct 1.2 0.3 0.0 0.0 4.5 0.01 1.6 7.4 2.6 2.3 0.0 33.3
Qwen2.5-32B-Instruct 61.9 24.5 9.7 11.8 17.9 0.03 3.8 19.7 6.0 4.3 3.6 28.3
Qwen3-4B (Thinking) 91.7 37.3 8.2 10.0 12.8 0.02 3.0 19.4 5.8 4.0 1.8 25.5
Qwen3-4B 0.0 0.0 0.0 0.0 15.8 0.09 12.4 13.8 5.8 2.9 6.7 15.2
Qwen3-8B (Thinking) 91.7 41.9 8.3 10.3 65.8 0.40 23.3 12.9 6.3 2.2 5.6 25.1
Qwen3-8B 26.2 10.1 2.6 3.8 139.5 2.52 113.3 8.0 6.3 2.3 4.1 23.8
Qwen3-32B (Thinking) 100.0 45.2 8.4 9.6 23.0 0.02 2.7 18.6 5.2 4.1 0.0 8.6
Qwen3-32B 82.1 35.0 13.2 11.9 22.7 0.42 14.8 15.0 7.0 3.2 6.5 20.3
Qwen3-30B-A3B (Thinking) 81.0 36.8 10.9 12.0 46.9 0.43 12.1 19.0 6.3 4.7 3.8 12.1
Qwen3-30B-A3B 46.4 16.9 6.1 6.7 54.2 0.64 16.7 18.5 6.4 4.8 1.9 23.1
Qwen3-4B-Instruct-2507 96.4 39.7 5.8 4.4 7.8 0.02 5.9 18.3 6.6 4.0 10.1 31.0
Qwen3-4B-Thinking-2507 97.6 35.0 10.4 13.3 22.6 0.01 2.1 19.6 5.9 4.9 0.0 25.0
Qwen3-30B-A3B-Instruct-2507 92.9 49.3 18.0 18.5 21.0 0.32 12.6 17.8 7.4 4.8 6.1 19.2
Qwen3-30B-A3B-Thinking-2507 100.0 46.5 15.9 20.1 31.7 0.02 3.4 19.9 4.6 3.9 0.0 39.6
LLaMA-3.1-8B-Instruct 98.8 25.8 2.8 3.3 4.4 0.02 2.7 12.3 8.0 6.3 58.2 10.9

Evaluation results on RAVine, with a maximum context length of 256k and the index built by gte-modernbert-base:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen3-4B-Instruct-2507 97.6 41.3 6.7 7.6 7.6 0.02 6.0 17.3 6.7 3.9 17.1 28.5
Qwen3-4B-Thinking-2507 100.0 34.1 15.6 16.5 20.5 0.02 2.1 17.8 5.6 4.6 0.0 0.0
Qwen3-30B-A3B-Instruct-2507 91.7 46.8 18.9 18.1 28.8 0.74 14.8 18.5 7.8 5.4 4.5 17.1
Qwen3-30B-A3B-Thinking-2507 100.0 46.2 14.6 17.5 29.4 0.02 3.3 19.6 4.8 4.1 0.0 33.7

Evaluation results on RAVine, with a maximum context length of 32k and the index built by BM25:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen2.5-7B-Instruct 22.6 6.4 4.4 6.3 6.9 0.02 2.9 24.6 12.3 6.9 27.1 29.2
Qwen2.5-32B-Instruct 70.2 26.2 13.0 17.1 16.6 0.04 3.9 24.7 12.9 6.9 5.1 49.5
Qwen3-4B (Thinking) 85.7 31.9 18.7 14.2 8.3 0.02 3.0 23.5 10.7 6.1 0.0 15.0
Qwen3-4B 0.0 0.0 0.0 0.0 4.8 0.04 5.1 21.8 8.0 4.8 0.0 39.3
Qwen3-8B (Thinking) 63.1 25.9 8.3 11.8 10.9 0.04 5.6 21.4 10.8 4.6 6.6 28.0
Qwen3-8B 15.5 7.9 3.5 4.5 6.1 0.14 13.2 18.5 9.3 3.9 1.8 34.6
Qwen3-32B (Thinking) 91.7 40.9 10.2 12.5 20.6 0.04 3.3 20.3 6.9 4.8 0.0 31.6
Qwen3-32B 57.1 26.7 13.2 14.5 10.1 0.09 7.5 23.6 13.8 6.3 9.8 34.0
Qwen3-30B-A3B (Thinking) 76.2 32.6 12.6 18.5 23.3 0.09 6.0 22.5 11.4 5.5 0.0 45.7
Qwen3-30B-A3B 58.3 23.2 12.3 12.7 9.1 0.13 9.1 17.9 11.3 4.5 1.1 48.9
Qwen3-4B-Instruct-2507 82.1 34.0 9.4 9.7 6.0 0.03 5.9 20.9 12.4 4.9 20.4 30.7
Qwen3-4B-Thinking-2507 100.0 34.5 13.8 16.8 17.6 0.01 2.2 24.7 8.4 5.5 20.0 60.0
Qwen3-30B-A3B-Instruct-2507 51.2 29.1 13.9 16.1 8.1 0.06 6.2 17.9 13.3 5.5 8.2 30.0
Qwen3-30B-A3B-Thinking-2507 97.6 45.8 17.2 22.9 27.0 0.02 3.0 25.1 9.3 6.4 0.0 46.0
LLaMA-3.1-8B-Instruct 97.6 22.1 3.0 2.5 4.5 0.02 2.6 21.1 14.8 8.6 46.8 25.5

Evaluation results on RAVine, with a maximum context length of 128k and the index built by BM25:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen2.5-7B-Instruct 7.1 1.4 0.0 0.0 21.2 0.01 1.8 10.5 3.2 2.1 66.7 0.0
Qwen2.5-32B-Instruct 73.8 28.7 16.8 20.4 16.8 0.04 3.9 26.9 15.1 7.5 5.3 43.0
Qwen3-4B (Thinking) 82.1 31.4 14.3 16.4 12.6 0.02 3.4 22.9 9.6 5.4 8.9 19.0
Qwen3-4B 0.0 0.0 0.0 0.0 12.2 0.09 9.3 18.1 8.7 4.1 6.2 26.2
Qwen3-8B (Thinking) 90.5 40.8 14.4 17.8 26.7 0.23 13.4 15.8 12.8 3.8 4.8 32.5
Qwen3-8B 19.0 6.3 5.2 5.4 39.2 1.87 81.0 7.4 12.2 1.4 5.0 38.4
Qwen3-32B (Thinking) 100.0 47.5 10.5 12.3 22.6 0.03 2.8 17.6 9.5 6.0 8.5 19.1
Qwen3-32B 76.2 31.1 13.6 17.6 20.6 0.40 13.3 17.3 15.5 4.8 9.2 29.7
Qwen3-30B-A3B (Thinking) 75.0 35.0 10.8 12.8 58.9 0.81 18.5 18.7 13.5 5.3 3.3 29.6
Qwen3-30B-A3B 50.0 18.5 13.1 14.7 35.3 0.37 12.3 20.5 10.0 4.7 6.5 27.1
Qwen3-4B-Instruct-2507 88.1 36.6 7.0 7.7 13.6 0.05 7.2 20.3 10.8 4.4 21.0 28.6
Qwen3-4B-Thinking-2507 100.0 34.8 17.2 19.7 17.2 0.01 2.1 24.1 8.9 5.6 0.0 0.0
Qwen3-30B-A3B-Instruct-2507 82.1 43.5 19.8 21.2 33.2 1.02 26.3 19.6 15.6 7.1 80.0 8.0
Qwen3-30B-A3B-Thinking-2507 100.0 47.5 22.5 27.7 27.4 0.03 3.2 23.0 8.9 5.9 0.0 51.4
LLaMA-3.1-8B-Instruct 97.6 23.4 4.6 5.0 9.5 0.02 2.6 23.2 15.6 10.2 28.6 28.6

Evaluation results on RAVine, with a maximum context length of 256k and the index built by BM25:

Report Quality Efficiency Search Fetch
Rate (โ†‘) Comp. (โ†‘) Rec. (โ†‘) Prec. (โ†‘) Latency (โ†“) Cost (โ†“) Turns Prec. (โ†‘) Rec. (โ†‘) Gain (โ†‘) URL Err. (โ†“) Prec. (โ†‘)
Qwen3-4B-Instruct-2507 91.7 37.8 5.5 6.7 35.1 0.04 6.7 20.3 13.9 5.1 18.4 34.9
Qwen3-4B-Thinking-2507 100.0 33.2 17.3 20.9 20.5 0.02 2.1 22.3 8.1 5.2 14.3 28.6
Qwen3-30B-A3B-Instruct-2507 81.0 42.0 22.7 22.0 51.5 2.39 38.5 20.0 12.4 5.9 74.6 4.3
Qwen3-30B-A3B-Thinking-2507 100.0 47.0 23.2 29.0 27.5 0.03 3.0 24.4 11.8 6.8 1.6 45.9

Citation

If you find RAVine useful in your research, please cite our paper:

@article{xu2025ravine,
  title={RAVine: Reality-Aligned Evaluation for Agentic Search},
  author={Xu, Yilong and Long, Xiang and Zheng, Zhi and Gao, Jinhua},
  journal={arXiv preprint arXiv:2507.16725},
  year={2025},
  url={https://arxiv.org/abs/2507.16725}
}

Contributors

ylXuuSwordFaith

Issues