RAVine (a Reality-Aligned eValuation framework for agentic LLMs with search), is a comprehensive evaluation system for agentic search, encompassing the web environment, benchmark datasets, and a novel evaluation method, serving as a full-process, reproducible, and goal-aligned evaluation sandbox.
- ๐ฏMore Precise: RAVine provides a more precise and attributable method for nuggets (claim-level ground truth) extraction, the final generated report evaluation is more accurate.
- โ๏ธMore Comprehensive: RAVine not only focuses on end-to-end result evaluation, but also designs detailed search process performance indicators.
- ๐ฐLower Cost: RAVine provides local search APIs and reduces the cost of calling LLM-Judge(Gemini-2.5-Flash) to ~0.01$ per evaluation data.
- ๐Full Sandbox: We have packaged the search tools and evaluation framework. You only need to provide an LLM to run or evaluate the agent search on RAVine.
We have uploaded the running evaluation data, nuggets, corpus index, etc. to hugging face. You can access them and modify the corresponding path variables in the running script:
- Queries & Nuggets: https://huggingface.co/datasets/sapphirex/RAVine-nuggets
- Raw Qrels: https://huggingface.co/datasets/sapphirex/RAVine-qrels
- Dense Index: https://huggingface.co/datasets/sapphirex/RAVine-dense-index
- URL-Docid Mapper: https://huggingface.co/datasets/sapphirex/RAVine-mapper
- Running logs (for reproduction): https://huggingface.co/datasets/sapphirex/RAVine-logs
First, install the operating environment. This project mainly uses two environments, and we recommend using uv to install the following two environments separately:
- Env for vllm: the vllm version we use is
0.9.0.1, runpip install vllm==0.9.0.1. - Env for
/src: the operating environment of our main program has been exported torequirements_agent.txt. Runuv pip install -r requirements_agent.txt.- The JDK version we use is
21.0.7, and we recommend that you use this version to support running the BM25 index based on PySerini.
- The JDK version we use is
Second, write the configuration file, which is related to the selection and setting of the model, operating environment, index, and file path. You can find examples at configs/ and write your own config here.
Third, run the vllm service and then run the main program. Example instructions are as follows:
bash scripts/server/vllm.sh your_config_file # run the server of agentic llm
bash scripts/evaluation/run.sh your_config_file
If you need to run multiple configurations at once, we provide the corresponding scripts. Please write the configurations to be run into /scripts/evaluation/run_all_eval.sh and then run it.
For more detailed steps, see Evaluation_and_reproduction.
Table Description:
- "Rate" denotes the Task Completion Rate.
- "Comp." refers to the score of Task Completeness.
- "Rec." and "Prec." represent Recall and Precision, respectively.
- "URL Err." denotes the URL Error.
- Latency is measured in seconds, and cost is measured in dollars.
- Symbols (โ) and (โ) indicate that higher or lower values are preferred, respectively.
- Models not marked with (Thinking) either run without thinking or lack support for the thinking mode.
- Bold values indicate the best performance for each corresponding metric in the column.
Evaluation results on RAVine, with a maximum context length of 32k and the index built by gte-modernbert-base:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen2.5-7B-Instruct | 19.0 | 6.8 | 1.9 | 1.8 | 7.1 | 0.01 | 3.1 | 18.7 | 5.7 | 4.7 | 8.8 | 19.8 |
| Qwen2.5-32B-Instruct | 71.4 | 23.0 | 14.9 | 16.5 | 40.3 | 0.03 | 4.0 | 21.1 | 6.5 | 4.4 | 1.4 | 28.7 |
| Qwen3-4B (Thinking) | 92.9 | 35.4 | 14.2 | 11.6 | 10.2 | 0.01 | 2.3 | 20.9 | 6.4 | 5.6 | 0.0 | 16.7 |
| Qwen3-4B | 0.0 | 0.0 | 0.0 | 0.0 | 10.1 | 0.04 | 6.2 | 21.3 | 6.0 | 4.7 | 1.3 | 28.3 |
| Qwen3-8B (Thinking) | 86.9 | 37.8 | 10.4 | 12.1 | 13.9 | 0.03 | 6.6 | 19.7 | 6.6 | 5.1 | 8.9 | 27.3 |
| Qwen3-8B | 28.6 | 12.4 | 4.8 | 6.1 | 11.2 | 0.06 | 9.3 | 19.3 | 5.9 | 5.0 | 2.4 | 23.8 |
| Qwen3-32B (Thinking) | 98.8 | 43.5 | 11.7 | 15.1 | 19.6 | 0.02 | 2.8 | 19.2 | 5.0 | 4.0 | 8.9 | 22.2 |
| Qwen3-32B | 85.7 | 38.0 | 12.8 | 12.6 | 14.6 | 0.08 | 8.5 | 19.1 | 6.3 | 5.0 | 8.1 | 20.2 |
| Qwen3-30B-A3B (Thinking) | 81.0 | 35.6 | 10.6 | 14.2 | 33.0 | 0.10 | 6.6 | 19.7 | 6.2 | 3.6 | 10.3 | 29.3 |
| Qwen3-30B-A3B | 77.4 | 30.9 | 11.3 | 14.2 | 15.7 | 0.07 | 7.3 | 16.8 | 6.2 | 3.4 | 0.6 | 30.4 |
| Qwen3-4B-Instruct-2507 | 95.2 | 39.4 | 7.6 | 7.0 | 8.3 | 0.02 | 5.9 | 17.5 | 6.5 | 3.8 | 13.8 | 27.6 |
| Qwen3-4B-Thinking-2507 | 98.8 | 36.6 | 12.1 | 15.4 | 23.0 | 0.02 | 2.1 | 17.8 | 5.5 | 4.4 | 0.0 | 40.0 |
| Qwen3-30B-A3B-Instruct-2507 | 83.3 | 43.1 | 17.8 | 17.5 | 15.8 | 0.08 | 8.6 | 17.6 | 7.3 | 4.5 | 15.6 | 21.2 |
| Qwen3-30B-A3B-Thinking-2507 | 96.4 | 44.1 | 15.6 | 19.0 | 29.7 | 0.02 | 3.5 | 18.1 | 4.2 | 3.5 | 0.0 | 31.5 |
| LLaMA-3.1-8B-Instruct | 96.4 | 24.0 | 3.1 | 3.1 | 7.3 | 0.02 | 2.7 | 12.1 | 8.8 | 6.6 | 36.8 | 15.8 |
Evaluation results on RAVine, with a maximum context length of 128k and the index built by gte-modernbert-base:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen2.5-7B-Instruct | 1.2 | 0.3 | 0.0 | 0.0 | 4.5 | 0.01 | 1.6 | 7.4 | 2.6 | 2.3 | 0.0 | 33.3 |
| Qwen2.5-32B-Instruct | 61.9 | 24.5 | 9.7 | 11.8 | 17.9 | 0.03 | 3.8 | 19.7 | 6.0 | 4.3 | 3.6 | 28.3 |
| Qwen3-4B (Thinking) | 91.7 | 37.3 | 8.2 | 10.0 | 12.8 | 0.02 | 3.0 | 19.4 | 5.8 | 4.0 | 1.8 | 25.5 |
| Qwen3-4B | 0.0 | 0.0 | 0.0 | 0.0 | 15.8 | 0.09 | 12.4 | 13.8 | 5.8 | 2.9 | 6.7 | 15.2 |
| Qwen3-8B (Thinking) | 91.7 | 41.9 | 8.3 | 10.3 | 65.8 | 0.40 | 23.3 | 12.9 | 6.3 | 2.2 | 5.6 | 25.1 |
| Qwen3-8B | 26.2 | 10.1 | 2.6 | 3.8 | 139.5 | 2.52 | 113.3 | 8.0 | 6.3 | 2.3 | 4.1 | 23.8 |
| Qwen3-32B (Thinking) | 100.0 | 45.2 | 8.4 | 9.6 | 23.0 | 0.02 | 2.7 | 18.6 | 5.2 | 4.1 | 0.0 | 8.6 |
| Qwen3-32B | 82.1 | 35.0 | 13.2 | 11.9 | 22.7 | 0.42 | 14.8 | 15.0 | 7.0 | 3.2 | 6.5 | 20.3 |
| Qwen3-30B-A3B (Thinking) | 81.0 | 36.8 | 10.9 | 12.0 | 46.9 | 0.43 | 12.1 | 19.0 | 6.3 | 4.7 | 3.8 | 12.1 |
| Qwen3-30B-A3B | 46.4 | 16.9 | 6.1 | 6.7 | 54.2 | 0.64 | 16.7 | 18.5 | 6.4 | 4.8 | 1.9 | 23.1 |
| Qwen3-4B-Instruct-2507 | 96.4 | 39.7 | 5.8 | 4.4 | 7.8 | 0.02 | 5.9 | 18.3 | 6.6 | 4.0 | 10.1 | 31.0 |
| Qwen3-4B-Thinking-2507 | 97.6 | 35.0 | 10.4 | 13.3 | 22.6 | 0.01 | 2.1 | 19.6 | 5.9 | 4.9 | 0.0 | 25.0 |
| Qwen3-30B-A3B-Instruct-2507 | 92.9 | 49.3 | 18.0 | 18.5 | 21.0 | 0.32 | 12.6 | 17.8 | 7.4 | 4.8 | 6.1 | 19.2 |
| Qwen3-30B-A3B-Thinking-2507 | 100.0 | 46.5 | 15.9 | 20.1 | 31.7 | 0.02 | 3.4 | 19.9 | 4.6 | 3.9 | 0.0 | 39.6 |
| LLaMA-3.1-8B-Instruct | 98.8 | 25.8 | 2.8 | 3.3 | 4.4 | 0.02 | 2.7 | 12.3 | 8.0 | 6.3 | 58.2 | 10.9 |
Evaluation results on RAVine, with a maximum context length of 256k and the index built by gte-modernbert-base:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen3-4B-Instruct-2507 | 97.6 | 41.3 | 6.7 | 7.6 | 7.6 | 0.02 | 6.0 | 17.3 | 6.7 | 3.9 | 17.1 | 28.5 |
| Qwen3-4B-Thinking-2507 | 100.0 | 34.1 | 15.6 | 16.5 | 20.5 | 0.02 | 2.1 | 17.8 | 5.6 | 4.6 | 0.0 | 0.0 |
| Qwen3-30B-A3B-Instruct-2507 | 91.7 | 46.8 | 18.9 | 18.1 | 28.8 | 0.74 | 14.8 | 18.5 | 7.8 | 5.4 | 4.5 | 17.1 |
| Qwen3-30B-A3B-Thinking-2507 | 100.0 | 46.2 | 14.6 | 17.5 | 29.4 | 0.02 | 3.3 | 19.6 | 4.8 | 4.1 | 0.0 | 33.7 |
Evaluation results on RAVine, with a maximum context length of 32k and the index built by BM25:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen2.5-7B-Instruct | 22.6 | 6.4 | 4.4 | 6.3 | 6.9 | 0.02 | 2.9 | 24.6 | 12.3 | 6.9 | 27.1 | 29.2 |
| Qwen2.5-32B-Instruct | 70.2 | 26.2 | 13.0 | 17.1 | 16.6 | 0.04 | 3.9 | 24.7 | 12.9 | 6.9 | 5.1 | 49.5 |
| Qwen3-4B (Thinking) | 85.7 | 31.9 | 18.7 | 14.2 | 8.3 | 0.02 | 3.0 | 23.5 | 10.7 | 6.1 | 0.0 | 15.0 |
| Qwen3-4B | 0.0 | 0.0 | 0.0 | 0.0 | 4.8 | 0.04 | 5.1 | 21.8 | 8.0 | 4.8 | 0.0 | 39.3 |
| Qwen3-8B (Thinking) | 63.1 | 25.9 | 8.3 | 11.8 | 10.9 | 0.04 | 5.6 | 21.4 | 10.8 | 4.6 | 6.6 | 28.0 |
| Qwen3-8B | 15.5 | 7.9 | 3.5 | 4.5 | 6.1 | 0.14 | 13.2 | 18.5 | 9.3 | 3.9 | 1.8 | 34.6 |
| Qwen3-32B (Thinking) | 91.7 | 40.9 | 10.2 | 12.5 | 20.6 | 0.04 | 3.3 | 20.3 | 6.9 | 4.8 | 0.0 | 31.6 |
| Qwen3-32B | 57.1 | 26.7 | 13.2 | 14.5 | 10.1 | 0.09 | 7.5 | 23.6 | 13.8 | 6.3 | 9.8 | 34.0 |
| Qwen3-30B-A3B (Thinking) | 76.2 | 32.6 | 12.6 | 18.5 | 23.3 | 0.09 | 6.0 | 22.5 | 11.4 | 5.5 | 0.0 | 45.7 |
| Qwen3-30B-A3B | 58.3 | 23.2 | 12.3 | 12.7 | 9.1 | 0.13 | 9.1 | 17.9 | 11.3 | 4.5 | 1.1 | 48.9 |
| Qwen3-4B-Instruct-2507 | 82.1 | 34.0 | 9.4 | 9.7 | 6.0 | 0.03 | 5.9 | 20.9 | 12.4 | 4.9 | 20.4 | 30.7 |
| Qwen3-4B-Thinking-2507 | 100.0 | 34.5 | 13.8 | 16.8 | 17.6 | 0.01 | 2.2 | 24.7 | 8.4 | 5.5 | 20.0 | 60.0 |
| Qwen3-30B-A3B-Instruct-2507 | 51.2 | 29.1 | 13.9 | 16.1 | 8.1 | 0.06 | 6.2 | 17.9 | 13.3 | 5.5 | 8.2 | 30.0 |
| Qwen3-30B-A3B-Thinking-2507 | 97.6 | 45.8 | 17.2 | 22.9 | 27.0 | 0.02 | 3.0 | 25.1 | 9.3 | 6.4 | 0.0 | 46.0 |
| LLaMA-3.1-8B-Instruct | 97.6 | 22.1 | 3.0 | 2.5 | 4.5 | 0.02 | 2.6 | 21.1 | 14.8 | 8.6 | 46.8 | 25.5 |
Evaluation results on RAVine, with a maximum context length of 128k and the index built by BM25:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen2.5-7B-Instruct | 7.1 | 1.4 | 0.0 | 0.0 | 21.2 | 0.01 | 1.8 | 10.5 | 3.2 | 2.1 | 66.7 | 0.0 |
| Qwen2.5-32B-Instruct | 73.8 | 28.7 | 16.8 | 20.4 | 16.8 | 0.04 | 3.9 | 26.9 | 15.1 | 7.5 | 5.3 | 43.0 |
| Qwen3-4B (Thinking) | 82.1 | 31.4 | 14.3 | 16.4 | 12.6 | 0.02 | 3.4 | 22.9 | 9.6 | 5.4 | 8.9 | 19.0 |
| Qwen3-4B | 0.0 | 0.0 | 0.0 | 0.0 | 12.2 | 0.09 | 9.3 | 18.1 | 8.7 | 4.1 | 6.2 | 26.2 |
| Qwen3-8B (Thinking) | 90.5 | 40.8 | 14.4 | 17.8 | 26.7 | 0.23 | 13.4 | 15.8 | 12.8 | 3.8 | 4.8 | 32.5 |
| Qwen3-8B | 19.0 | 6.3 | 5.2 | 5.4 | 39.2 | 1.87 | 81.0 | 7.4 | 12.2 | 1.4 | 5.0 | 38.4 |
| Qwen3-32B (Thinking) | 100.0 | 47.5 | 10.5 | 12.3 | 22.6 | 0.03 | 2.8 | 17.6 | 9.5 | 6.0 | 8.5 | 19.1 |
| Qwen3-32B | 76.2 | 31.1 | 13.6 | 17.6 | 20.6 | 0.40 | 13.3 | 17.3 | 15.5 | 4.8 | 9.2 | 29.7 |
| Qwen3-30B-A3B (Thinking) | 75.0 | 35.0 | 10.8 | 12.8 | 58.9 | 0.81 | 18.5 | 18.7 | 13.5 | 5.3 | 3.3 | 29.6 |
| Qwen3-30B-A3B | 50.0 | 18.5 | 13.1 | 14.7 | 35.3 | 0.37 | 12.3 | 20.5 | 10.0 | 4.7 | 6.5 | 27.1 |
| Qwen3-4B-Instruct-2507 | 88.1 | 36.6 | 7.0 | 7.7 | 13.6 | 0.05 | 7.2 | 20.3 | 10.8 | 4.4 | 21.0 | 28.6 |
| Qwen3-4B-Thinking-2507 | 100.0 | 34.8 | 17.2 | 19.7 | 17.2 | 0.01 | 2.1 | 24.1 | 8.9 | 5.6 | 0.0 | 0.0 |
| Qwen3-30B-A3B-Instruct-2507 | 82.1 | 43.5 | 19.8 | 21.2 | 33.2 | 1.02 | 26.3 | 19.6 | 15.6 | 7.1 | 80.0 | 8.0 |
| Qwen3-30B-A3B-Thinking-2507 | 100.0 | 47.5 | 22.5 | 27.7 | 27.4 | 0.03 | 3.2 | 23.0 | 8.9 | 5.9 | 0.0 | 51.4 |
| LLaMA-3.1-8B-Instruct | 97.6 | 23.4 | 4.6 | 5.0 | 9.5 | 0.02 | 2.6 | 23.2 | 15.6 | 10.2 | 28.6 | 28.6 |
Evaluation results on RAVine, with a maximum context length of 256k and the index built by BM25:
| Report Quality | Efficiency | Search | Fetch | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rate (โ) | Comp. (โ) | Rec. (โ) | Prec. (โ) | Latency (โ) | Cost (โ) | Turns | Prec. (โ) | Rec. (โ) | Gain (โ) | URL Err. (โ) | Prec. (โ) | |
| Qwen3-4B-Instruct-2507 | 91.7 | 37.8 | 5.5 | 6.7 | 35.1 | 0.04 | 6.7 | 20.3 | 13.9 | 5.1 | 18.4 | 34.9 |
| Qwen3-4B-Thinking-2507 | 100.0 | 33.2 | 17.3 | 20.9 | 20.5 | 0.02 | 2.1 | 22.3 | 8.1 | 5.2 | 14.3 | 28.6 |
| Qwen3-30B-A3B-Instruct-2507 | 81.0 | 42.0 | 22.7 | 22.0 | 51.5 | 2.39 | 38.5 | 20.0 | 12.4 | 5.9 | 74.6 | 4.3 |
| Qwen3-30B-A3B-Thinking-2507 | 100.0 | 47.0 | 23.2 | 29.0 | 27.5 | 0.03 | 3.0 | 24.4 | 11.8 | 6.8 | 1.6 | 45.9 |
If you find RAVine useful in your research, please cite our paper:
@article{xu2025ravine,
title={RAVine: Reality-Aligned Evaluation for Agentic Search},
author={Xu, Yilong and Long, Xiang and Zheng, Zhi and Gao, Jinhua},
journal={arXiv preprint arXiv:2507.16725},
year={2025},
url={https://arxiv.org/abs/2507.16725}
}