HyGen is an efficient LLM serving system with elastic online-offline request co-location. Further details of HyGen are available in our paper accepted by NeurIPS 2025.
HyGen has been tested using CUDA 11.8 on A100, A40 and A5000 GPUs.
git clone https://github.com/Marmot-C/HyGen.gitpip install -e . --extra-index-url https://flashinfer.ai/whl/cu118/torch2.3/Here's an example of running the efficient hybrid serving system provided by HyGen. The traces used in our experiments are all publicly accessible. For traces that does not provide explicit token counts, we've used Llama2 tokenizer to process the trace.
python -m hygen.benchmark.main \
--output_dir ./benchmark_output \
--seed 42 \
--time_limit 18000 \
--num_replicas 1 \
--write_json_trace True \
--worker_config_gpu_memory_utilization 0.95 \
--model_config_model meta-llama/Llama-2-7b-chat-hf \
--model_config_max_model_len 4096 \
--parallel_config_pipeline_parallel_size 1 \
--parallel_config_tensor_parallel_size 1 \
--online_request_generator_config_type TRACE \
--trace_request_generator_config_num_requests 512 \
--trace_request_generator_config_trace_file ./data/processed_traces/azure_llm_inference.csv \
--trace_request_length_generator_config_max_tokens 1024 \
--trace_request_generator_config_qps 1.0 \
--offline_request_generator_config_type OFFLINE_SYNTHETIC \
--offline_synthetic_request_generator_config_num_requests 512 \
--offline_length_generator_config_type OFFLINE_TRACE \
--offline_trace_request_length_generator_config_trace_file ./data/processed_traces/arxiv_summarization_filtered_stats_llama2_tokenizer.csv \
--offline_trace_request_length_generator_config_prefill_scale_factor 1 \
--offline_trace_request_length_generator_config_decode_scale_factor 1 \
--offline_trace_request_length_generator_config_max_tokens 2048 \
--offline_interval_generator_config_type STATIC \
--offline_scheduler_config_type HYGEN \
--hygen_scheduler_config_max_num_seqs 48 \
--hygen_scheduler_config_chunk_size 512 \
--hygen_scheduler_config_p99_tbt $LATENCY_LIMIT \
--hygen_scheduler_config_a1=$COEFFICIENT_A1 \
--hygen_scheduler_config_b1=$COEFFICIENT_B1 \
--hygen_scheduler_config_c1=$COEFFICIENT_C1 \
--hygen_scheduler_config_a2=$COEFFICIENT_A2 \
--hygen_scheduler_config_b2=$COEFFICIENT_B2 \
--hygen_scheduler_config_c2=$COEFFICIENT_C2 \
--hygen_scheduler_config_d=$COEFFICIENT_DTo achieve SLO attainment with HyGen, the latency limit for given SLOs and the coefficients for HyGen's latency predictor can be obtained using the following command.
python hygen.benchmark.latency_search.main \
--command-path $COMMAND_PATH \
--output-dir $OUTPUT_DIR \
--ttft-slo-quantile $TTFT_SLO_QUANTILE \ # TTFT SLO quantile, e.g. 0.99 for P99, 0 for mean
--ttft-slo-value $TTFT_SLO_VALUE \
--tbt-slo-quantile $TBT_SLO_QUANTILE \ # TBT SLO quantile
--tbt-slo-value $TBT_SLO_VALUEIf you find our work useful, please kindly consider citing our paper:
@misc{sun2025hygenefficientllmserving,
title={HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location},
author={Ting Sun and Penghan Wang and Fan Lai},
year={2025},
eprint={2501.14808},
archivePrefix={arXiv},
primaryClass={cs.DC},
url={https://arxiv.org/abs/2501.14808},
}
HyGen's codebase is based on Sarathi-Serve and vLLM.