Sunt-ing/HyGen

★ 0Forks 0GitHub ↗Compare

README

HyGen

HyGen is an efficient LLM serving system with elastic online-offline request co-location. Further details of HyGen are available in our paper accepted by NeurIPS 2025.

Setup

CUDA

HyGen has been tested using CUDA 11.8 on A100, A40 and A5000 GPUs.

Clone Repository

git clone https://github.com/Marmot-C/HyGen.git

Install HyGen

pip install -e . --extra-index-url https://flashinfer.ai/whl/cu118/torch2.3/

Usage

Basic Usage

Here's an example of running the efficient hybrid serving system provided by HyGen. The traces used in our experiments are all publicly accessible. For traces that does not provide explicit token counts, we've used Llama2 tokenizer to process the trace.

python -m hygen.benchmark.main \
    --output_dir ./benchmark_output \
    --seed 42 \
    --time_limit 18000 \
    --num_replicas 1 \
    --write_json_trace True \
    --worker_config_gpu_memory_utilization 0.95 \
    --model_config_model meta-llama/Llama-2-7b-chat-hf \
    --model_config_max_model_len 4096 \
    --parallel_config_pipeline_parallel_size 1 \
    --parallel_config_tensor_parallel_size 1 \
    --online_request_generator_config_type TRACE \
    --trace_request_generator_config_num_requests 512 \
    --trace_request_generator_config_trace_file ./data/processed_traces/azure_llm_inference.csv \
    --trace_request_length_generator_config_max_tokens 1024 \
    --trace_request_generator_config_qps 1.0 \
    --offline_request_generator_config_type OFFLINE_SYNTHETIC \
    --offline_synthetic_request_generator_config_num_requests 512 \
    --offline_length_generator_config_type OFFLINE_TRACE \
    --offline_trace_request_length_generator_config_trace_file ./data/processed_traces/arxiv_summarization_filtered_stats_llama2_tokenizer.csv \
    --offline_trace_request_length_generator_config_prefill_scale_factor 1 \
    --offline_trace_request_length_generator_config_decode_scale_factor 1 \
    --offline_trace_request_length_generator_config_max_tokens 2048 \
    --offline_interval_generator_config_type STATIC \
    --offline_scheduler_config_type HYGEN \
    --hygen_scheduler_config_max_num_seqs 48 \
    --hygen_scheduler_config_chunk_size 512 \
    --hygen_scheduler_config_p99_tbt $LATENCY_LIMIT \
    --hygen_scheduler_config_a1=$COEFFICIENT_A1 \
    --hygen_scheduler_config_b1=$COEFFICIENT_B1 \
    --hygen_scheduler_config_c1=$COEFFICIENT_C1 \
    --hygen_scheduler_config_a2=$COEFFICIENT_A2 \
    --hygen_scheduler_config_b2=$COEFFICIENT_B2 \
    --hygen_scheduler_config_c2=$COEFFICIENT_C2 \
    --hygen_scheduler_config_d=$COEFFICIENT_D

SLO-aware Profiler

To achieve SLO attainment with HyGen, the latency limit for given SLOs and the coefficients for HyGen's latency predictor can be obtained using the following command.

python hygen.benchmark.latency_search.main \
    --command-path $COMMAND_PATH \
    --output-dir $OUTPUT_DIR \
    --ttft-slo-quantile $TTFT_SLO_QUANTILE \  # TTFT SLO quantile, e.g. 0.99 for P99, 0 for mean
    --ttft-slo-value $TTFT_SLO_VALUE \
    --tbt-slo-quantile $TBT_SLO_QUANTILE \ # TBT SLO quantile
    --tbt-slo-value $TBT_SLO_VALUE

Citation

If you find our work useful, please kindly consider citing our paper:

@misc{sun2025hygenefficientllmserving,
      title={HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location}, 
      author={Ting Sun and Penghan Wang and Fan Lai},
      year={2025},
      eprint={2501.14808},
      archivePrefix={arXiv},
      primaryClass={cs.DC},
      url={https://arxiv.org/abs/2501.14808}, 
}

Acknowledgment

HyGen's codebase is based on Sarathi-Serve and vLLM.

Contributors

Marmot-C

Issues