Flexible GPU fractionalization — run more workloads per GPU. Transparently intercepts HIP APIs to enforce per-process VRAM limits and compute (CU) restrictions — no kernel driver modifications, no daemon, no config files.
Download a release (no build required):
wget -q https://github.com/saienduri/hipflex/releases/download/v0.1.0/libhipflex.soOr build from source:
cargo build --release -p hipflex
# Output: target/release/libhipflex.so# Limit each GPU to 4 GiB
FH_MEMORY_LIMIT=4GiB LD_PRELOAD=./libhipflex.so python3 train.py
# Restrict to 38 of 304 CUs on MI325X (no memory limit)
FH_CU_RANGE=0-37 LD_PRELOAD=./libhipflex.so python3 train.py
# Limit memory AND restrict CUs
FH_MEMORY_LIMIT=24GiB FH_CU_RANGE=0-37 LD_PRELOAD=./libhipflex.so python3 train.pyAccepts bytes (137438953472), SI (128G, 512MB), binary (128GiB, 512MiB), or fractional (1.5G).
See the demo guide for full working examples with PyTorch, vLLM, and SGLang on MI325X.
┌──────────────────────────────────────────────────────────────────┐
│ Application (PyTorch, JAX, etc.) │
│ hipMalloc(&ptr, size) hipMemGetInfo() rocm-smi query │
└──────────┬──────────────────────┬──────────────────┬────────────┘
│ │ │
│ PLT / dlsym │ PLT / dlsym │ dlsym
▼ ▼ ▼
┌──────────────────────────────────────────────────────────────────┐
│ libhipflex.so │
│ │
│ Three interception paths: │
│ 1. LD_PRELOAD exports (27 symbols) — PLT callers (PyTorch) │
│ 2. dlsym override — dlopen/dlsym callers (SMI tools) │
│ 3. Frida GUM inline hooks — internal libamdhip64 calls │
│ │
│ ┌─────────────────┐ ┌──────────────┐ ┌────────────────────┐ │
│ │ Alloc/Free Hooks │ │ Info Spoofing │ │ SMI Spoofing │ │
│ │ 15 alloc + 7 free│ │ hipMemGetInfo │ │ rocm-smi, amd-smi │ │
│ │ │ │ hipDevTotal │ │ (dlsym override) │ │
│ └────────┬─────────┘ │ hipGetDevProp │ └────────────────────┘ │
│ │ └──────────────┘ │
│ ▼ │
│ ┌────────────────────────────────────────────┐ │
│ │ Limiter (reserve-then-allocate) │ │
│ │ │ │
│ │ 1. fetch_add(size) on SHM counter │ │
│ │ 2. over limit? → fetch_sub, deny │ │
│ │ 3. call real HIP API │ │
│ │ 4. native fail? → fetch_sub, rollback │ │
│ │ 5. track (pointer → device, size) │ │
│ └────────┬───────────────────────────────────┘ │
└──────────┬──────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────────┐
│ Shared Memory (/dev/shm/hipflex) │
│ │
│ per-device: pod_memory_used (atomic u64), mem_limit, UUID │
│ per-process: PID slot table with non-HIP overhead per device │
└──────────────────────────────────────────────────────────────────┘
Non-HIP memory (code objects, page tables, scratch buffers) is tracked via KFD sysfs and reconciled into the effective limit. Crashed or killed processes are automatically reaped so stale overhead doesn't reduce capacity. Each GPU gets independent accounting — FH_MEMORY_LIMIT applies per device. Optional compute restriction via FH_CU_RANGE sets HSA_CU_MASK before the HIP runtime starts, hardware-limiting which Compute Units a process can use. Structured logging via tracing covers VRAM usage, overhead, and slot state with configurable levels and file rotation.
All configuration is via environment variables.
| Variable | Description |
|---|---|
FH_MEMORY_LIMIT |
Per-GPU memory limit. Accepts bytes, SI, binary, or fractional sizes. |
FH_CU_RANGE |
Restrict GPU Compute Units, applied uniformly to all visible GPUs. Format: start-end (inclusive, e.g., 0-37 for 38 CUs). Sets HSA_CU_MASK and spoofs multiProcessorCount. Works independently or with FH_MEMORY_LIMIT. |
FH_ENABLE_HOOKS |
Set to false to disable all hooking (library becomes a no-op). Default: true. |
FH_HIP_LIB_PATH |
Override path to libamdhip64.so. |
FH_SHM_PATH |
Override shared memory directory. Default: /dev/shm/hipflex. |
| Variable | Description |
|---|---|
FH_ENABLE_LOG |
Set to off, 0, or false to disable logging. |
FH_LOG_PATH |
Log destination. stderr for stderr, or a file/directory path. Default: /tmp/hipflex/hipflex.log.* (daily rotation). |
FH_LOG_LEVEL |
Tracing filter directive (e.g. debug, hipflex=trace). Default: info. |
crates/
hipflex/ # Main cdylib — hooks, limiter, reconciliation, reaping, device mapping
hipflex-internal/ # Shared internals — SHM types, proc slots, logging, hook manager
hipflex-macro/ # Proc macro for hook function generation
hipflex-fuzz/ # Property-based fuzzer (proptest) + simulated limiter
tests/
gpu-tests/ # Python GPU conformance tests (requires AMD GPU + Docker)
scripts/
run-gpu-tests.sh # Test runner (fuzzer + GPU tests)
cargo test -p hipflex -p hipflex-fuzz -p hipflex-internalRuns unit tests, property-based fuzz tests (accounting invariants, concurrent stress, edge cases, multi-process lifecycle), and SHM compatibility checks.
bash scripts/run-gpu-tests.shSee scripts/run-gpu-tests.sh for setup instructions and options (--fuzz-only, --gpu-only, --full).
- Demo — working examples with PyTorch, vLLM, and SGLang, plus accounting validation
- Design — architecture, operating modes, atomics, SHM layout, failure modes
- Hook coverage — full inventory of hooked HIP/SMI APIs, exclusions, and known gaps