Shared CI for Julia package benchmarks. A reusable workflow runs a repo's BenchmarkTools suite on every push and pull request, compares a PR against its base on the same runner, and publishes a trend dashboard to GitHub Pages.
The workflow and the dashboard live here once. Each consuming repo keeps its own benchmark scripts and suite, which it also runs locally.
- Push to the default branch: runs the suite, appends the result to the
gh-pagesbranch underbenchmarks/results/, rebuilds the index, and publishes the dashboard. - Pull request: runs the suite on the PR head and on the base commit (same runner, to neutralise hardware variance), writes a Markdown comparison to the job summary, and posts a single link comment on the PR.
- Dashboard: a static page at
https://<owner>.github.io/<repo>/benchmarks/that fetches the stored JSON and charts each benchmark over time. It reads its repo slug and title from aconfig.jsonthe publish step writes, so the page carries no per-repo values. When the suite includes acalibrationbenchmark, the time charts scale each run by it to cancel runner-speed drift between days.
Add .github/workflows/benchmark.yml to your repo:
name: Benchmark
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
jobs:
benchmark:
uses: MichaelHatherly/julia-benchmark-action/.github/workflows/benchmark.yml@v1
permissions:
contents: write
pull-requests: writeThe dashboard publishes to GitHub Pages from the gh-pages branch. Enable Pages
for that branch once in repository settings.
A benchmark project at benchmark/ (override the path with the project input):
benchmark/Project.toml— depends onBenchmarkTools,JSON, and your package.benchmark/benchmarks.jl— definesconst SUITE = BenchmarkTools.BenchmarkGroup()and populates it.benchmark/run.jl— runsSUITEand writes the result JSON. Invoked both in CI and locally (julia --project=benchmark benchmark/run.jl out.json).benchmark/compare.jl— definescompare_and_report(baseline_json, current_json, output_md).
The CI runs these scripts directly; it does not supply them. Local development uses the same scripts with no extra setup.
run.jl takes an optional output path as its first argument and writes a JSON
document with this shape (the dashboard and compare.jl read it):
{
"timestamp": "2026-06-30T10:00:00Z",
"git": { "commit": "...", "branch": "...", "dirty": false },
"benchmarks": {
"parse/julia": {
"time_ns": { "median": 1234.0, "minimum": 1000.0, "mean": 1300.0 },
"memory_bytes": 4096,
"allocations": 12
}
}
}Nested BenchmarkGroups flatten to slash-joined keys (parse/julia), which become
the dashboard's benchmark names.
A shared runner has no performance isolation, so the same code can clock 20% slower
on a busy day. Add a calibration benchmark to SUITE: a fixed, allocation-free
compute kernel of a pinned size that does identical work every run. Its wall-clock is
a direct read of runner speed.
The dashboard divides it out of the time charts, and compare.jl can normalize a PR's
base-vs-head times against it. The verdict still rests on allocations and memory, which
are deterministic; calibration only de-noises the time signal. Pin the kernel's size
forever, since changing it rebases every stored result against a different clock. The
key is excluded from the benchmark picker.
| Input | Default | Description |
|---|---|---|
julia-version |
'1' |
Julia version to run under. |
project |
benchmark |
Path to the benchmark project. |
develop-path |
'.' |
Path to the package developed into the project. |
timeout-minutes |
50 |
Job timeout. |