Whisper.jl is a Julia package for automatic speech recognition, based on OpenAI's Whisper model. It wraps whisper.cpp, a C/C++ implementation of the model, through whisper_cpp_jll. Model weights are downloaded on first use.
transcribe takes a model name and a vector of audio samples at 16 kHz, mono, with values in [-1, 1], and returns the transcript:
using Whisper
result = transcribe("base.en", audio) # audio::Vector{Float32}, 16 kHz monoFor more than one call, load the model once and reuse it — loading dominates for short clips:
ctx = WhisperContext("large-v3-turbo") # downloads ~1.5 GB the first time
result = transcribe(ctx, audio)
for s in segments(ctx) # timestamps of the last transcription
println(s.t0, " – ", s.t1, ": ", s.text) # seconds
end
close(ctx) # or leave it to the GCOptions cover the common whisper.cpp settings:
transcribe(ctx, audio;
language = "auto", # or an ISO 639-1 code; multilingual models only
translate = false, # translate to English (multilingual models)
sampling = :beam, # :greedy (default) or :beam
beam_size = 5,
n_threads = 8,
initial_prompt = "Names: Rackauckas, ModelingToolkit", # primes the vocabulary
)
Whisper.detected_language(ctx) # after language = "auto"Whisper.jl takes samples, not files. Any loader works; this resamples to 16 kHz mono with LibSndFile.jl and SampledSignals.jl:
using FileIO, LibSndFile, SampledSignals
s = load("speech.ogg")
n = round(Int, length(s) * (16000 / samplerate(s)))
out = SampleBuf(zeros(Float32, n, nchannels(s)), 16000) # zeroed: the resampler may write < n frames
written = write(SampleBufSink(out), SampleBufSource(s)) # resample
data = out.data[1:min(written, n), :]
audio = nchannels(s) == 1 ? vec(data) : vec(sum(data, dims = 2)) ./ nchannels(s)For a 16 kHz mono WAV, WAV.wavread gives usable samples directly.
Use the name; the weights are fetched from HuggingFace on first use and cached by DataDeps. Models with an .en suffix are English-only and slightly more accurate for English; the others are multilingual.
| Model | Download | Notes |
|---|---|---|
tiny, tiny.en |
74 MB | fastest, lowest quality |
base, base.en |
141 MB | |
small, small.en |
465 MB | |
medium, medium.en |
1.4 GB | |
large-v1, large-v2 |
2.9 GB | |
large-v3 |
2.9 GB | best quality |
large-v3-turbo |
1.5 GB | about 8× faster than large-v3, similar quality |
available_models() lists them. A path to a ggml .bin file is accepted in place of a name.
whisper_cpp_jll ships CUDA builds for Linux (x86_64 and aarch64, including Jetson) and Metal for Apple Silicon. The right build is selected automatically at install time when CUDA_Runtime_jll (installed by CUDA.jl) is present, and WhisperContext uses the GPU by default when one is available — nothing to configure. On a CPU-only build use_gpu = true silently falls back to the CPU.
To see which backend is in use:
Whisper.log_level!(:info) # whisper.cpp logs backend and model details on load
ctx = WhisperContext("base.en")WhisperContext(...; use_gpu = false) forces the CPU; gpu_device = n selects a GPU.
To use your own whisper.cpp build (a different CUDA version, ROCm, Vulkan, ...), point Julia at it with an Overrides.toml for the whisper_cpp_jll artifact. The override must be a full cmake --install prefix of the same whisper.cpp version, since the bindings are generated from its headers.
The whole whisper.cpp C API is available as Whisper.LibWhisper, generated by Clang.jl from whisper.h (gen/generator.jl). whisper_full_params is exposed as an opaque struct with generated pointer accessors; see transcribe in src/Whisper.jl for how to fill it.
- Streaming / real-time transcription (whisper.cpp's
streamexample is a separate program, not a library API). - Word-level timestamps (
token_timestamps), speaker diarization (tdrz), and grammar-constrained decoding are reachable throughLibWhisperbut have no high-level API.