aviks/Whisper.jl

Implementation of OpenAI Whisper model based on whisper.cpp

★ 52Forks 6JuliaGitHub ↗Compare

README

Whisper

Build Status

Whisper.jl is a Julia package for automatic speech recognition, based on OpenAI's Whisper model. It wraps whisper.cpp, a C/C++ implementation of the model, through whisper_cpp_jll. Model weights are downloaded on first use.

Quick start

transcribe takes a model name and a vector of audio samples at 16 kHz, mono, with values in [-1, 1], and returns the transcript:

using Whisper

result = transcribe("base.en", audio)   # audio::Vector{Float32}, 16 kHz mono

For more than one call, load the model once and reuse it — loading dominates for short clips:

ctx = WhisperContext("large-v3-turbo")          # downloads ~1.5 GB the first time
result = transcribe(ctx, audio)
for s in segments(ctx)                          # timestamps of the last transcription
    println(s.t0, " – ", s.t1, ": ", s.text)    # seconds
end
close(ctx)                                      # or leave it to the GC

Options cover the common whisper.cpp settings:

transcribe(ctx, audio;
    language = "auto",        # or an ISO 639-1 code; multilingual models only
    translate = false,        # translate to English (multilingual models)
    sampling = :beam,         # :greedy (default) or :beam
    beam_size = 5,
    n_threads = 8,
    initial_prompt = "Names: Rackauckas, ModelingToolkit",   # primes the vocabulary
)
Whisper.detected_language(ctx)                  # after language = "auto"

Loading audio

Whisper.jl takes samples, not files. Any loader works; this resamples to 16 kHz mono with LibSndFile.jl and SampledSignals.jl:

using FileIO, LibSndFile, SampledSignals

s = load("speech.ogg")
n = round(Int, length(s) * (16000 / samplerate(s)))
out = SampleBuf(zeros(Float32, n, nchannels(s)), 16000)          # zeroed: the resampler may write < n frames
written = write(SampleBufSink(out), SampleBufSource(s))          # resample
data = out.data[1:min(written, n), :]
audio = nchannels(s) == 1 ? vec(data) : vec(sum(data, dims = 2)) ./ nchannels(s)

For a 16 kHz mono WAV, WAV.wavread gives usable samples directly.

Models

Use the name; the weights are fetched from HuggingFace on first use and cached by DataDeps. Models with an .en suffix are English-only and slightly more accurate for English; the others are multilingual.

Model Download Notes
tiny, tiny.en 74 MB fastest, lowest quality
base, base.en 141 MB
small, small.en 465 MB
medium, medium.en 1.4 GB
large-v1, large-v2 2.9 GB
large-v3 2.9 GB best quality
large-v3-turbo 1.5 GB about 8× faster than large-v3, similar quality

available_models() lists them. A path to a ggml .bin file is accepted in place of a name.

GPU

whisper_cpp_jll ships CUDA builds for Linux (x86_64 and aarch64, including Jetson) and Metal for Apple Silicon. The right build is selected automatically at install time when CUDA_Runtime_jll (installed by CUDA.jl) is present, and WhisperContext uses the GPU by default when one is available — nothing to configure. On a CPU-only build use_gpu = true silently falls back to the CPU.

To see which backend is in use:

Whisper.log_level!(:info)        # whisper.cpp logs backend and model details on load
ctx = WhisperContext("base.en")

WhisperContext(...; use_gpu = false) forces the CPU; gpu_device = n selects a GPU.

To use your own whisper.cpp build (a different CUDA version, ROCm, Vulkan, ...), point Julia at it with an Overrides.toml for the whisper_cpp_jll artifact. The override must be a full cmake --install prefix of the same whisper.cpp version, since the bindings are generated from its headers.

Lower level

The whole whisper.cpp C API is available as Whisper.LibWhisper, generated by Clang.jl from whisper.h (gen/generator.jl). whisper_full_params is exposed as an opaque struct with generated pointer accessors; see transcribe in src/Whisper.jl for how to fill it.

Not yet supported

  • Streaming / real-time transcription (whisper.cpp's stream example is a separate program, not a library API).
  • Word-level timestamps (token_timestamps), speaker diarization (tdrz), and grammar-constrained decoding are reachable through LibWhisper but have no high-level API.

Contributors

aviksChrisRackauckasjpsamarooclaude

Issues