brucek/soundcheck

macOS CLI for comparing the technical quality of audio files

★ 0Forks 0PythonGitHub ↗Compare

README

soundcheck

soundcheck is a small macOS command-line tool for comparing the technical properties of two or more audio files. It combines ffprobe metadata with ffmpeg loudness measurements, broad spectral-band observations, silence checks, and optional spectrograms.

It helps choose among versions by separating encoding advantages from mastering/dynamics evidence. Recommendations are tentative and explain their assumptions: technical measurements cannot prove which version sounds best or identify its original source. For important choices, use full analysis and a level-matched listening comparison.

Requirements

  • macOS and Python 3.10+
  • Homebrew ffmpeg (includes ffprobe): brew install ffmpeg
  • pipx

Install

From this repository:

pipx install -e .
soundcheck --help

pipx installs the executable on your PATH, so no .bash_aliases entry is needed. If soundcheck is not found, run pipx ensurepath, open a new shell, and try again.

To use a project-local environment instead:

python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/soundcheck file1.m4a file2.m4a

Update after editing

For the editable pipx installation, source edits take effect immediately. After changing packaging metadata or dependencies, refresh it with:

pipx reinstall soundcheck

For a non-editable future installation, use pipx reinstall soundcheck after updating the source or version.

Usage

soundcheck "Song source A.m4a" "Song source B.webm"
soundcheck *.m4a
soundcheck --spectrogram *.m4a
soundcheck --spectrogram --output ./analysis file1.m4a file2.mp3
soundcheck --full "important master A.flac" "important master B.flac"
soundcheck --json file1.m4a file2.mp3 > comparison.json
soundcheck --verbose file1.m4a file2.mp3

The report shows aligned metadata and loudness/dynamics tables, a tentative overall preference, and pairwise advantages and tradeoffs. Useful pairwise advice remains visible even when a batch has no single winner. --json writes exactly one JSON document to standard output, making it suitable for scripts.

Before analysing, soundcheck prints the fully resolved input list and then reports each file's probe, loudness/dynamics, spectrum, silence, and optional spectrogram stages. These updates go to standard error, so a JSON pipeline remains clean:

soundcheck --json *.m4a > comparison.json

By default, every decode-heavy measurement uses the first 120 seconds: loudness and dynamics, high-frequency bands, leading-silence detection, and optional spectrograms. This keeps large batches predictable. Metadata, file size, and duration always cover the complete file. The report and JSON identify this sampled analysis explicitly; trailing silence is only checked in full mode.

Use --full when comparing a smaller set of important versions. It processes the entire audio file for loudness, spectrum, spectrograms, and both leading/trailing silence. That is more representative for long tracks, mixes, or files whose beginning is unrepresentative, but naturally takes longer.

--spectrogram creates one 1800×1000 logarithmic-frequency PNG per input from the same analysis window (the first two minutes by default, whole file with --full). By default it uses a fresh temporary directory and prints that directory. Use --output DIR to keep images in a chosen location. Existing images are not overwritten.

What the report means

The metadata table reports the container, codec, audio-stream bitrate, sample rate, meaningful source bit depth when available, channels/layout, duration, and file size. When the audio bitrate is missing, a single-stream audio-only container may supply an estimate marked ~; container rates from files with artwork, video or multiple streams are not substituted for audio bitrate. A container is not a quality rating: .m4a may contain AAC or ALAC, while .webm may contain Opus or Vorbis.

The audio table distinguishes several measurements:

Measurement What it tells you How the comparison uses it
Integrated LUFS Average perceived loudness over the analysis window Gives an attenuation amount for level matching; louder does not win
True peak, dBTP Estimated reconstructed waveform peak Flags peaks above 0 dBTP as playback overload risk, not proof of existing clipping
LRA, LU Longer-term variation in loudness Descriptive; a larger value does not automatically mean a better master
PLR, dB True peak minus integrated LUFS Peak-to-average contrast; can help identify differences in peak limiting
Crest factor, dB Sample peak minus unweighted RMS level A second measure of peak-to-average contrast, measured before normalization
RMS, dBFS Average signal power expressed as a level Context for crest factor; it is not perceptually weighted like LUFS

PLR and crest factor are unchanged by a simple gain adjustment, within measurement/quantization tolerances and provided the audio remains above loudness gates and does not clip. If both favor the same file by at least 2 dB, the report suggests that file for greater transient contrast, consistent with less peak limiting if the musical content matches. These are comparison heuristics, not an audibility standard or a standardized DR score. EQ, isolated peaks, edits, silence, and stereo balance can affect them. LRA is intentionally separate: it describes variation over longer intervals and can be misleading for very short clips.

Source peak/RMS statistics are collected in the existing loudness pass, before loudnorm processes the audio. No additional file decode is needed for these metrics. Silent or unmeasurable values appear as — in text and null in JSON.

Measurement references: FFmpeg astats, FFmpeg loudnorm, and EBU Tech 3342 on Loudness Range.

For compatible sample rates, soundcheck samples broad 2 kHz bands above 12 kHz. A sharp, sustained falloff may be reported around 16 kHz or 18–20 kHz. That is an observation—not proof of a bad encode or a prior transcode. Music with little treble, noise reduction, or intentional production choices can have the same appearance. Likewise, energy near the Nyquist frequency does not prove an audible advantage.

Spectrograms are useful for inspecting cutoffs, lossy-to-lossy transcoding clues, and odd artifacts. They are not a hearing test, and they cannot establish the original source. Examine a broad part of a track rather than a single transient.

Comparing codecs and sources

Do not rank files by bitrate alone. AAC, MP3, Opus, and Vorbis have different efficiencies; bitrate is most useful within the same codec, channel layout, and similar sample rate. Higher sample rates are not inherently better either.

Lossless codecs (such as FLAC, PCM WAV and ALAC) preserve their input, but that input may already have been a low-bitrate lossy file. A high-rate MP3 can similarly be a transcode from a weaker source.

The decision rules make those limits explicit:

  • Same lossy codec/profile, sample rate and channels: a bitrate increase of at least 20% and 24 kbps supports a low-confidence encoding preference. For example, 320 kbps MP3 gets a tentative preference over 192 kbps MP3 when there is no conflicting evidence. Small nominal/VBR fluctuations do not determine a winner.
  • Lossless versus lossy: matching sample rate and channels can support a lossless preservation preference, conditional on comparable original sources. It cannot establish provenance or recover detail lost earlier.
  • Different lossy codecs: bitrate alone does not determine the choice. A 256 kbps AAC is not automatically worse than a 320 kbps MP3.
  • Lossless versus lossless: bitrate, sample rate and bit depth alone do not establish a fidelity advantage.
  • Dynamics: corroborating PLR and crest-factor differences support a preference for greater transient contrast. LRA or loudness alone never selects a winner.
  • Tradeoffs: if encoding favors one file and dynamics favor another, no overall choice is made. A materially lower apparent spectral ceiling at 16 kHz or below in an otherwise preferred candidate also triggers a tradeoff. A falloff around 18–20 kHz is reported but does not by itself veto the preference; the broad-band check is too coarse to justify that. High-frequency energy alone never earns a quality score.
  • Comparability: different durations (over one second), channel configurations, analysis windows, or leading offsets prevent ranking that pair. Similar metadata does not prove the same recording; recommendations assume you supplied matching musical content.
  • Batches: a file must have an unopposed preference against every other input to be the overall choice. Ties and incomparable files retain their pairwise explanations.

A loudness difference prompts level matching instead of suppressing all recommendations. Sampled measurements and unknown source provenance keep confidence low. Analysis failures prevent an overall pair preference while preserving any useful metadata observations.

JSON preserves the top-level recommendation fields and adds comparisons, whose preferences identify files by their input paths. New fields under each file's loudness include sample_peak_dbfs, rms_dbfs, crest_factor_db, and peak_to_loudness_ratio_db; bitrate_source identifies an audio-stream rate or container estimate.

YouTube and yt-dlp downloads

YouTube commonly serves already-lossy audio. For comparisons, retain the original stream where practical instead of converting it again:

yt-dlp -f ba -o '%(title)s.%(ext)s' 'VIDEO_URL'
soundcheck "downloaded source.webm" "another source.m4a"

Converting a downloaded Opus/WebM stream to AAC, MP3, WAV, or FLAC cannot recover missing detail. WAV/FLAC output may be useful for compatibility, but it remains a lossless copy of a lossy source. Use soundcheck's report and spectrograms to look for corroborating clues, then make the final choice by level-matched listening on the target player (Apple Music, VLC, rekordbox, or Pioneer hardware).

Development

python3 -m pip install -e '.[test]'
python3 -m pytest

The tests use no large audio fixtures. Integration tests generate short synthetic audio with ffmpeg and skip only when ffmpeg/ffprobe are unavailable.

Limitations and future directions

This version does not fingerprint, align, or ABX-test audio. It cannot reliably distinguish all masters, prove a transcoding history, detect every kind of distortion, or determine playback compatibility. Matching-duration files can still contain different recordings. Full analysis avoids an unrepresentative opening but does not solve alignment or provenance. Waveform alignment, comparison of gain-matched residuals, and blind audition tools would be useful next steps for distinguishing identical masters from different processing.

Contributors

brucek

Issues