raw audio → preprocess → analyze → reduce_noise → apply_gain → compress → synthesize → postprocess → enhanced audio
audio → remove_dc_offset (high-pass @ 20Hz, butterworth order=1) → apply_preemphasis (coeff=0.97) → balanced audio
- Input: raw audio with DC offset
- Output: centered audio (mean ≈ 0) with boosted high frequencies
- Parameters:
DC_CUTOFF_FREQ=20.0 Hz,PREEMPHASIS_COEFF=0.97
audio → STFT (n_fft=2048, hop=512, window='hann') → split by frequency → [6 bands: 125-250Hz, 250-500Hz, 500-1kHz, 1k-2kHz, 2k-4kHz, 4k-8kHz] + phase
- Input: preprocessed audio (sample_rate=16000 Hz)
- Output: 6 complex STFT matrices (one per band) + phase information
- Parameters:
N_FFT=2048,HOP_LENGTH=512,WINDOW='hann' - Bands:
[(125,250), (250,500), (500,1000), (1000,2000), (2000,4000), (4000,8000)]
each band → vad_energy_zcr (energy + zcr thresholds=0.0) → estimate_noise_power (20% quietest frames or VAD=0) → wiener_filter (gain=SNR/(SNR+1), floor=0.001) → denoised band
- Input: 6 noisy bands + original audio (for VAD)
- Output: 6 cleaned bands with reduced background noise
- Parameters:
VAD_THRESHOLD=0.02,NOISE_FLOOR=0.01, floor gain=0.001 - Method: Wiener filter
gain = signal_power / (signal_power + noise_power)
each band → get_band_hearing_loss (interpolate from audiogram) → estimate_band_level (RMS to dB SPL, centered @ 65dB) → compute_nal_nl2_gain (k=0.5, ref_level=65dB, compression_ratio=2.0) → apply_insertion_gain → amplified band
- Input: 6 cleaned bands + hearing loss profile
- Output: 6 frequency-specific amplified bands
- Audiogram:
{250:30dB, 500:35dB, 1000:40dB, 2000:45dB, 4000:55dB, 8000:60dB} - Parameters: compensation factor k=0.5, reference_level=65dB SPL,
COMPRESSION_RATIO=2.0 - Formula:
gain = 0.5 * hearing_loss - (input_level - 65) / 2.0
each band → envelope_detector (attack=5ms, release=100ms) → compute_compression_gain (threshold=50dB, ratio=3:1, knee=10dB) → apply makeup_gain (+5dB) → compressed band
- Input: 6 amplified bands
- Output: 6 compressed bands (soft sounds louder, loud sounds comfortable)
- Parameters:
threshold_db=50.0compression_ratio=3.0(3:1)attack_time_ms=5.0release_time_ms=100.0makeup_gain_db=5.0knee_width_db=10.0
- Formula:
gain_reduction = (input - threshold) * (1 - 1/ratio)for input > threshold
6 compressed bands + phase → recombine_bands (reconstruct full STFT) → ISTFT (hop=512) → single time-domain audio
- Input: 6 processed bands + original phase
- Output: single audio signal (may exceed ±1.0)
- Parameters:
HOP_LENGTH=512
audio → soft_limit (tanh, threshold=0.95, gain=1.0) → normalize_audio (target=-3dB, headroom=0.1) → safe audio
- Input: synthesized audio (possibly clipping)
- Output: safe audio for playback (peak ≤ 0.707)
- Parameters:
soft_limit_threshold=0.95soft_limit_gain=1.0target_level_db=-3.0(0.707 linear)headroom=0.1(10% safety margin)
- Formula:
output = tanh(gain * input), then scale to10^(-3/20) * 0.9 = 0.636
preprocess() → analyze() → reduce_noise_bands() → apply_gain_all_bands() → apply_wdrc_all_bands() → synthesize() → postprocess()
[time-domain 16kHz mono]
→ [time-domain cleaned]
→ [6 STFT bands (complex)]
→ [6 denoised bands]
→ [6 amplified bands]
→ [6 compressed bands]
→ [time-domain reconstructed]
→ [safe time-domain output]
- Sample Rate: 16000 Hz
- Frequency Bands: 6 bands (125-250, 250-500, 500-1k, 1k-2k, 2k-4k, 4k-8k Hz)
- STFT: n_fft=2048, hop=512, window='hann' (75% overlap)
- Hearing Loss: Moderate (30-60 dB HL across frequencies)
- Compression: 3:1 ratio @ 50dB threshold, attack=5ms, release=100ms
- Output Level: -3 dBFS (peak ≈ 0.707)
Input: Raw speech audio with background noise
Output: Enhanced audio optimized for hearing loss profile
Processing: 7 stages, 6 frequency bands, multiband dynamic processing
Sample Rate: 16 kHz
Total Latency: ~128ms (2048 samples FFT + overlap-add)
Goal: Restore audibility and comfort for hearing-impaired listeners