cigi10/ampiear

software based hearing aid for moderate sensorineural hearing loss written in python

★ 0Forks 0PythonGitHub ↗Compare

README

AmpiEar

Overview Pipeline

raw audio → preprocess → analyze → reduce_noise → apply_gain → compress → synthesize → postprocess → enhanced audio

Stages

1. Preprocess

audio → remove_dc_offset (high-pass @ 20Hz, butterworth order=1) → apply_preemphasis (coeff=0.97) → balanced audio
  • Input: raw audio with DC offset
  • Output: centered audio (mean ≈ 0) with boosted high frequencies
  • Parameters: DC_CUTOFF_FREQ=20.0 Hz, PREEMPHASIS_COEFF=0.97

2. Analysis Filter Bank

audio → STFT (n_fft=2048, hop=512, window='hann') → split by frequency → [6 bands: 125-250Hz, 250-500Hz, 500-1kHz, 1k-2kHz, 2k-4kHz, 4k-8kHz] + phase
  • Input: preprocessed audio (sample_rate=16000 Hz)
  • Output: 6 complex STFT matrices (one per band) + phase information
  • Parameters: N_FFT=2048, HOP_LENGTH=512, WINDOW='hann'
  • Bands: [(125,250), (250,500), (500,1000), (1000,2000), (2000,4000), (4000,8000)]

3. Noise Reduction

each band → vad_energy_zcr (energy + zcr thresholds=0.0) → estimate_noise_power (20% quietest frames or VAD=0) → wiener_filter (gain=SNR/(SNR+1), floor=0.001) → denoised band
  • Input: 6 noisy bands + original audio (for VAD)
  • Output: 6 cleaned bands with reduced background noise
  • Parameters: VAD_THRESHOLD=0.02, NOISE_FLOOR=0.01, floor gain=0.001
  • Method: Wiener filter gain = signal_power / (signal_power + noise_power)

4. Insertion Gain (NAL-NL2)

each band → get_band_hearing_loss (interpolate from audiogram) → estimate_band_level (RMS to dB SPL, centered @ 65dB) → compute_nal_nl2_gain (k=0.5, ref_level=65dB, compression_ratio=2.0) → apply_insertion_gain → amplified band
  • Input: 6 cleaned bands + hearing loss profile
  • Output: 6 frequency-specific amplified bands
  • Audiogram: {250:30dB, 500:35dB, 1000:40dB, 2000:45dB, 4000:55dB, 8000:60dB}
  • Parameters: compensation factor k=0.5, reference_level=65dB SPL, COMPRESSION_RATIO=2.0
  • Formula: gain = 0.5 * hearing_loss - (input_level - 65) / 2.0

5. WDRC Compression

each band → envelope_detector (attack=5ms, release=100ms) → compute_compression_gain (threshold=50dB, ratio=3:1, knee=10dB) → apply makeup_gain (+5dB) → compressed band
  • Input: 6 amplified bands
  • Output: 6 compressed bands (soft sounds louder, loud sounds comfortable)
  • Parameters:
    • threshold_db=50.0
    • compression_ratio=3.0 (3:1)
    • attack_time_ms=5.0
    • release_time_ms=100.0
    • makeup_gain_db=5.0
    • knee_width_db=10.0
  • Formula: gain_reduction = (input - threshold) * (1 - 1/ratio) for input > threshold

6. Synthesis Filter Bank

6 compressed bands + phase → recombine_bands (reconstruct full STFT) → ISTFT (hop=512) → single time-domain audio
  • Input: 6 processed bands + original phase
  • Output: single audio signal (may exceed ±1.0)
  • Parameters: HOP_LENGTH=512

7. Post-processing

audio → soft_limit (tanh, threshold=0.95, gain=1.0) → normalize_audio (target=-3dB, headroom=0.1) → safe audio
  • Input: synthesized audio (possibly clipping)
  • Output: safe audio for playback (peak ≤ 0.707)
  • Parameters:
    • soft_limit_threshold=0.95
    • soft_limit_gain=1.0
    • target_level_db=-3.0 (0.707 linear)
    • headroom=0.1 (10% safety margin)
  • Formula: output = tanh(gain * input), then scale to 10^(-3/20) * 0.9 = 0.636

Quick Reference

Function Flow

preprocess() → analyze() → reduce_noise_bands() → apply_gain_all_bands() → apply_wdrc_all_bands() → synthesize() → postprocess()

Data Transformations

[time-domain 16kHz mono] 
  → [time-domain cleaned] 
  → [6 STFT bands (complex)] 
  → [6 denoised bands] 
  → [6 amplified bands] 
  → [6 compressed bands] 
  → [time-domain reconstructed] 
  → [safe time-domain output]

Key Parameters

  • Sample Rate: 16000 Hz
  • Frequency Bands: 6 bands (125-250, 250-500, 500-1k, 1k-2k, 2k-4k, 4k-8k Hz)
  • STFT: n_fft=2048, hop=512, window='hann' (75% overlap)
  • Hearing Loss: Moderate (30-60 dB HL across frequencies)
  • Compression: 3:1 ratio @ 50dB threshold, attack=5ms, release=100ms
  • Output Level: -3 dBFS (peak ≈ 0.707)

Summary

Input: Raw speech audio with background noise
Output: Enhanced audio optimized for hearing loss profile
Processing: 7 stages, 6 frequency bands, multiband dynamic processing
Sample Rate: 16 kHz
Total Latency: ~128ms (2048 samples FFT + overlap-add)
Goal: Restore audibility and comfort for hearing-impaired listeners

Contributors

cigi10

Issues