mkaflowski
### Description With many reference voices, the generated output sounds heavily compressed and has audible crackling, as if the signal is overdriven or clipping. The reference recordings themselves are clean, with no distortion. ### Steps to reproduce 1. Use the attached reference audio as the voice prompt. 2. Generate with **default settings** using this text (Polish): > Mieszkam z Vergilem w Rotterdamie, niedaleko miejsca, w którym dorastał. Jest bardzo... nowoczesny. Zniszczony na wojnie, a potem... odbudowany. Dobre miejsce na nowy początek. 3. Listen to the output. ### Expected behavior Clean speech whose quality is comparable to the reference recording. ### Actual behavior - The output sounds over-compressed and has reduced dynamics. - Crackling and clipping-like distortion is audible, especially on louder or stressed syllables. - This happens with many different voices, not just one. - With the higher CFG it gets worse ### Attachments - `reference.wav`: original voice prompt - `generated.wav`: output with default settings ### Environment - VoxCPM2 version / commit: <!-- e.g. 2.x.x or commit hash --> - Installation: <!-- pip / from source --> - OS: <!-- e.g. Windows 11 / Ubuntu 24.04 --> - GPU / CUDA: <!-- e.g. RTX 4070, CUDA 12.4 --> - PyTorch: <!-- version --> - Python: <!-- version --> ### Additional notes - Is this a known issue with non-English (Polish) text? - Are there recommended settings to reduce it, such as CFG value, inference timesteps, or normalization of the reference audio? Original: [Rufius_I_live_with_Vergil_in.wav](https://github.com/user-attachments/files/32683287/Rufius_I_live_with_Vergil_in.wav) Generated: [Rufius_I_live_with_Vergil_in.wav](https://github.com/user-attachments/files/32683291/Rufius_I_live_with_Vergil_in.wav)