WhisperKit large-v2 produces much worse results than whisper-diarization

#528 · open · 2 comments

View on GitHub ↗

Arche151

Diarization I'm seeing noticeably worse results with Whisper large-v2 via WhisperKit in MacWhisper compared with https://github.com/MahmoudAshraf97/whisper-diarization using large-v2. With WhisperKit I constantly get: - Wrong language detection - Skipped sentences - Straight up hallucinations With whisper-diarization, German is detected correctly and the transcription is much more complete and reliable on the same recordings. Since both use large-v2, could this be related to WhisperKit's model conversion, segmentation/VAD, or decoding settings? I am using WhisperKit via MacWhisper by Jordi Bruin btw.

Comments

atiorh

Hello @Arche151! Could you please share an example audio if possible? Whisper Large v2 is generally great but notoriously hard to control and may lead to unexpected results. WhisperKit's implementation has been hardened by 2.5 years of production usage across millions of users but it is always possible that we may have a bug or a recent macOS/iOS upgrade may have led to a previously unobserved issue. It will be great to get to the bottom of this.

Arche151

@atiorh Thanks for getting back to me so quickly! I need to find out whether, I'm allowed to share the examples with you, since they have been produced in a work context. I'll get back to you when I was able do do that. :)