I'm seeing the timestamp drift by i.e. almost by 10s. See following [audio.mp3](https://github.com/user-attachments/files/31648756/audio_example.mp3):
> Así es, así es, también músico, estaban yendo a ese [...]
That happens in the audio at 02:05 (mm:ss), while the model predicts at `01:54` the start of that phrase:
<img width="1098" height="126" alt="Image" src="https://github.com/user-attachments/assets/2ad1be83-2d80-4c2b-8a6b-139b378aee80" />
Me too. When the audio doesn't start immediately, there's a "silence" segment before it, and the first few transcriptions will be shifted by 10 seconds.
00:00 ~ 05:39:No output, correct. (Music playing and no one speaking.)
05:39 ~ 07:30:Transcriptions be 10 seconds slower.
As transcription progresses, the 07:30 mark is the dividing line.
7:30 ~ 23:36:From this point onwards, the model output is very accurate until the end.
Audio from podcast: [
VTuber學|關於我推的學問 ft. 貓頭鷹出版社 王正緯](https://www.youtube.com/live/N9boWvU-KkA)