etancen
Feeding a window of pure digital silence into `OpenMOSS-Team/MOSS-Transcribe-Diarize` returns a complete, confident sentence that is not in the audio: ``` I'm sorry, I can't assist with that request. I'm a Qwen-1 model developed by Qwen-Omni, a new generation of Qwen models. I don't have the right to access that information. ``` Measured while running a windowed real-time transcription pipeline (window 20 s, hop 5 s, bf16, HF backend, RTX A5500 Laptop) over a 70-second clip whose 30–45 s region is exactly zero samples. The window covering 31.9–41.9 s returned that text, and the pipeline committed it as an ordinary 10-second transcript segment. Audio of that window, checked with `soundfile`: ``` rms = 0.00000 peak = 0.00000 ``` Speech elsewhere in the same file transcribes correctly, so this looks specific to silence rather than to the file or the decoding settings. Is there a recommended way to ask the model for "no speech here" instead of assistant-style boilerplate? A windowed/streaming pipeline hits silent windows regularly — the design's coverage guarantee forces one through every `window - hop` seconds — so a model-side answer (or an explicit "empty" token) would help every streaming use, not just ours. For now we drop the result whenever the window's audio is silent, which is a workaround at the pipeline level rather than a fix at the model level. Context: this came out of the real-time mode proposed in #58.