Early-exit on forced prefix fires even without `promptTokens` — same class of bug as #372/PR #497, but that fix doesn't cover this path

#525 · open · 2 comments

View on GitHub ↗

gabrieImoreira

# Título Early-exit on forced prefix fires even without `promptTokens` — same class of bug as #372 / PR #497, but that fix doesn't cover this path # Corpo Related to #372 and PR #497 (thank you @yangzichao for the root-cause writeup there) — but this reproduces the same failure mode **without `promptTokens` set at all**, which PR #497 explicitly leaves untouched ("The no-prompt path is untouched — checks keep their existing indices, so default behavior is unchanged"). ## Symptom Calling `transcribe(audioArray:decodeOptions:)` once per VAD-detected speech window of a real ~87min pt-BR meeting recording (no `promptTokens`, `large-v3-v20240930_turbo` and non-turbo both reproduce), a large fraction of otherwise-valid speech windows return **completely empty text** — `avgLogprob == 0.0`, `temperature == 0.0` (the very first attempt is treated as a trivially-complete success), no error thrown, no temperature fallback triggered. In one real recording, 100/100 and 87/91 windows on the two tracks came back empty at temperature 0. This is audio-content dependent per window (not duration, not concurrency, not `chunkingStrategy`, not `wordTimestamps` — all ruled out live), which points at the decode loop reacting to what the model predicts, not a config issue. ## Root cause (same mechanism as #372, different forced token) In `TextDecoder.decodeText`'s main loop: ```swift let isSegmentCompleted = sampleResult.completed || currentTokens.count >= Constants.maxTokenContext - 1 || isFirstTokenLogProbTooLow ``` `sampleResult.completed` (EOT was the model's top prediction) is evaluated **unconditionally**, including at forced-prefix positions — same defect #372 found for the `promptTokens` case. Without `promptTokens`, `usePrefillCache` (default `true`) prefills SOT+language+task via the separate KV-cache-prefill model, so `prefilledIndex` (= `cacheLength[0]`) is `3`. But the prefill sequence is 4 tokens (`[SOT, language, task, timestamp]`) — the 4th (timestamp/no-timestamps token) is *not* covered by that prefill and still goes through the main loop, at `tokenIndex == prefilledIndex == 3`. If the model's prediction for the *next* token right after being forced through that timestamp position happens to be EOT, `sampleResult.completed` fires and the loop breaks before a single real content token is sampled — for a window VAD already confirmed contains speech. PR #497's fix is scoped to exactly the `promptTokens` case: ```swift let hasPromptTokens = !(options.promptTokens?.isEmpty ?? true) let ignoresSampledPrediction = hasPromptTokens && isPrefill ... let isSegmentCompleted = (sampleResult.completed && !ignoresSampledPrediction) || ... ``` `isPrefill = tokenIndex < initialPromptIndex - 1`, which is also `false` at `tokenIndex == 3` in the no-prompt case (`initialPromptIndex - 1 == 3`) — so even applying #497 as-is wouldn't change anything here; `hasPromptTokens` alone already keeps `ignoresSampledPrediction` false for every no-prompt call. ## Suggested fix Generalize the guard so a forced-prefix position never drives `isSegmentCompleted`, regardless of `promptTokens`: ```swift let ignoresSampledPrediction = tokenIndex < initialPromptIndex ``` (or equivalent — the point is "any position still forcing a known token, prompt or not, shouldn't have its trailing prediction treated as segment-completion"). This matches the reference Whisper implementation's behavior of only running completion checks from `sample_begin` onward, universally, not just when a prompt is present. ## Workaround found, not a real fix Re-decoding an empty-result window at a non-zero temperature (0.4–0.8) recovers *some* of the lost windows (sampling escapes the spurious EOT at the same forced position), but it's unreliable — re-running the identical window at the identical temperature gives different pass/fail results run to run, and on one real recording a two-step temperature ladder only recovered ~5% of the empty windows. Not something to build a real pipeline on top of. ## Repro environment - argmax-oss-swift 0.18.0 (`e2adabbe7d98dc4d0ab9a5b75424ecc42a9cdbef`) - macOS 26.6.1, Apple M4 Pro - Models: `openai_whisper-large-v3-v20240930_turbo` and `openai_whisper-large-v3-v20240930` (non-turbo) — both reproduce - Language: pt (Portuguese), `detectLanguage: false` - No `promptTokens`, `chunkingStrategy: .none`, called via `transcribe(audioArray:)` once per pre-segmented (external VAD) speech window Happy to share a redacted/synthetic repro harness if useful — the real audio is a work meeting I can't share directly.

Comments

ZachNagengast

Hi @gabrieImoreira thanks for the report, looking into this but any audio / configs you can provide that would reproduce would be helpful as well!

ZachNagengast

I also just noticed your repro is using v0.18.0 >argmax-oss-swift 0.18.0 (e2adabbe7d98dc4d0ab9a5b75424ecc42a9cdbef) Are you able to reproduce this on the current v1.1.0?