dennisonbertram
### Work type Bug / regression ### Observed behavior The completion line keeps counting after the run has finished. `Stop()` freezes nothing, so the duration the user reads grows for as long as the line is displayed: ``` tick 0: "· Worked for 5.0s" tick 1: "· Worked for 5.3s" tick 2: "· Worked for 5.6s" tick 3: "· Worked for 5.9s" ``` The completion line is shown for `completionFramesDefault` (10) ticks at the 120 ms tick rate, so the final figure a user sees is inflated by up to ~1.2 s over the true run duration. For a short run that is a large relative error — a 0.4 s run reports as 1.6 s. Reproduction rate: 100%. ### Expected behavior `Worked for <duration>` states how long the run took. The value is fixed at the moment the run stops and does not change while the line is on screen. ### Reproduction ```go m := New(0).Start() m.startTime = time.Now().Add(-5 * time.Second) m = m.Stop(100) m.View(80) // "· Worked for 5.0s" time.Sleep(300 * time.Millisecond) m = m.Tick() m.View(80) // "· Worked for 5.3s" <- still climbing ``` ### User and operational impact Affected users: everyone using the TUI — this line appears after every run. Severity: low-moderate. Nothing breaks, but the one number the TUI reports about run duration is wrong, and wrong in a direction that flatters nothing and confuses timing comparisons. Anyone using it to judge whether a change made runs faster is reading a figure with up to 1.2 s of noise added after the fact. No data or security implications. ### Suspected seam and search evidence `cmd/harnesscli/tui/components/spinner/model.go`. - `View` renders the completion state as `return m.CompletionLine(m.ElapsedSeconds())`. - `ElapsedSeconds()` returns `time.Since(m.startTime).Seconds()` — a live clock read, with no notion of the run having ended. - `Stop(tokens)` sets `active=false`, `done=true`, `tokens`, and `completionFrames`. It records no stop time. So every re-render during the completion window recomputes the elapsed time against the wall clock. The spinner ticks throughout that window (that is what counts `completionFrames` down), so the value is recomputed roughly every 120 ms. Callers searched: `ElapsedSeconds` is called from `View` (completion branch) and is exported, so it may have external callers — `grep -rn "ElapsedSeconds" cmd/` before changing its semantics. The safer change is to freeze at `Stop` and have the completion path read the frozen value, leaving `ElapsedSeconds` meaning "time since start" for a live spinner. Provenance: found by `openai-gpt-oss-120b` ($0.07/M input) via the Surplus proxy during a cheap-model evaluation, then confirmed locally by rendering the completion line across its display window. `gpt-6-astra` ($10/M) reviewed the same file and did not report it — it raised a related but differently-framed point about elapsed time being non-deterministic in snapshot tests, which is a false positive, since those tests set `startTime` explicitly. ### Blast-radius impact map Callers and data flow: `Stop`, `View`'s completion branch, and `CompletionLine`. Confined to the spinner package. Config/env/defaults: none. API/CLI/wire formats/tools: `CompletionLine(seconds float64)` keeps its signature — it already takes the duration as an argument, which is what makes this fixable without changing the API. `ElapsedSeconds()` keeps its current meaning. Persistence/schema/cache: none. Concurrency/lifecycle: the `Model` stays an immutable value; freezing means storing one more field at `Stop`. Security/auth/permissions/privacy: none. TUI/web/macOS/other clients: TUI only. Provider/model/tool catalog: none. Deployment/observability/runbooks: none. Compatibility: the displayed number changes — it stops growing. That is the fix, not a regression. Existing tests/fixtures: `completion_test.go` asserts the completion line contains `"Worked for"` and a duration, and sets `startTime` explicitly. Those keep passing. The committed snapshots render a fixed `5.0s` and are unaffected because they capture a single frame. Documentation: `docs/logs/engineering-log.md`. ### Regression test first `cmd/harnesscli/tui/components/spinner/completion_test.go`: `TestCompletionDurationIsFrozenAtStop` — start a spinner with `startTime` 5 s in the past, `Stop`, capture the rendered line, then advance the clock (sleep) and `Tick` several times within the completion window, asserting every subsequent render is byte-identical to the first. Red output before the fix: `"· Worked for 5.3s"` where `"· Worked for 5.0s"` was expected. Why it proves the bug: it asserts the property that matters — the reported duration does not depend on when you look at it — rather than asserting a particular number, which would be timing-dependent and flaky. False-positive controls: keep the existing assertions that the line still contains `"Worked for"` and a plausible duration, so a fix that freezes the value at zero or drops it entirely fails. Also assert the value still reflects the real elapsed time at stop (≈5 s), not a constant. ### Fix boundaries In scope: - Record the elapsed duration at `Stop` and render that in the completion line. - Two cleanups in the same file, from the same review, kept explicitly small: - `shortenLabel` takes an unused `full` parameter (found by `gpt-5-nano`; confirmed — the name appears only in the signature). Remove it. - `Tick` indexes `holds[m.step]` while advancing `m.step` modulo `len(pulse)`. The two slices must stay the same length; a guard or a comment making the invariant explicit prevents a future edit turning it into an out-of-range panic. Out of scope: - Changing `ElapsedSeconds()`'s meaning for a live spinner. - The completion window length, the cancel hint, the label ladder, and the pulse cadence. - `WithStyles` taking a value and storing its address — also raised by the review, but the copy is deliberate and the pattern is fine for an immutable value type. Not a defect. ### Diagnostic and observability evidence Before: render the completion line four times across the display window with sleeps between; the duration climbs 5.0 → 5.3 → 5.6 → 5.9. After: all four renders are identical. ### Verification plan - Red: run the new test before the fix; record the climbing value. - Green: same test after. - Targeted: `go test ./cmd/harnesscli/tui/components/spinner/... -race -v`. - Full regression: `go test ./cmd/... ./internal/...`. - Real path: run a real TUI turn and watch the completion line settle on one number, since the whole point is what a human reads after a run. ### Rollout and rollback Single PR to `main`; picked up on the next `scripts/install.sh`. No migration or persisted state. Rollback is reverting the commit. ### Documentation and handoff `docs/logs/engineering-log.md`: the defect, the cause (a live clock read on every re-render of a finished run), and the provenance — found by a $0.07/M model that a $10/M model missed on the same file, which is the more useful lesson than the bug itself. ### Definition of done - [ ] Test written first, observed failing on a climbing duration - [ ] Duration frozen at `Stop` and identical across the completion window - [ ] Value still reflects real elapsed time, not a constant - [ ] Unused `full` parameter removed - [ ] `pulse`/`holds` length invariant made explicit - [ ] `ElapsedSeconds()` semantics unchanged for a live spinner - [ ] `go test ./cmd/harnesscli/tui/... -race` green; full regression green - [ ] Real TUI turn watched, not only asserted - [ ] Engineering log records the provenance