[P1] --timeout remains blocked until the Linux payload duration and underreports elapsed time

#20 · closed · 1 comments

View on GitHub ↗

Azathothas

Tested the published **wsl-toolkit-v1.3.0 Windows amd64** binary on Windows 11/WSL2 on 2026-09-10. SHA-256: `71b2ef9b0cf50da0acda45e7cc9293185a2b18d8aa403812e20ed29222d54424`, matching the release manifest. Related: #6. ## Summary `--timeout` eventually returns exit 124, but it does not bound caller wall time: the CLI remains blocked until approximately the original Linux payload duration. Its reported job/matrix durations are also much shorter than the time the caller actually waited. ## Reproduction and measurements With a healthy helper/base, measured using `System.Diagnostics.Stopwatch` around each process: ```powershell & $wtk run --image alpine --timeout 2s -c 'printf "run-timeout-start\n"; sleep 8' & $wtk matrix --images alpine --timeout 2s -c 'printf "matrix-timeout-start\n"; sleep 8' ``` Repeated run result: - process exit: 124 - CLI-reported duration: 4.422 s - caller wall time: 11.981 s Matrix result: - row: timeout / exit 124 - row-reported duration: 3.765 s - summary: `in 9s` - caller wall time: 11.514 s - process exit: 1, per matrix policy An earlier `sleep 20` / `--timeout 2s` similarly kept the caller blocked for about 20 seconds. The payload never printed its post-sleep marker, so termination is eventually recognized, but the deadline is not a usable wall-time bound. ## Source observation The run context wraps `wsl.exe`; Windows cancellation attempts `taskkill /PID ... /T /F` and uses a one-second `WaitDelay`. `Runner.Run` records duration immediately after `wsl.Exec` and then performs container cleanup. The observed remaining delay appears to involve the WSL/Linux descendant or stream relay outliving the canceled host process. The report should not assume that taskkill itself is absent. ## Expected The process should return close to the requested deadline plus a small documented cleanup grace, regardless of the payload's original sleep duration. Reported job and fleet durations should cover the same end-to-end interval the caller experiences. Add wall-clock assertions for both direct/helper run and matrix, including a child process that inherits stdout/stderr.

Comments

Azathothas

Fixed in `wsl-toolkit-v2.0.0`, by `WSL-45` in `TODO/wsl-toolkit-go.md`. Two separate defects, and the report named both. The deadline bounded the **child** and not the caller, so the process stayed blocked until roughly the payload's own duration; and `duration_ns` stopped at the container's death, so the number a caller read was shorter than the time they waited. The deadline is the caller's wall time now, and `duration_ns` is the interval the caller waited, cleanup included. Most of the missing seconds were `podman rm` on a running container: measured here at 10.63s, against 0.56s with `-t 0`. Re-derived against the published v2.0.0 binary, 2026-09-10: ```text $ wsl-toolkit run --image alpine --timeout 2s -c 'printf "start\n"; sleep 8' exit=124 caller waited 4.74s duration_ns says 4.37s timed_out=True ``` Against `v1.3.0` the same command reported 4.42s and held the caller for about 15s. `duration_ns` changing meaning is a break and `docs/consumers.md` carries the row.