[Bug]: Every worker-pool thread runs its own RollLogFileTask on the shared tunarr.log (ENOENT race at midnight); followed by 6 worker threads spinning at ~65% CPU each until restart

#2193 · open · 1 comments

View on GitHub ↗

smidley

### This issue respects the following points: - [x] This is a **bug**, not a question or a configuration issue. - [x] This issue is **not** already reported on GitHub (searched; #546 from 2024 is about truncation, not the worker pool). - [x] I'm using an up to date version of Tunarr (2026.9.2). ### Description of the bug Two things, the first provable from the code, the second observed on my server and possibly caused by the first. **1. Log rolling runs once per thread on the same file.** `RollingLogDestination.initDestination()` creates a `ScheduledTask(RollLogFileTask, …)` for each destination instance, and each worker-pool thread builds its own logger/destination for the same `tunarr.log`. With the default pool (8 workers) that is nine independent roll tasks firing at the same cron tick on the same files. `roll()` copies `tunarr.log` to `tunarr.log.tmp`, truncates, renames the numbered files and then renames `.tmp` to `.1` — so the threads race on `.tmp` and on the numbered files. Tonight's run produced exactly that: - worker: `Task RollLogFileTask ran in 38.39 ms and failed` → `ENOENT: rename '/config/tunarr/logs/tunarr.log.tmp' -> '/config/tunarr/logs/tunarr.log.1'` - `Error rotating /config/tunarr/logs/tunarr.log.4` (twice), `Error rotating /config/tunarr/logs/tunarr.log.1` - `Error while deleting log file /config/tunarr/logs/tunarr.log.5 … ENOENT` Afterwards the logs directory held `tunarr.log`, `tunarr.log.1` (5.6 KB) and `tunarr.log.3`, no `.2`. Before the restart the process had 15 open handles on `tunarr.log` and 6 on `tunarr.log.3` (one SonicBoom per thread). This looks like the same class of problem #2175 fixed for `settings.json` ("workers bootstrap through the same entry point"): the roll task should probably run on the main thread only, or workers should log through the parent. **2. Worker threads spinning at full CPU with nothing to do.** About 14 hours after the container was recreated on 2026.9.2, the `tunarr` process was at 540–606% CPU (docker stats) with no streams, no ffmpeg children, no HTTP connections, and every scheduled task idle (`/api/tasks` all `running: false`). Per-thread sampling from the host showed 6 of the 8 worker threads (consecutive TIDs in the worker block) each at 60–75% CPU, roughly 40% of it system time, `/proc/<tid>/syscall` = `running` on every sample, ~750 involuntary context switches per second per thread, zero page faults, zero bytes read or written. The main thread, libuv workers and Meilisearch were idle. That pattern is an event loop spinning (0-delay timer/immediate or similar), not real work. Memory was flat at ~1.5 GiB. Nothing was logged while it spun. A container restart cleared it (0.2% CPU since, same 2026.9.2 image). I could not get a CPU profile: the packaged binary does not enable the inspector on SIGUSR1 (it terminates the process instead), and the container has no strace/perf. The only worker-side event in the log during the process's life was the failed midnight roll above, and CPU-time accounting (63.5 CPU-hours over a 14-hour life at ~5–6 cores) is consistent with the spin starting around that midnight tick, but I cannot prove the link. If the roll failure can leave a worker's scheduler or SonicBoom in a hot loop, that would explain it; if not, treat item 2 as a separate report and I will add data if it recurs after the next midnight roll. ### Reproduction steps 1. Docker install, worker pool at the default size (8 workers started at boot). 2. Logging settings: log roll enabled, `every 1 day`, `maxFileSizeBytes` 1048576, `rolledFileLimit` 3 (`/api/system/settings` → `logging.logRollConfig`). 3. Let the process run across midnight (roll tick). 4. Watch the container log at 00:00 local and `ls /config/tunarr/logs`. 5. For item 2: check per-thread CPU of the `tunarr` process a few hours later (`docker stats`, then `top -H -p <pid>` or `/proc/<pid>/task/*/stat`). ### What is the current _bug_ behavior? Nine roll tasks collide on the same files at midnight (ENOENT rename/unlink, "Error rotating" warnings, numbered files skipped). On this server, 6 worker threads then spun at ~65% CPU each for ~10 hours with no work, until the container was restarted. ### What is the expected _correct_ behavior? One roll per file per tick (main thread only, as with the settings/migration work in #2175), no ENOENT noise, and idle workers at ~0% CPU. ### Tunarr version Latest ### FFMPEG encoder type QSV (QuickSync) ### Deployment Type Docker ### What operating system are you using? Unraid 7.3.2 (Docker, bridge network, `ghcr.io/chrisbenincasa/tunarr` 2026.9.2 = image rev `7496a6dc`; `/api/version` → `{"tunarr":"2026.9.2","ffmpeg":"7.1.1","nodejs":"22.20.0"}`; Jellyfin media source; 25 HLS channels) ### Full server logs ```shell 2026-09-28T20:32:22.714Z [info]: Running startup task SeedSystemDevicesStartupTask 2026-09-28T20:32:22.716Z [info]: Running startup task RefreshLibrariesStartupTask 2026-09-28T20:32:26.678Z [info]: Starting Meilisearch service... 2026-09-28T20:32:26.692Z [info]: Meilisearch service started on port 43661 2026-09-28T20:32:30.433Z [info]: Tunarr worker started {"worker":true} (x8) 2026-09-29T00:00:00.113Z [warn]: Task RollLogFileTask ran in 38.39 ms and failed {"worker":true} err: { "type": "Error", "message": "ENOENT: no such file or directory, rename '/config/tunarr/logs/tunarr.log.tmp' -> '/config/tunarr/logs/tunarr.log.1'", "stack": Error: ENOENT: no such file or directory, rename '/config/tunarr/logs/tunarr.log.tmp' -> '/config/tunarr/logs/tunarr.log.1' at Object.renameSync (node:fs:1020:11) at die.roll (/snapshot/dist/bundle.cjs:952:38802) at pie.runInternal (/snapshot/dist/bundle.cjs:952:39626) at pie.run (/snapshot/dist/bundle.cjs:952:35293) at Nl.runJobInternal (/snapshot/dist/bundle.cjs:952:34227) at Nl.run (/snapshot/dist/bundle.cjs:952:33644) at vV.job (/snapshot/dist/bundle.cjs:952:33263) at vV.invoke (/snapshot/dist/bundle.cjs:55:10831) at /snapshot/dist/bundle.cjs:55:7254 at Timeout._onTimeout (/snapshot/dist/bundle.cjs:55:6666) at listOnTimeout (node:internal/timers:588:17) at process.processTimers (node:internal/timers:523:7) "errno": -2, "code": "ENOENT", "syscall": "rename", "path": "/config/tunarr/logs/tunarr.log.tmp", "dest": "/config/tunarr/logs/tunarr.log.1" } Error rotating /config/tunarr/logs/tunarr.log.4 2026-09-29T00:00:01.615Z [info]: Building guide info for channel 038a7708-... {"category":"scheduling"} (x25 channels) 2026-09-29T00:00:01.627Z [info]: Scanning collections for jellyfin media source server (ID = 96aa2943-...) Error rotating /config/tunarr/logs/tunarr.log.4 Error rotating /config/tunarr/logs/tunarr.log.1 Error while deleting log file /config/tunarr/logs/tunarr.log.5 Error: ENOENT: no such file or directory, unlink '/config/tunarr/logs/tunarr.log.5' at Object.unlinkSync (node:fs:1952:11) at /snapshot/dist/bundle.cjs:952:39334 at attemptSync (/snapshot/dist/bundle.cjs:845:407517) at die.checkFileRemoval (/snapshot/dist/bundle.cjs:952:39236) at die.roll (/snapshot/dist/bundle.cjs:952:38936) at pie.runInternal (/snapshot/dist/bundle.cjs:952:39626) ... 2026-09-29T00:00:02.633Z [info]: XMLTV Updated at 2026-09-29T00:00:02-07:00 {"task":"UpdateXmlTvTask"} # ~10 h later, nothing streaming, all tasks idle: $ docker stats --no-stream Tunarr -> 606% / 604% / 602% CPU, 1.52 GiB, 72 pids $ per-thread (5 s window, 100% = 1 core): tid 1315089 tunarr user 23% sys 53% tid 1315088 tunarr user 38% sys 36% tid 1315085 tunarr user 32% sys 37% tid 1315083 tunarr user 20% sys 46% tid 1315086 tunarr user 20% sys 44% tid 1315084 tunarr user 18% sys 43% main thread / libuv-worker / DelayedTaskScheduler: 0% io: rchar 0 wchar 0 read_bytes 0 write_bytes 0 ; minflt 0 ; ~750 nonvoluntary ctx switches/s per hot thread /proc/<tid>/syscall: "running" 30/30 samples on each hot thread $ /api/tasks: every task running:false ; no ffmpeg processes ; no established TCP connections to :8000 # after `docker restart`: 0.2% CPU, workers 0-7 started normally. ```

Comments

chrisbenincasa

Thanks for the detailed report — the per-thread breakdown made this much easier to chase. **Part 1 (log rolling) is confirmed and fixed in #2195.** Every worker thread built its own log roller, so all nine threads rolled `tunarr.log` at the same tick and raced on the same files. That's exactly the ENOENT and "Error rotating" output you saw. After the fix, only the main thread rolls, and workers just append to the live file. While in there we also found that size-based rolling (`maxFileSizeBytes`) had never actually run. It works now, so with your settings `tunarr.log` will roll at 1 MiB as well as at midnight. **Part 2 (the CPU spin) we couldn't reproduce or trace.** We looked for a way a failed roll could leave a worker's scheduler or log writer in a hot loop and didn't find one. The fix above removes the roll race as a suspect. If the spin comes back on a build with #2195, please post the same per-thread numbers, the time it started, and whatever the log shows around then. That would tell us it has a separate cause. We'll leave the issue open for that.