### Operating system
💠Windows 11
### Runtime method
📦 PyApp executables, 🔥 Source code (git main)
### Project version
_No response_
### Python version
Auto-managed
### GPU Brand
_No response_
### GPU Model
_No response_
### Description
On Windows 11, when using ShaderFlow’s export feature to render a scene to video, the rendering process reaches 100% but never terminates. The application appears stuck, waiting for FFmpeg to exit. Examining scene.py shows that FFmpeg writes logs to stderr, and ShaderFlow’s code is never reading from that pipe. This leads to a deadlock: FFmpeg blocks trying to write, while the application is blocked in process.wait().
│DepthFlow├┤0'11.052├┤INFO │ ▸ (Module 1 • DepthScene ) Waiting for FFmpeg process to finish encoding (Queued writes, codecs lookahead, buffers, etc)
However, the process never actually finishes encoding because FFmpeg has filled the stderr pipe buffer, causing a complete stall.
Steps to Reproduce
Set up a ShaderFlow scene that exports a long video (5 seconds or more with 60 fps)
Observe that it reaches frame 100% with logs indicating everything is rendered.
The application never exits from the call to process.wait().
Console shows:
"Waiting for FFmpeg process to finish encoding (Queued writes, codecs lookahead, buffers, etc)"
Quick fix:
Attach stderr to the console (no pipe)
`
stderr=None: Inherits your Python process’s stderr, so logs appear in the terminal
stderr=subprocess.DEVNULL: Discards all logs
stderr=sys.stderr: Same effect as None if you want to unify logs
For example:
python
Copy
Edit
self.process = self.scene.ffmpeg.popen(
stdin=PIPE,
stderr=None, # or stderr=sys.stderr or stderr=subprocess.STDOUT
)
`
That way, there is no separate buffer for your script to drain. FFmpeg can just write its log output normally, so no deadlock occurs.
### Traceback
```shell
```
Hmmph. I tested the code in a Windows machine two days ago with a friend. Also hanged, but it was the current git main code, and I was about to try it myself as it's fine on Linux and macOS and could have been due a recent refactor.
I'm pretty sure it worked on the current PyApp releases (v0.8.0), otherwise there would be massive user complaints before. Must have been some broken build or commit in BtbN's [FFmpeg Builds](https://github.com/BtbN/FFmpeg-Builds), as I'm downloading the latest release if not found in the system, and this version did _not_ set `stderr` to the subprocess pipe, so I doubt it's the source of error.
Could you try downloading FFmpeg elsewhere (like WinGet) or some older [build](https://github.com/BtbN/FFmpeg-Builds/releases) of BtbN (1w+ old), having it on PATH and deleting the folder where it downloaded stuff? (IIRC `%LocalAppData%\BrokenSource`)
Will try it myself later today too!
Chatgpt o1 found the problem, the explanation makes sense and the suggested fix did fix the problem.
I will just link the full conversation with his detailed explanation so you can have a look:
[https://chatgpt.com/share/67d8588d-14d0-8008-a239-0a3f0fd2bcb2](https://chatgpt.com/share/67d8588d-14d0-8008-a239-0a3f0fd2bcb2)
Just tested here on Windows 11, no halts with any combination of (v0.8.0 executable • git main code) with (winget • latest BtbN) FFmpeg, both (`stderr=PIPE` • `stderr=None`) for git main. Oh well, don't know what's wrong ðŸ«
I agree that o1 _thinks_ there's a deadlock, that is well said in Python [subprocess docs](https://docs.python.org/3/library/subprocess.html#subprocess.Popen.wait) on `.wait()` under certain conditions, and could be the case given that it works for you without catching it. Btw, their solution to use [`Popen.communicate()`](https://docs.python.org/3/library/subprocess.html#subprocess.Popen.communicate) isn't the way here: any write with `communicate` expects a read, but we're piping all video frames blindly without any synchronization
The PyApp v0.8 uses `-loglevel error` and git main `-loglevel info` for FFmpeg; while it's my bad that it should be `error` on git main, there's simply not enough status output in small renders to fill the stderr buffer where FFmpeg writes actual info, as `stdout` is reserved for pipe outputs (unless it's unbuffered, which it isn't by default, or is somehow spammimg messages)
The reason for catching stderr is purely for logging errors on the WebUI, if it causes less issues not having it, I'm fine removing
Will call my friend and try to understand what's going different in his machine, hoping to get back with answers! 😓
Done, refactored the code and added a potential fix in https://github.com/BrokenSource/ShaderFlow/commit/b369dcae53d4d0b6cee6cd3b849f896cd2b24c70, while still catching `stderr` and `stdout` outputs!
Could you test it with either:
- [Astral/uv](https://docs.astral.sh/uv/getting-started/installation/) (v0.6.8+): `uvx --from git+https://github.com/BrokenSource/DepthFlow depthflow gradio`
- Recloning or updating what you had with `git pull --recurse-submodules`, activating venv, `depthflow gradio`
Tested on Windows and I don't see more halts even if spamming stderr with `-loglevel verbose` 🙂
#### Explanation (fun one)
Yes, ultimately it was the `os.pipe()` created by `subprocess.Popen` with both `stdin=PIPE` and `stderr=PIPE` getting "full" in a weird interaction - FFmpeg was stalling for more data in stdin to complete a frame (backlogging chunks) and attempting to write in stderr; but the pipes bufsize seems to be shared, so it blocked itself and everyone from read/writes
The solution was simple, instead of relying on `PIPE` or removing it, we just send a new file for stderr/stdout as so:
```python
from tempfile import TemporaryFile
@define
class ExportingHelper:
def popen(self) -> None:
self.stderr = TemporaryFile(mode="r+b")
self.stdout = TemporaryFile(mode="r+b")
self.process = self.ffmpeg.popen(stdin=PIPE,
stdout=self.stdout, stderr=self.stderr)
```
Then, only `stdin` uses `os.pipe()` required for IPC without named pipes, while others are now independent without limits. Reading goes `self.stderr.seek(0)` followed by `self.stderr.read()`, so I get my logs or in-memory renders 🙂
Let me know if it works, maybe a `NamedTemporaryFile` is needed, but I can't reproduce the halt anymore locally