Conductor workflow stalled after a sub-workflow

#3885 · open · 3 comments

View on GitHub ↗

rajeshwar-nu

Hi Team, I am experiencing this issue in latest version of conductor <https://github.com/Netflix/conductor/issues/3491> Stack 1. orkesio/orkes-conductor-community:1.1.11 2. Redis for workflow execution - <http://docker.io/bitnami/redis:7.0.8-debian-11-r0|docker.io/bitnami/redis:7.0.8-debian-11-r0> 3. Postgres for workflow persistence - <http://ghcr.io/cloudnative-pg/postgresql:15.3|ghcr.io/cloudnative-pg/postgresql:15.3> **Description of issue** A workflow get stuck in `RUNNING` state right after completion of a `SUBWORKFLOW`. This was observed in multiple workflows we have, all having `subworkflow`. The issue is erratic, it only happens for a few executions. I have attached 3 images for 3 sample failures ![workflow1 (1)](https://github.com/Netflix/conductor/assets/42337391/53548e09-2225-4879-8926-880a6084899e) ![workflow2 (1)](https://github.com/Netflix/conductor/assets/42337391/12eacdf2-d71c-475e-819c-c64216912e94) ![workflow3 (1)](https://github.com/Netflix/conductor/assets/42337391/26f6919f-bb2e-4d21-a7e9-0a26c6bd6049) The problem gets fixed when we `pause` and `resume` , after which it completes normally [Slack Message](https://orkes-conductor.slack.com/archives/C02KJ820XPW/p1701414993322349)

Comments

manan164

Hi @rajeshwar-nu , Are these subworkflow retried or restarted?

rajeshwar-nu

Hey @manan164 , no they are not.

appunni-old

@rajeshwar-nu do they have double underscore in the name ?