Comments (11)
IIRC, there is no way to set job status to failed. Perhaps we could add StartingDeadlineSeconds(similar to ActiveDeadlineSeconds.) in runPolic, which specifies the time a job can remain Created. If the job is not running by the end of this period, it would be set to failed.
from training-operator.
@hywell-h Did you try using backoffLimit
?
from training-operator.
@hywell-h Did you try using
backoffLimit
?
when pod gose into ImagePullBackOff
, the status.phase
of pod is Pending
, and restart count of pod is zero;So backoffLimit
will not work, because pod has not reached the Completed
or Failed
state yet.
from training-operator.
@hywell-h Did you try using
backoffLimit
?when pod gose into
ImagePullBackOff
, thestatus.phase
of pod isPending
, and restart count of pod is zero;SobackoffLimit
will not work, because pod has not reached theCompleted
orFailed
state yet.
Oh, that makes sense.
Did you try using ActiveDeadlineSeconds
? But I understand the ActiveDeadlineSeconds
wouldn't be resolved the root cause.
To give failure handling more flexibility, I guess that the training-operator needs to migrate to the batch/v1 Job from plain pods management, and then support JobPodFailurePolicy.
from training-operator.
This sounds like the motivation behind kubernetes/kubernetes#122300
from training-operator.
This sounds like the motivation behind kubernetes/kubernetes#122300
@kannon92 Yeah, that's right, but as I remember correctly, configurable pod-stuck errors was rejected by SIG-Node, right?
from training-operator.
Did you try using
ActiveDeadlineSeconds
? But I understand theActiveDeadlineSeconds
wouldn't be resolved the root cause.
Yes, the pod status is always pending and it does not trigger the ActiveDeadlineSeconds
.
from training-operator.
@tenzen-y there was no rejection.
from training-operator.
@tenzen-y there was no rejection.
Ah, I see. As I recheck the issue, the next action seems to raise this topic in the SIG Node meeting.
We probably can move this forward in the next cycle since the current is under the enhancement freeze for the 1.30.
from training-operator.
Yea its mostly that I moved to Red Hat and I'm not as focused on batch related items so I don't really have the bandwidth to lead that one.
from training-operator.
Yea its mostly that I moved to Red Hat and I'm not as focused on batch related items so I don't really have the bandwidth to lead that one.
No worries. As I'm still there (WG Batch), let me check the enhancement in the next cycle if I can find enough time.
from training-operator.
Related Issues (20)
- Support MLX on Kubernetes with Kubeflow HOT 2
- Migrate to controller-runtime logger HOT 3
- Support CertManager for the Webhook cert generation HOT 1
- Unable to start elastic PyTorchJob example HOT 5
- Commonize webhook validations at the some points
- Update developer documentation for arm HOT 1
- Aunpun1.00 HOT 1
- Update pytorch launcher component in Kubeflow Pipelines repository HOT 2
- Update developer guide to handle missing training-operator-webhook-cert HOT 2
- Job Status is failed, when scale-in ps. HOT 4
- Failed K8s nodes leave jobs hanging indefinitely HOT 3
- Update examples for `train` API HOT 1
- [Question] Training Operator v1.8 Release Date HOT 1
- Why manifests/base/service.yaml does not include webhook server port (443) in version 1.7.0~1.5.0? HOT 7
- Not getting Kubeflow Training SDK v1.7 when installing `kubeflow-training` HOT 13
- Flaky Test: [It] should create desired Pods and Services: Distributed TFJob (4 workers, 2 PS) is succeeded
- MPIJob requires service names for the pods. HOT 3
- Add DeepSpeed Example with MPI Operator HOT 9
- chore(style): provide type for `STORAGE_INITIALIZER_VOLUME` constant
- fix(compatability): match-case syntax only compatible with Python3.10 HOT 5
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from training-operator.