Comments (9)
from kserve.
I've been thinking about the model storage problem. Re-downloading from S3/GCS is somewhat expensive. It would be kind of nice to shared a shared model cache that lives somewhere, maybe as a shared mountable volume.
We could then have some process for loading the model to some local volume and then pointing for all model servers to read from there. We could encapsulate TensorRT details onto that volume. Thoughts?
from kserve.
Looks like there is complication of syncing from S3/GCS model storage, TensorRT does polling and then adds/removes the model versions(https://docs.nvidia.com/deeplearning/sdk/tensorrt-inference-server-master-branch-guide/docs/model_repository.html#modifying-the-model-repository).
from kserve.
from kserve.
Definitely agreed, we can disable this at this stage, however there is real prod use case for continuous training and active learning where this feature can be useful and manual canary is not an option, the rollout risk is managed somewhere else in the train-serve loop.
from kserve.
I was thinking of downloading in an init container and exposing a mount to the tensorRT container, however as noted in #129 we need to figure out how to enhance or workaround knative here...
from kserve.
Why do you need to expose a mount. If you download in the init container, the pod shares a disk, right?
from kserve.
pod shares a disk, right
They may share the disk on the node if that is what you mean, but each container has its own filesystem so you would still need mount support
from kserve.
closed via #148
from kserve.
Related Issues (20)
- Add Oracle Cloud Infrastructure (OCI) Object Storage as a storage agent
- VirtualService regex match should be case insensitive
- Python SDK for KServe and Kubeflow Pipelines can not be installed at the same time HOT 2
- UnicodeDecodeError for grpcurl request with Bytes column in DataFrame HOT 7
- error setting up interface service HOT 2
- mlflow model cannot be loaded HOT 8
- stop using `gcr.io/kubebuilder/kube-rbac-proxy` before `18 March 2025` (image being deleted) HOT 1
- add Xinfernece ( an inference platform which integrated transformers, vllm, and llama.cpp as engines,) runtime for LLM Serving Runtime HOT 5
- Completion fails when echo is true with vLLM backend
- protobuf version conflict while trying to integrate with kfp HOT 2
- Client fails to list clusterservingruntimes HOT 2
- Not able to access torchserve custom metrics after deploying inference service on kserve
- The request to InferenceService is sent twice
- Getting timeout failed to failed to call webhook: Post "https://kserve-webhook-server-service.default.svc:443/mutate-serving-kserve-io-v1beta1-inferenceservice?timeout=10s" HOT 7
- Multi-Lora support
- fake client returns no kind "ClusterServingRuntimeList" is registered for version "serving/v1alpha1" HOT 1
- duplicated hosts error after configuring the additional domains HOT 1
- Document missing content
- Can't find version compatibility matrix for KServe HOT 4
- Make inference to models using Istio and keycloak avoiding session cookies HOT 2
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from kserve.