Hi <@U045RQTTBEC>, I’m currently working on the G...
# general
g
Hi @limited-pizza-33551, I’m currently working on the GPU View stack for Portainer. The goal is to consume DCGM Exporter metrics and surface the relevant GPU information and utilization metrics directly in the Portainer dashboard. Does Rancher have anything similar in terms of GPU monitoring/visibility that I could use as a reference for the dashboard design and the metrics they expose? I’ve already built a script that collects and reports the GPU stack details across the Kubernetes cluster. You can take a look here: github.com/ShubhamTatvamasi/gpu-operator/blob/…/gpu-cluster-report.sh Any references or examples from Rancher would be really helpful as I shape the GPU view. Thanks
😑 1
b
If you're using something like vllm all that's already exposed via Prometheus and hooks directly into rancher-monitoring but there's gotta be some pod/workload with access to the GPU. That's a fairly standard way to get metrics and isn't going to be rancher specific. The product isn't opinionated like that as I understand it.
I think you're better off looking at rancher as a collection of vanilla clusters. How would you tell which ones have GPUs?
g
Already figured it out. Thanks