This message was deleted.
# harvester
a
This message was deleted.
b
Namely the annotations on the nodes were wrong/off.
Copy code
$ kubectl describe node gpu2 | grep -i nvidia
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>:  0
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>:  0
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>  2                  2
There were 4 cards on the node, but it was only registering 2 of them.
I had to remove all the duplicate entries from the
permittedHostDevices
section (
kubectl edit kubevirt kubevirt -n harvester-system
) then restart rke2 on the node... (
systemctl restart rke2-agent || systemctl restart rke2-server
) then kick the pci pod (
kubectl delete pod $(kubectl get pods -n harvester-system -o wide | grep gpu2 | grep pci |awk '{print $1}') -n harvester-system
) to get it to reflect properly:
Copy code
$ kubectl describe node gpu2 | grep -i nvidia
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>:  4
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>:  4
  <http://nvidia.com/GV100GL_TESLA_V100_SXM2_32GB|nvidia.com/GV100GL_TESLA_V100_SXM2_32GB>  3
I'm not sure if I should open another bug because of the cascade efffect or just expand the existing one.
Eh. I tacked it onto the existing bug