We have a local cluster running k3s and a number o...
# general
b
We have a local cluster running k3s and a number of RKE2 clusters, created with it, which it controls. When going to the Cluster Dashboard of a cluster and looking at the Capacity: Pods, the suggested total is not always in keeping with the number of workers x 110, which we left as the default. It can be some multiple of 110 greater than it should be. Any ideas what might be causing it or how to settle it back down to reporting the correct maximum?
b
We changed this on our nodes intentionally and I'm trying to remember the exact setting we set.
There might have been multiple ones, but essentially it came down to IP pools.
b
It seems there's a mismatch with what is reported in the managed cluster and what the local cluster is expecting. For one cluster with 3 control planes and 3 workers, the total is show as 440. When I run kubectl get nodes.management.cattle.io -n c-m-w2z77sbr -o wide on local, I get 7 machines listed. I've got to persuade local to lay some machines to rest. Expired. No more, Have ceased to be. The managed clusters know they've gone, local keeps hanging on.
b
Well that sounds different from what I expected.
You have autoscaling turned on? How are these provisioned? Harvester? GCP? Custom? Elemental?
b
Manually, through vSphere.
b
So they're custom clusters? (I know there's a vSphere provider too iirc)
b
The machines are built on vSphere, The clusters are built through Rancher, then the machines are registered.
b
Right, but you can do that with the elemental and/or vSphere Providers.
b
All manually done - DHCP wasn't on our system and it seemed that was required to build the machines themselves through Rancher.
b
image.png
So it can say "Imported" or "Custom"
Under Provider
on your cluster View
in Rancher
b
It's custom. The cluster is built through Rancher using Terraform. After that, the machines are built through vSphere, again using Terraform. They're entirely pure and innocent at the time with neither side knowing anything of the other. Then the curl for registration is copied from Rancher and applied to each of the servers, some with control plane/etcd, some with worker. After that, we have a working RKE2 cluster.
b
Yeah that's cool, I think technically you can do your basic workflow with Custom or Imported, but the some of the mechanisms are different, which might play a part here.
So it sounds like the nodes.management.cattle.io doesn't match downstream. You're running the cattle.io nodes against Rancher's kubectl right? What does
kubectl get nodes
show on the control plane of the downstream cluster?
b
It shows the nodes we're expecting, just the 6 of them (where there are 6).
b
And what happens when you delete the cattle.io object from upstream?
(I'd also check
kubectl get machines -A
and look for stale entries in the Rancher cluster)
b
I've never tried: what has happened is that we've taken nodes out by draining them and removing them from the cluster via Cluster Management, rebuilt them and run the registration curl on them anew. Pardon my ignorance of deleting cattle.io.
b
You shouldn't have to
But if they're stale or stuck somehow, we'll need to clean them up to get the right numbers
You might end up having to patch the finalizers (delete them) to get them to actually delete
And sorry I short handed
<http://cattle.io|cattle.io>
from
<http://nodes.management.cattle.io|nodes.management.cattle.io>
I've had stale objects get stuck in the past, but never really anything I could re-produce for a bug report.
b
No stale entries, all listed as running. I've got 69 machines listed and... 79 in the cluster servers. I so love maths.
b
kubectl get <http://nodes.management.cattle.io|nodes.management.cattle.io> -n c-m-w2z77sbr
shows 7 nodes instead of 6 right?
b
Thanks for all the help. I'll see if I can check on the machine names compared to what is actually in place
b
I thought that's what you mentioned.
b
Yes, that's it.
b
Yeah just find the stale one and delete it as a next step I think
kubectl delete <http://nodes.management.cattle.io|nodes.management.cattle.io> <badnode> -n c-m-w2z77sbr
b
Yes, I'm expecting that to be the right way of doing it. I just have to make sure it's the right machine or it could be embarrasing.
b
Yep!
b
Thanks again.
b
good luck!
b
A finale on this:
kubectl get <http://nodes.management.cattle.io|nodes.management.cattle.io> -n <cluster-id> -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.nodeName}{"\n"}{end}' | sort -k 2
will show if you have more than one machine name per node name. They will have different UIDs, but I couldn't find a way to work out which one was active. Instead, I deleted any node that had more than one machine name, rebuilt it and re-registered it. Maybe not sleek and refined, but it worked.
🦜 1
🎉 1
🐿️ 1