Hello friends I need some help I'm running rancher...
# vsphere
c
Hello friends I need some help I'm running rancher v2.14.2 and vSphere 7 I am able to deploy a cluster. Rancher creates the VMs on vSphere but the cluster is stuck in "Updating" and the nodes stuck in "Reconciling". I can ping each node between them and also the rancher nodes Provisioning Log keeps showing
Copy code
[INFO ] configuring bootstrap node(s) rke-lab-internal-pool1-hnhtg-jfr67: Waiting for Cluster control plane to be initialized, waiting for probes: calico, etcd, kube-apiserver, kube-controller-manager, kube-scheduler
[INFO ] configuring bootstrap node(s) rke-lab-internal-pool1-hnhtg-jfr67: Waiting for Cluster control plane to be initialized, waiting for probes: calico, kube-apiserver, kube-controller-manager
[INFO ] configuring bootstrap node(s) rke-lab-internal-pool1-hnhtg-jfr67: Waiting for Cluster control plane to be initialized, waiting for probes: calico
from within the first node
Copy code
root@rke-lab-internal-pool1-hnhtg-jfr67:~# kubectl get nodes
NAME                                 STATUS   ROLES                AGE   VERSION
rke-lab-internal-pool1-hnhtg-jfr67   Ready    control-plane,etcd   19m   v1.35.6+rke2r1
rke-server.service is running
Copy code
systemctl status rke2-server.service --no-pager -l
● rke2-server.service - Rancher Kubernetes Engine v2 (server)
     Loaded: loaded (/usr/local/lib/systemd/system/rke2-server.service; enabled; preset: enabled)
     Active: active (running) since Fri 2026-07-03 18:34:53 UTC; 2min 37s ago
 Invocation: 55d3f198141f4d09ae4ff91283074d6b
       Docs: <https://github.com/rancher/rke2#readme>
   Main PID: 1642 (rke2)
      Tasks: 227
     Memory: 637.9M (peak: 724.2M)
        CPU: 2min 17.114s
     CGroup: /system.slice/rke2-server.service
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Connected to proxy" url="<wss://172.25.150.26:9345/v1-rke2/connect>"
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Remotedialer connected to proxy" url="<wss://172.25.150.26:9345/v1-rke2/connect>"
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applied manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-snapshot-controller.yaml\"" object=kube-system/rke2-snapshot-controller reason=AppliedManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applying manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-snapshot-validation-webhook.yaml\"" object=kube-system/rke2-snapshot-validation-webhook reason=ApplyingManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applied manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-snapshot-validation-webhook.yaml\"" object=kube-system/rke2-snapshot-validation-webhook reason=AppliedManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applying manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-traefik-crd.yaml\"" object=kube-system/rke2-traefik-crd reason=ApplyingManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applied manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-traefik-crd.yaml\"" object=kube-system/rke2-traefik-crd reason=AppliedManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applying manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-traefik.yaml\"" object=kube-system/rke2-traefik reason=ApplyingManifest type=Normal
Jul 03 18:34:55 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:55Z" level=info msg="Event occurred" apiVersion=<http://k3s.cattle.io/v1|k3s.cattle.io/v1> fieldPath= kind=Addon logger=rke2/deploy message="Applied manifest at \"/var/lib/rancher/rke2/server/manifests/rke2-traefik.yaml\"" object=kube-system/rke2-traefik reason=AppliedManifest type=Normal
Jul 03 18:34:56 rke-lab-internal-pool1-hnhtg-jfr67 rke2[1642]: time="2026-07-03T18:34:56Z" level=info msg="Event occurred" apiVersion= fieldPath= kind=Node logger=rke2 message="Deferred node password secret validation complete" object=rke-lab-internal-pool1-hnhtg-jfr67 reason=NodePasswordValidationComplete type=Normal
What am I missing here? 😔
OK, some new info looks like the issue is with
deployment.apps/calico-typha
Copy code
2026-07-03 18:54:08.350 [ERROR][1] typha/health.go 389: Health endpoint failed, trying to restart it... error=listen tcp: lookup localhost on 192.168.27.8:53: no such host
192.168.27.8 is my network DNS server why is calico-typha it searching for
localhost
in there?!
I just added
"cluster-dns": "127.0.0.1",
to Networking to no avail One important information I'm getting the IP Address from a "Network Protocol Profile" on vsphere and using this cloud-config to apply it gist.github.com/gschanuel/4817b239bc3ef163f2ee011d575bd464
b
This was one you solved in the "general" tab?