This message was deleted.
# harvester
a
This message was deleted.
🫠 1
☝️ 1
πŸ€” 1
b
Lemme log into a node and check
At least one of my cp nodes it looks fine
Interesting... so 2 of them are normal (almost no load) and the last node is at 100
I restarted the service on that node, but it's still pegged at 100
vs:
those are both
top -p $(systemctl status rke2-server.service  |grep PID |awk '{print $3}')
from different etcd nodes. I'm happy to help replicate or troubleshoot with you if there's a bug or ticket or something.
I don't really see anything going on via journald that seems to be a smoking gun:
Copy code
Feb 25 20:50:26 dh4 systemd[1]: Started Rancher Kubernetes Engine v2 (server).
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Reconciling ETCDSnapshotFile resources"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Starting dynamiclistener CN filter node controller with SANs: [10.16.2.150 127.0.0.1 ::1 localhost dh4 10.16.2.154 10.53.0.1 kubernetes kubernetes.default kubernetes.defaul>
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Starting managed etcd node metadata controller"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Tunnel server egress proxy mode: agent"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Reconciliation of ETCDSnapshotFile resources complete"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Reconciling ETCDSnapshotFile resources"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Starting <http://k3s.cattle.io/v1|k3s.cattle.io/v1>, Kind=Addon controller"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Creating deploy event broadcaster"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Starting /v1, Kind=Node controller"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Cluster dns configmap already exists"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Reconciliation of ETCDSnapshotFile resources complete"
Feb 25 20:50:26 dh4 rke2[1210221]: time="2026-02-25T20:50:26Z" level=info msg="Labels and annotations have been set successfully on node: dh4"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Started tunnel to 10.16.2.152:9345"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Started tunnel to 10.16.2.154:9345"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Started tunnel to 10.16.2.151:9345"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Stopped tunnel to 127.0.0.1:9345"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connecting to proxy" url="<wss://10.16.2.152:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Proxy done" err="context canceled" url="<wss://127.0.0.1:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connecting to proxy" url="<wss://10.16.2.154:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="error in remotedialer server [400]: websocket: close 1006 (abnormal closure): unexpected EOF"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connecting to proxy" url="<wss://10.16.2.151:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Handling backend connection request [dh4]"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connected to proxy" url="<wss://10.16.2.154:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Remotedialer connected to proxy" url="<wss://10.16.2.154:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connected to proxy" url="<wss://10.16.2.152:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Remotedialer connected to proxy" url="<wss://10.16.2.152:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Connected to proxy" url="<wss://10.16.2.151:9345/v1-rke2/connect>"
Feb 25 20:50:28 dh4 rke2[1210221]: time="2026-02-25T20:50:28Z" level=info msg="Remotedialer connected to proxy" url="<wss://10.16.2.151:9345/v1-rke2/connect>"
Feb 25 20:53:11 dh4 rke2[1210221]: time="2026-02-25T20:53:11Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:58520: write: connection reset by peer"
Feb 25 20:53:41 dh4 rke2[1210221]: time="2026-02-25T20:53:41Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:53800: write: connection reset by peer"
Feb 25 20:54:11 dh4 rke2[1210221]: time="2026-02-25T20:54:11Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:47896: write: connection reset by peer"
Feb 25 20:54:41 dh4 rke2[1210221]: time="2026-02-25T20:54:41Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:54208: write: connection reset by peer"
Feb 25 20:57:41 dh4 rke2[1210221]: time="2026-02-25T20:57:41Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:41282: write: connection reset by peer"
Feb 25 20:58:21 dh4 rke2[1210221]: time="2026-02-25T20:58:21Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:59944: write: connection reset by peer"
Feb 25 20:59:11 dh4 rke2[1210221]: time="2026-02-25T20:59:11Z" level=warning msg="Proxy error: write failed: write tcp 127.0.0.1:9345->127.0.0.1:50006: write: connection reset by peer"
Feb 25 21:00:28 dh4 rke2[1210221]: time="2026-02-25T21:00:28Z" level=info msg="Reconciling ETCDSnapshotFile resources"
Feb 25 21:00:28 dh4 rke2[1210221]: time="2026-02-25T21:00:28Z" level=info msg="Reconciliation of ETCDSnapshotFile resources complete"
p
yeah pegged at 100 is exactly what I'm seeing. very strange. I'm probably just gonna go back to v1.7 ... I did think it might be zfs stuff but wouldn't make sense really
ditto on
Proxy error: write failed
also btw
b
p
hah nice find!
b
You can thank my co-worker @miniature-salesclerk-33951
πŸ™Œ 1
🫑 1
You might want to check what version 1.7.0 ships as it could be affected too.
p
it seems to be just that version. we have a few 1.7.0 that are fine. I swapped out rke2 and cpu seems normal now. looks like 1,34.4 works too
πŸ™Œ 1
m
That's good to know. One of the issues said it would be patched later in February but I didn't see that explicitly in the release notes for it
πŸŽ‰ 1
b
We also stumbled upon this one bug, but in guest clusters. I didn't expect this in the Harvester image. Gosh. For the guest clusters we simply upgraded RKE2, but this is not really possible in Harvester. What did you do to work around it? It should just affect 1 CPU on the first RKE2 node, so maybe it's "acceptable"?
πŸ™ˆ 1