narrow-guitar-87575
07/16/2026, 8:52 AMetcd-snapshot-schedule-cron: '20 */6 * * *'
is there a way to set, that snapshot shall not be done on the leader?
ratio: at the time of snapshot, the etcd cluster is unavailable for a short time and some leader elections fail. (e.g., rancher leader election).boundless-salesclerk-29005
07/16/2026, 9:11 AMboundless-salesclerk-29005
07/16/2026, 9:12 AMnarrow-guitar-87575
07/16/2026, 10:23 AM2026-07-11T12:00:35.244208163+02:00 stderr F E0711 10:00:35.242218 43 leaderelection.go:445] "Failed to update lease optimistically, falling back to slow path" err="Put \"<https://1>
0.43.0.1:443/apis/coordination.k8s.io/v1/namespaces/kube-system/leases/cattle-controllers?timeout=15m0s\": context deadline exceeded" lock="kube-system/cattle-controllers"
2026-07-11T12:00:35.24448038+02:00 stderr F E0711 10:00:35.242289 43 leaderelection.go:452] "Error retrieving lease lock" err="context deadline exceeded" lock="kube-system/cattle-controllers"
2026-07-11T12:00:35.244491511+02:00 stderr F I0711 10:00:35.242309 43 leaderelection.go:299] "Failed to renew lease" lock="kube-system/cattle-controllers" err="context deadline exceeded"
it happens exactly in time of etcd doing snapshot. Snap size is about 500MB.boundless-salesclerk-29005
07/16/2026, 11:16 AMetcd-db-metric-poll-interval and etcd-compaction-interval but I wonder if that will have any impact at all.
There are also election related defaults you can increase in kube-controller-manager: kubernetes.io/docs/reference/command-line-tools-reference/kube-controller-manager
But I would first confirm there are no frequent leader changes or lease related problems due to disk/network latency, CPU or memory pressure or frequent restarts (evictions, crashes)narrow-guitar-87575
07/16/2026, 11:26 AM2026-07-11T11:55:51.446420274+02:00 stderr F {"level":"info","ts":"2026-07-11T09:55:51.446247Z","caller":"mvcc/hash.go:157","msg":"storing new hash","hash":2268017395,"revision":1971370137,"compact-revision":1971364818}
2026-07-11T11:59:44.103476757+02:00 stderr F {"level":"info","ts":"2026-07-11T09:59:44.103089Z","caller":"etcdserver/server.go:2217","msg":"triggering snapshot","local-member-id":"ef12a22198b381bf","local-member-applied-index":2183595556,"local-member-snapshot-index":2183585555,"local-member-snapshot-count":10000,"snapshot-forced":false}
2026-07-11T11:59:44.106370649+02:00 stderr F {"level":"info","ts":"2026-07-11T09:59:44.106225Z","caller":"etcdserver/server.go:2262","msg":"saved snapshot to disk","snapshot-index":2183595556}
2026-07-11T12:00:02.911064579+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:02.910806Z","caller":"fileutil/purge.go:88","msg":"purged","path":"/var/lib/rancher/rke2/server/db/etcd/member/snap/0000000000000086-00000000822642ce.snap"}
2026-07-11T12:00:42.21567136+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:42.215454Z","caller":"wal/wal.go:825","msg":"created a new WAL segment","path":"/var/lib/rancher/rke2/server/db/etcd/member/wal/0000000000014b55-00000000822710c3.wal"}
2026-07-11T12:00:50.523398813+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:50.523128Z","caller":"mvcc/index.go:194","msg":"compact tree index","revision":1971375438}
2026-07-11T12:00:51.534562147+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:51.534399Z","caller":"mvcc/kvstore_compaction.go:70","msg":"finished scheduled compaction","compact-revision":1971375438,"took":"914.428521ms","hash":325786831,"current-db-size-bytes":506068992,"current-db-size":"506 MB","current-db-size-in-use-bytes":230494208,"current-db-size-in-use":"230 MB"}
2026-07-11T12:00:51.534598906+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:51.534481Z","caller":"mvcc/hash.go:157","msg":"storing new hash","hash":325786831,"revision":1971375438,"compact-revision":1971370137}
2026-07-11T12:00:59.725137219+02:00 stderr F {"level":"info","ts":"2026-07-11T10:00:59.724944Z","caller":"fileutil/purge.go:88","msg":"purged","path":"/var/lib/rancher/rke2/server/db/etcd/member/wal/0000000000014b50-000000008225c0a6.wal"}
2026-07-11T12:01:13.57256393+02:00 stderr F {"level":"info","ts":"2026-07-11T10:01:13.572327Z","caller":"etcdserver/server.go:2217","msg":"triggering snapshot","local-member-id":"ef12a22198b381bf","local-member-applied-index":2183605557,"local-member-snapshot-index":2183595556,"local-member-snapshot-count":10000,"snapshot-forced":false}
2026-07-11T12:01:13.576075152+02:00 stderr F {"level":"info","ts":"2026-07-11T10:01:13.575918Z","caller":"etcdserver/server.go:2262","msg":"saved snapshot to disk","snapshot-index":2183605557}
2026-07-11T12:01:32.912669563+02:00 stderr F {"level":"info","ts":"2026-07-11T10:01:32.912432Z","caller":"fileutil/purge.go:88","msg":"purged","path":"/var/lib/rancher/rke2/server/db/etcd/member/snap/0000000000000086-00000000822669df.snap"}
2026-07-11T12:04:42.641566805+02:00 stderr F {"level":"info","ts":"2026-07-11T10:04:42.641340Z","caller":"wal/wal.go:825","msg":"created a new WAL segment","path":"/var/lib/rancher/rke2/server/db/etcd/member/wal/0000000000014b56-0000000082273fb3.wal"}
2026-07-11T12:04:59.767781681+02:00 stderr F {"level":"info","ts":"2026-07-11T10:04:59.767572Z","caller":"fileutil/purge.go:88","msg":"purged","path":"/var/lib/rancher/rke2/server/db/etcd/member/wal/0000000000014b51-00000000822603f7.wal"}
2026-07-11T12:05:50.529221937+02:00 stderr F {"level":"info","ts":"2026-07-11T10:05:50.528973Z","caller":"mvcc/index.go:194","msg":"compact tree index","revision":1971382954}
2026-07-11T12:05:51.513946107+02:00 stderr F {"level":"info","ts":"2026-07-11T10:05:51.512740Z","caller":"mvcc/kvstore_compaction.go:70","msg":"finished scheduled compaction","compact-revision":1971382954,"took":"940.708031ms","hash":1607211870,"current-db-size-bytes":506068992,"current-db-size":"506 MB","current-db-size-in-use-bytes":259751936,"current-db-size-in-use":"260 MB"}
2026-07-11T12:05:51.514014685+02:00 stderr F {"level":"info","ts":"2026-07-11T10:05:51.512837Z","caller":"mvcc/hash.go:157","msg":"storing new hash","hash":1607211870,"revision":1971382954,"compact-revision":1971375438}boundless-salesclerk-29005
07/16/2026, 11:42 AMnarrow-guitar-87575
07/16/2026, 11:47 AMboundless-salesclerk-29005
07/16/2026, 11:49 AMboundless-salesclerk-29005
07/16/2026, 11:52 AMnarrow-guitar-87575
07/16/2026, 11:52 AMcreamy-pencil-82913
07/16/2026, 5:00 PMcreamy-pencil-82913
07/16/2026, 5:01 PMnarrow-guitar-87575
07/16/2026, 5:07 PMkubectl get pods -n cattle-system
NAME READY STATUS RESTARTS AGE
rancher-ff7f7d8b4-4cftm 2/2 Running 1 (6d14h ago) 8d
rancher-ff7f7d8b4-7zrdv 2/2 Running 1 (5d7h ago) 7d3h
rancher-ff7f7d8b4-fz7h8 2/2 Running 1 (2d8h ago) 8d
all randomly, but the restart time always matches etcd snapshot time on some nodenarrow-guitar-87575
07/16/2026, 10:56 PMcreamy-pencil-82913
07/17/2026, 12:13 AMiostat -d -x 1 running and then trigger a manual snapshot with rke2 etcd-snapshot save . I suspect your disk is struggling more than you expect.creamy-pencil-82913
07/17/2026, 12:14 AMnarrow-guitar-87575
07/17/2026, 8:33 AMDevice r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util
md0 0.00 0.00 0.00 0.00 0.00 0.00 4.00 16.00 0.00 0.00 0.00 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
md1 2.00 8.00 0.00 0.00 0.00 4.00 41.00 176.00 0.00 0.00 0.00 4.29 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 9.60
md124 0.00 0.00 0.00 0.00 0.00 0.00 239.00 237056.00 0.00 0.00 224.17 991.87 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 53.58 48.00
md125 0.00 0.00 0.00 0.00 0.00 0.00 58.00 264.00 0.00 0.00 6.28 4.55 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.36 15.20
md126 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
md127 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
sda 0.00 0.00 0.00 0.00 0.00 0.00 300.00 237326.50 5.00 1.64 26.01 791.09 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 7.80 46.80
sdb 0.00 0.00 0.00 0.00 0.00 0.00 306.00 240626.50 4.00 1.29 23.79 786.36 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 7.28 46.20
sdc 2.00 8.00 0.00 0.00 0.00 4.00 33.00 157.50 6.00 15.38 0.09 4.77 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
sdd 0.00 0.00 0.00 0.00 0.00 0.00 8.00 53.50 8.00 50.00 0.12 6.69 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
and only 2 prints, i.e., all below 3 seconds which should be within all timeouts.creamy-pencil-82913
07/17/2026, 8:36 AMnarrow-guitar-87575
07/17/2026, 8:39 AM"apply request took too long","took":"162.399707ms","expected-duration":"100ms"
so just slightly above threshold.