Hi,, Restoring a Cluster when all master are down ...
# rke2
q
Hi,, Restoring a Cluster when all master are down Cluster was having unhealthy master hence removed all master and started adding master as per docs. I followed all the restore steps and the cluster is healthy now. However, the restored cluster doesn't contain any application resources - all app namespaces, PVs, and workloads are missing. I selected the correct etcd snapshot. Should an etcd restore also recover application namespaces/PVs, or does it restore only the cluster infrastructure? ranchermanager.docs.rancher.com/how-to-guides/…/restore-rancher-launched-kubernetes-clusters-from-backup#…
c
etcd snapshot contains everything that existed in the cluster at that point in time.
Are you sure you followed the steps correctly?
Also, I am curious why your first response to having a single unhealthy master node is to delete all the nodes and start over from a snapshot… that seems like a drastic response.
q
yes i am attempting to recreate a issue which occurred in prod. Due to a human error, by mistake all master nodepools were deleted and we were in a situation on how to recover the cluster including master. Hence following the steps shared earlier by you. I am sure i followed same steps as per docs. Is it fine to share the logs/snippets here?
c
sure
q
Hi @creamy-pencil-82913 - Attached list of steps executed & master logs. I have executed same set of steps N number of time however every time i endup with the checksum error and then i had apply the tweak as mentioned in doc, not sure if that is creating a conflict. Rest all steps remains are per rancher doc Please review and help.
@creamy-pencil-82913 - Not urgent.. when you have some time, please review and share your suggestions.
c
you made the snapshot folder on the node and then triggered the restore… but you didn’t actually put the snapshot in that folder? And then you patched the secret to ignore the fact that it was failing to restore because you didn’t actually give it a snapshot to restore.
Ideally you would keep your snapshots on S3, so that it can just pull the snapshot from the bucket. If you are working with only local snapshots, and delete all the nodes, you need to put a snapshot file back on the node by hand. You seem to have missed that part entirely - you make the snapshot dir where you WOULD put the snapshot to be restored… but then don’t do that. And then when it fails to restore (because the file isnt there) you just hack the plan secret to skip restoring. So you end up with a new cluster, instead of a restored snapshot.
q
Thanks @creamy-pencil-82913 for your support as always! So far in all my test-case i was testing only using S3 and not with local snapshot restore. By checking attached logs which i shared - "master-vm-logs.txt", i have found the cause for the checksum error & worker vm reconciliation issue. Have closed this - github.com/rancher/rancher/issues/54774 Below is the error I encountered. After triggering the S3 snapshot restore from the UI, the S3 endpoint CA certificate was not created automatically under /var/lib/rancher/rke2/etc/config-files/ causing the restore to fail with a checksum/S3 client initialization error. As a workaround, I manually created the S3 endpoint CA certificate before initiating the restore. The S3 restore then completed successfully and restored the cluster without any issues. Same i tried for Local ETCD snapshot restore was well, worked perfectly! //// Jul 16 131517 carbon-sadc-sbx-postgres-v3-np1-w1c2-m-5c9c9-xjxgr rancher-system-agent[5658]: time="2026-07-16T131517-04:00" level=info msg="[c9c94de7455b36cb858845cbfbb6d70bc39c87fec6ac47dfc8a88cac484c6f7c_4stderr] time=\"2026-07-16T131517-04:00\" level=info msg=\*"Attempting to create new S3 client for endpoint=\*\\"bkpcohprds.com\\\" bucket=\\\"etcd-backup-sbx\\\" folder=\\\"rke2-sadc-sbx-p> Jul 16 131517 carbon-sadc-sbx-postgres-v3-np1-w1c2-m-5c9c9-xjxgr rancher-system-agent[5658]: time="2026-07-16T131517-04:00" level=info msg="[c9c94de7455b36cb858845cbfbb6d70bc39c87fec6ac47dfc8a88cac484c6f7c_4stderr] time=\"2026-07-16T131517-04:00\" level=info msg=\"Shutdown request received\"" Jul 16 131517 carbon-sadc-sbx-postgres-v3-np1-w1c2-m-5c9c9-xjxgr rancher-system-agent[5658]: time="2026-07-16T131517-04:00" level=info msg="[c9c94de7455b36cb858845cbfbb6d70bc39c87fec6ac47dfc8a88cac484c6f7c_4stderr] time=\"2026-07-16T131517-04:00\" level=fatal msg=\"Error: starting kubernetes: failed to start cluster: start managed database: failed to initialize S3 client: open /var/lib/rancher/rke2/etc/config-files/>
c
How did you configure that custom ca cert? In the rancher UI?
q
yes. 1. At global level - Rancher UI Cloud Creds - adding s3 secrets with certs - attached screenshot 2. Another area is, when i provision each downstream cluster using helm chart i add under etcd-> S3 certs as shown below. Another screenshot attached is from downstream rancher cluster which shows etcd section with certs
Copy code
etcd:
      s3:
        bucket: etcd-backup-sbx
        cloudCredentialName: cattle-global-data:s3cohesitycredentials
        endpoint: bkpcohprdsasysctrl.com:3000
        endpointCA: |-
          -----BEGIN CERTIFICATE-----
          MIIHfDCCBWSgAwIBAgITSAAAACMcmL6NV4VFEgAAAAAAIzANBgkqhkiG9w0BAQsF
          ADAZMRcwFQYDVQQDEw5NU0NBUk9PVC1MT1dFUzAeFw0yMzA0MjcxMzM4NDNaFw0z
          MTA0MjIxODAwMTJaMEYxEzARBgoJkiaJk/IsZAEZFgNjb20xFTATBgoJkiaJk/Is
          lAGfXbkbMmpzDl1PBEt+pA35UEmn01rgvnh037dfabGCguWtYPlipFrmcfcMn86X
          -----END CERTIFICATE-----
          -----BEGIN CERTIFICATE-----
          MIIFDTCCAvWgAwIBAgIQUE41gd1Bi4NOni4fBSNDojANBgkqhkiG9w0BAQsFADAZ
          MRcwFQYDVQQDEw5NU0NBUk9PVC1MT1dFUzAeFw0xNTA0MjIxNzUwMTRaFw0zMTA0
          MjIxODAwMTJaMBkxFzAVBgNVBAMTDk1TQ0FST09ULUxPV0VTMIICIjANBgkqhkiG
          9w0BAQEFAAOCAg8AMIICCgKCAgEA749p7swlIcNR+kstAwZNOKEIRY8k5nAt1pFO
          tOW9S/NjO/2qt2Z/0UpF090vtCUe3iJoAtGPcQvpjRPaxP/P9MTNKTu7mPywDLGa
          9j2zBQ0vn0Q9P0qi9/P1sKr+saLAqSpVFDbTf73uOjMKD1OuYIaDfaISK8w1UKGV
          Cg==
          -----END CERTIFICATE-----
        folder: rke2-sadc-sbx-postgres-v4
        region: sbx
      snapshotRetention: 14
      snapshotScheduleCron: 0 */12 * * *
Once cluster is up, within VM - Rancher auto creates the 50-rancher.yaml & /var/lib/rancher/rke2/etc/config-files/ and place the cert. I don't specify which path to add cert anywhere in helm/config, looks rancher auto does. From VM [root@carbon-sadc-sbx-postgres-v4-np1-w1c1-m-8c5br-lslpx config.yaml.d]# pwd /etc/rancher/rke2/config.yaml.d [root@carbon-sadc-sbx-postgres-v4-np1-w1c1-m-8c5br-lslpx config.yaml.d]# cat 50-rancher.yaml | grep etcd "etcd-arg": [ "etcd-expose-metrics": true, "etcd-s3": true, "etcd-s3-access-key": "GPuoJRFwsuV", "etcd-s3-bucket": "etcd-backup-sbx", "etcd-s3-endpoint": "bkpcohprdsasysctrl.com:3000", "etcd-s3-endpoint-ca": "/var/lib/rancher/rke2/etc/config-files/s3-endpoint-ca-19131.crt", "etcd-s3-folder": "rke2-sadc-sbx-postgres-v4", "etcd-s3-region": "sbx", "etcd-s3-secret-key": "iI0wJ_Uw", "etcd-snapshot-retention": 14, "etcd-snapshot-schedule-cron": "0 */12 * * *", "node-role.kubernetes.io/etcd:NoExecute" [root@carbon-sadc-sbx-postgres-v4-np1-w1c1-m-8c5br-lslpx config.yaml.d]# ls -ltr /var/lib/rancher/rke2/etc/config-files/s3-endpoint-ca-19131.crt -rw-------. 1 root root 4467 Jul 20 10:54 /var/lib/rancher/rke2/etc/config-files/s3-endpoint-ca-19131.crt Half snippet of Cluster Spec from local UCM
Copy code
❯ k get cluster -n fleet-default carbon-sadc-sbx-postgres-v4 -o yaml
apiVersion: provisioning.cattle.io/v1
kind: Cluster
metadata:
  annotations:
    field.cattle.io/creatorId: u-6bionprqup
    meta.helm.sh/release-name: carbon-sadc-sbx-postgres-v4
    meta.helm.sh/release-namespace: fleet-default
    platform_name: postgres
    provisioning.cattle.io/management-cluster-display-name: carbon-sadc-sbx-postgres-v4
  creationTimestamp: "2026-07-20T14:46:00Z"
  finalizers:
  - wrangler.cattle.io/provisioning-cluster-remove
  - wrangler.cattle.io/rke-cluster-remove
  - wrangler.cattle.io/cloud-config-secret-remover
  generation: 3
  labels:
    app.kubernetes.io/managed-by: Helm
    environment: sbx
  name: carbon-sadc-sbx-postgres-v4
  namespace: fleet-default
  resourceVersion: "1187307836"
  uid: 3581b301-bfe3-412d-983d-da20811f8998
spec:
  cloudCredentialSecretName: cattle-global-data:vsphere-creds
  enableNetworkPolicy: false
  kubernetesVersion: v1.33.8+rke2r1
  localClusterAuthEndpoint:
    enabled: true
  rkeConfig:
    additionalManifest: |-
      ---
      apiVersion: helm.cattle.io/v1
      kind: HelmChartConfig
      metadata:
        name: rke2-calico
        namespace: kube-system
      spec:
        valuesContent: |-
          installation:
            calicoNetwork:
              mtu: 1450
    chartValues: null
    dataDirectories: {}
    etcd:
      s3:
        bucket: etcd-backup-sbx
        cloudCredentialName: cattle-global-data:s3cohesitycredentials
        endpoint: bkpcohprdsasysctrl.com:3000
        endpointCA: |-
          -----BEGIN CERTIFICATE-----
          MIIHfDCCBWSgAwIBAgITSAAAACMcmL6NV4VFEgAAAAAAIzANBgkqhkiG9w0BAQsF
          ADAZMRcwFQYDVQQDEw5NU0NBUk9PVC1MT1dFUzAeFw0yMzA0MjcxMzM4NDNaFw0z
          MTA0MjIxODAwMTJaMEYxEzARBgoJkiaJk/IsZAEZFgNjb20xFTATBgoJkiaJk/Is
          lAGfXbkbMmpzDl1PBEt+pA35UEmn01rgvnh037dfabGCguWtYPlipFrmcfcMn86X
          -----END CERTIFICATE-----
          -----BEGIN CERTIFICATE-----
          MIIFDTCCAvWgAwIBAgIQUE41gd1Bi4NOni4fBSNDojANBgkqhkiG9w0BAQsFADAZ
          MRcwFQYDVQQDEw5NU0NBUk9PVC1MT1dFUzAeFw0xNTA0MjIxNzUwMTRaFw0zMTA0
          MjIxODAwMTJaMBkxFzAVBgNVBAMTDk1TQ0FST09ULUxPV0VTMIICIjANBgkqhkiG
          9w0BAQEFAAOCAg8AMIICCgKCAgEA749p7swlIcNR+kstAwZNOKEIRY8k5nAt1pFO
          tOW9S/NjO/2qt2Z/0UpF090vtCUe3iJoAtGPcQvpjRPaxP/P9MTNKTu7mPywDLGa
          9j2zBQ0vn0Q9P0qi9/P1sKr+saLAqSpVFDbTf73uOjMKD1OuYIaDfaISK8w1UKGV
          Cg==
          -----END CERTIFICATE-----
        folder: rke2-sadc-sbx-postgres-v4
        region: sbx
      snapshotRetention: 14
      snapshotScheduleCron: 0 */12 * * *
    machineGlobalConfig:
      cluster-cidr: 192.168.0.0/17
      cni: calico
      disable:
      - rke2-ingress-nginx
      - rke2-snapshot-controller
      - rke2-snapshot-controller-crd
      - rke2-snapshot-validation-webhook
      etcd-arg:
      - quota-backend-bytes=5368709120
      - election-timeout=5000
      - heartbeat-interval=1000
      etcd-expose-metrics: true
      kube-apiserver-arg:
      - service-account-lookup=true
      - audit-log-path=/var/log/kubernetes/audit.log
      - audit-log-maxage=30
      - audit-log-maxbackup=10
      - audit-log-maxsize=100
      - enable-admission-plugins=NamespaceLifecycle,LimitRanger,ServiceAccount,DefaultStorageClass,DefaultTolerationSeconds,MutatingAdmissionWebhook,ValidatingAdmissionWebhook,ResourceQuota,NodeRestriction,Priority,TaintNodesByCondition,PersistentVolumeClaimResize,EventRateLimit,DenyServiceExternalIPs
      - admission-control-config-file=/etc/rancher/rke2/admission/admission-config.yaml
      - audit-policy-file=/etc/rancher/rke2/audit/audit-policy.yaml
      kube-apiserver-extra-mount:
      - /etc/rancher/rke2/admission:/etc/rancher/rke2/admission:ro
      - /etc/rancher/rke2/audit:/etc/rancher/rke2/audit:ro
      kube-controller-manager-arg:
      - node-cidr-mask-size-ipv4=26
      - bind-address=0.0.0.0
      kube-scheduler-arg:
      - bind-address=0.0.0.0
      kubelet-arg:
      - tls-cipher-suites=TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384,TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256,TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305,TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305
      service-cidr: 192.168.128.0/17
      system-default-registry: e-car-onprem-docker-virtual.docker.com
      write-kubeconfig-mode: 384