Hi everyone, I'd like to share here the incident r...
# harvester
h
Hi everyone, I'd like to share here the incident report regarding a massive failure that I had last night, maybe someone is experiencing the same scenario, happening on version 1.7.1
👀 1
🐿️ 1
b
Please share the improvements you did to the deployment to solve this problem.
h
no improvements, still no idea why the 3 controlplane were unresponding at the same time
b
We're not using longhorn v2, but we've had similar situations when we lost layer3 or layer4 networking between nodes. Since this is specifically related to the #CC2UQM49Y part of the stack, you might consider cross posting there, specially because I think there's a lot of harvester folks that are still running the v1 engine defaults here.
b
Imho this looks like your nodes were under heavy memory pressure. Did you see more oom-kills, e.g. for virt-manager pods or qemu processes? Longhorn pods should be killed last, since they have (or should have) a priority class set. OTOH Longhorn v2 is not GA yet, so imho you're doing beta tests for Suse here.