This message was deleted.
# harvester
a
This message was deleted.
g
The solution is to manually unbind the nvme driver:
Copy code
for d in 0000:c2:00.0 0000:c3:00.0 0000:c4:00.0 0000:c5:00.0; do
  echo $d > /sys/bus/pci/drivers/nvme/unbind
done
and restart instance manager
The errors I have: •
VFIO group is not viable! Not all devices in IOMMU group bound to VFIO or unbound
•
Driver cannot attach the device
•
No controller was found with provided trid
• then Longhorn fails with
No such device
for
0000:c2:00.0
,
0000:c3:00.0
,
0000:c4:00.0
, and
0000:c5:00.0
IOMMU group is obviously correct
Have anyone else encountered this problem? I know Longhorn v2 is going to be GA soon
l
@salmon-doctor-9726
g
would it be possible for us to have a support bundle?
g
I've rebooted the node on staging environment and the problem is there too. Even though there are different servers, regions, disks.
I think I can do it, were should I send you the bundle?
@great-bear-19718
I am doing it now, I'll DM you the bundle.
Again, running:
for d in 0000:c3:00.0 0000:c4:00.0; do echo $d > /sys/bus/pci/drivers/nvme/unbind; done
and restarting instance-manager fixes the issue
👍 1
g
how did you add these disks?
👍 1
they are not added via blockdevices interface
Copy code
k get blockdevices -A
NAMESPACE         NAME                               TYPE   DEVPATH        MOUNTPOINT   NODENAME      PROVISIONPHASE   AGE
longhorn-system   0265fb713e13d847e807b759f12f53eb   disk   /dev/nvme1n1   null         harvester-0   Unprovisioned    6d7h
longhorn-system   0db1f7f502b9f8e5e5c8c063da7d6b90   disk   /dev/nvme1n1   null         harvester-1   Unprovisioned    6d7h
longhorn-system   273d6daf065cc400bc89c46bda4e1dfe   disk   /dev/nvme0n1   null         harvester-1   Unprovisioned    6d7h
longhorn-system   2ef8424a3200c8735232978406005568   disk   /dev/nvme1n1   null         harvester-2   Unprovisioned    6d7h
longhorn-system   cb03e2c390a6553fc14d349338f6950e   disk   /dev/nvme0n1   null         harvester-2   Unprovisioned    6d7h
longhorn-system   eb08e12bd77132b851a4db7ebb8916cb   disk   /dev/nvme0n1   null         harvester-0   Unprovisioned    6d7h
g
The devices are not available in the UI, therefore I am adding them directly longhorn resources:
Copy code
apiVersion: <http://longhorn.io/v1beta2|longhorn.io/v1beta2>
kind: Node
This one:
Copy code
apiVersion: longhorn.io/v1beta2
kind: Node
metadata:
  name: harvester-0
  namespace: longhorn-system
spec:
  allowScheduling: true
  disks:
    default-disk-e9cb6bd301feed3b:
      allowScheduling: true
      diskDriver: ''
      diskType: filesystem
      evictionRequested: false
      path: /var/lib/harvester/defaultdisk
      storageReserved: 0
      tags: []
    spdk-c1:
      allowScheduling: true
      diskDriver: nvme
      diskType: block
      evictionRequested: false
      path: 0000:c3:00.0
      storageReserved: 0
      tags: []
    spdk-c2:
      allowScheduling: true
      diskDriver: nvme
      diskType: block
      evictionRequested: false
      path: 0000:c4:00.0
      storageReserved: 0
      tags: []
  evictionRequested: false
  instanceManagerCPURequest: 0
  name: harvester-0
  tags: []
I don't have them provisioned using blockdevices, but after patching, they appear in the list of nodes and work
should I change the provisioning method?
g
yes please, that is the supported method
this will ensure changes survive reboots
g
Like I mentioned, the disks are not present on the 1.7.1 Harvester in the UI: https://docs.harvesterhci.io/v1.7/host/#multi-disk-management
Is there a resource I can patch to have a similar effect? Or an endpoint I can try to use?
After upgrading to 1.8.0 and testing it using this method: https://docs.harvesterhci.io/v1.7/advanced/longhorn-v2 The result is the same. I need to manually unbind the nvme driver, becuase SPDK fails to do it. The error:
Copy code
EAL: 0000:c4:00.0 VFIO group is not viable! Not all devices in IOMMU group bound to VFIO or unbound
And latter:
Copy code
[longhorn-instance-manager] time="2026-05-07T12:31:00.519299507Z" level=warning msg="Unbinding NVMe disk 0000:c4:00.0 since failed to attach" func="nvme.(*DiskDriverNvme).DiskCreate.func1" file="nvme.go:49" error="error sending message, id 5866, method bdev_nvme_attach_controller, params {Name:ccf428fb-1730-43ab-9dc1-4897902a30ab NvmeTransportID:{Trtype:PCIe Adrfam: Traddr:0000:c4:00.0 Trsvcid: Subnqn:} Hostaddr: Hostsvcid: CtrlrLossTimeoutSec:30 ReconnectDelaySec:2 FastIOFailTimeoutSec:15 Multipath:disable}: {\"code\": -19,\"message\": \"No such device\"}"