Skip to content

Build1 publisher3 min readPublished

qm start said nothing, the UI said stopped, and only host dmesg named the passthrough failure

A Proxmox 8 GPU passthrough that passed every checklist item still refused to boot. The only signal was an IRQ and vfio interrupt remapping burst in host dmesg, and the reported culprit was an HBA.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying qm start said nothing, the UI said stopped, and only host dmesg named the passthrough failure
Photo: nvidia.com

What happened

  • The plan was to PCIe passthrough an NVIDIA RTX A4000 from a Proxmox 8 host into an Ubuntu 24.04 VM, giving a Filecoin sealing worker direct GPU access for the compute phases that need it.
  • The setup followed the standard sequence: VT-d on the Xeon, IOMMU enabled at the kernel line, OVMF BIOS on the VM, vfio-pci bound to the device IDs, and machine type set to q35 for PCIe support.
  • qm start returned and the VM did not come up: not a kernel panic, not a host lockup, and not a 'device not found' error surfaced to the operator.
  • QEMU exited from the host's perspective and left nothing behind: no running guest, no console output, no obvious pointer at what had failed.
  • The Proxmox web UI showed the VM as stopped, exactly as it had before start was pressed.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A Proxmox 8 host refused to bring up an Ubuntu 24.04 VM after an NVIDIA RTX A4000 was passed through to it for a Filecoin sealing worker, and the only place the failure was legible was host dmesg [1][6]. The errors named the GPU; the author's account concludes the actual culprit was an HBA elsewhere in the box [9].

The build was the one every guide prescribes: VT-d on the Xeon, IOMMU enabled at the kernel line, OVMF on the VM, vfio-pci bound to the device IDs, machine type q35 for PCIe [2]. Then `qm start` returned and the guest did not come up [3]. No kernel panic, no host lockup, no "device not found" surfaced to the operator [3]. QEMU exited from the host's perspective and left nothing behind: no running guest, no console output [4]. The web UI showed the VM as stopped, exactly the state it was in before start was pressed [5]. The host stayed completely fine throughout [7].

The signal did exist, in a different terminal. `dmesg -w` on the host produced a burst of IRQ allocation failures and vfio interrupt remapping errors at VM startup, then silence [6]. That is the entire diagnostic surface, and it is emitted by the layer nobody was watching. The author notes the pattern read more like a misconfigured VM than a hardware conflict, which is why the cause survived so long [8].

It survived three sessions across about a week, roughly ninety minutes each, each one eliminating a correct-looking explanation without producing the answer [10] - call it four and a half hours of diagnosis [20]. Session one confirmed the obvious: `lspci -nnk -d 10de:24b0` reported `Kernel driver in use: vfio-pci`, with no nouveau and no nvidia [11], and the VM config carried the correct BDF, q35, and OVMF [12]. Everything about the device being passed was right. A checklist verifies the device you are thinking about, and the IOMMU topology contains devices you are not.

The reason this mattered is throughput arithmetic, not benchmark pride. Filecoin's SDR phase is memory-latency bound, CPU-heavy, and immune to GPU acceleration by protocol design, taking three to four hours per sector on this hardware [13]. TreeRC tree building is the GPU-eligible half: two to three hours per sector on CPU, roughly fifteen to twenty minutes on the A4000 [14]. Sealing without a GPU is therefore about five to seven hours per sector end to end [18], and the card removes somewhere between 100 and 165 minutes from each one [19]. The larger payoff is pipelining: SDR on one sector while the GPU does TreeRC on the previous one makes throughput gated by SDR alone [16], which the author puts at the difference between two sectors a day and five or six for a small provider [17], a 2.5x to 3x change [22] and consistent with the six-to-eight-per-day ceiling that a three-to-four-hour SDR implies [21].

Two operational takeaways. Start `dmesg -w` before you touch `qm start`, because the Proxmox UI and the `qm` return path carried no information about this failure at all [3][5][6]. And when interrupt remapping errors appear, enumerate the whole IOMMU topology rather than re-auditing the target card, which session one already proved clean [11][12]. The supplied account stops partway through that first session, so the mechanism tying the HBA to the interrupt errors is not in the material here: the HBA is the reported culprit, not a demonstrated one [9].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories