If you’re running Proxmox on an HP EliteDesk or any mini PC with a cheap NVMe SSD and your server keeps freezing every couple of days, this post might save you a month of headaches.
The Problem
My Proxmox server is an HP EliteDesk 800 G3 TWR. Intel i5-7500, 24GB RAM, Vi3000 256GB NVMe. It runs a bunch of containers and two VMs (TrueNAS and Home Assistant). Nothing crazy.
Every 1 to 5 days it would just… stop. The machine stayed on, fans spinning, power LED lit. But no SSH, no web UI, no ping. Nothing. The only fix was pulling the power and rebooting.
And the logs? Completely useless. journalctl just stopped mid-entry. No errors, no warnings, no kernel panic. Like someone hit pause on the entire system.
Everything I Tried (That Didn’t Work)
This went on for over a month. Here’s the list of things I ruled out:
- Temperatures – CPU was 37C, NVMe was 28C. Not even close to thermal issues.
- NVMe health – SMART passed, zero errors.
- RAM – Ran memtester on 21GB, two full passes. Clean.
- Power supply – Machine stayed on during every crash. UPS was fine.
- Kernel version – Same crash on kernel 6.5 and 6.8.
- NIC EEE bug – Disabled Energy Efficient Ethernet on the Intel I219-LM. Crashed again after 3.8 days.
- NIC ring buffers and GSO – Tuned them. No difference.
- CPU C-states – Already limited with
intel_idle.max_cstate=1. - PCIe power management – Already had
pcie_aspm=off.
I was pretty convinced it was the Intel NIC at this point. The I219-LM has a long history of bugs with the e1000e driver. So I bought a USB gigabit ethernet adapter, plugged it in, moved the cable over, and took the onboard NIC completely out of the picture.
It crashed again after 5 days. With the onboard NIC physically disconnected.
That’s when I knew the NIC was never the problem.
The Clue That Changed Everything
Early on I set up a hardware watchdog (Intel iTCO, 30 second timeout). The idea is simple: systemd pings the watchdog every few seconds. If the kernel completely locks up, systemd can’t ping it, and the hardware forces a reboot after 30 seconds.
But when the next crash happened, the watchdog didn’t reboot the machine. It just sat there frozen for over an hour until I manually power cycled it.
That means the kernel was still alive. systemd was still running, still feeding the watchdog. But everything else was dead. Cron stopped, journald stopped, all the VMs and containers were frozen. And the HDD activity light on the front panel was solid white instead of blinking. Zero disk I/O.
So the kernel was alive, but nothing could read or write to disk. Interesting.
The Answer: NVMe Power Saving
NVMe drives have a feature called APST (Autonomous Power State Transition). It lets the drive go into low-power sleep states when it’s idle. On drives with good firmware, this works fine. On budget drives like my Vi3000, the firmware has a bug where the drive goes to sleep and doesn’t wake up.
When that happens:
- The root filesystem just disappears
- Every process that tries to touch disk gets stuck waiting forever
- The kernel keeps running because it’s already loaded in memory
- systemd keeps feeding the watchdog because that’s just a memory-mapped write, no disk needed
- Eventually even network stuff dies because services need to write logs, read configs, etc.
This is actually a known bug in how systemd handles the watchdog. It never checks if the disk is actually working. It just checks if it can write to a memory address. There’s an open issue about it (systemd #21083) that hasn’t been fixed.
I also found a thread on the Proxmox forums from someone with the exact same hardware (HP EliteDesk 800 G3 + NVMe) having random NVMe disconnections. That sealed it for me.
The Fix
One kernel parameter:
nvme_core.default_ps_max_latency_us=0
This tells the kernel to never let the NVMe enter any power saving state. You can apply it right now without rebooting:
echo 0 > /sys/module/nvme_core/parameters/default_ps_max_latency_us
To make it survive reboots, add it to GRUB:
# Edit /etc/default/grub, add to the GRUB_CMDLINE_LINUX_DEFAULT line:
GRUB_CMDLINE_LINUX_DEFAULT="quiet nvme_core.default_ps_max_latency_us=0"
# Then run:
update-grub
Verify it’s set:
cat /sys/module/nvme_core/parameters/default_ps_max_latency_us
# Should show: 0
Did It Work?
Before the fix, the server never lasted more than 5.5 days. Usually crashed every 2-3 days.
After the fix: 23 days and counting. Not a single freeze.
The only downside is the NVMe uses a tiny bit more power at idle. We’re talking fractions of a watt. I’ll take it.
How to Tell If This Is Your Problem
You probably have the same issue if:
- Your Linux server freezes randomly every few days
- The machine stays on but is completely unreachable
- There are no errors in the logs, they just stop
- The HDD/NVMe activity light goes solid during the freeze
- You’re using a budget NVMe (Vi3000, Kingston A400/NV1, WD Green SN350, or similar cheap drives)
This seems especially common on Intel systems with 100/200 series chipsets (Skylake/Kaby Lake era).
Quick Reference
| Symptom | Random freezes every 1-5 days, no logs, machine stays on |
| Cause | NVMe APST bug – drive goes to sleep and doesn’t wake up |
| Fix | nvme_core.default_ps_max_latency_us=0 in kernel cmdline |
| Downside | Slightly higher idle power (fractions of a watt) |
If this saved your server too, drop a comment. It took a month of debugging, a USB NIC I didn’t need, a hardware watchdog, custom monitoring scripts, and way too many forced reboots to figure this out. Hopefully you found this post before going through all that.



Leave a Reply