--- # HA-managed guests. DELIBERATELY SHORT: the cluster's stated posture is "no HA, # guests are disposable" (see ../../../CLAUDE.md). Anything listed here is an # exception that has earned it, and the reason belongs in the comment. pve_ha_resources: # vyos-rtr. It is the gateway for BOTH SDN VNets and an OSPF speaker, so losing # it takes labnet+retronet offline (the main LAN is unaffected — the NEC IX is # its gateway). Disk is on pve-rg (DRBD, place-count 2), so it can start on any # node; pve3 would attach diskless. # NOTE this buys crash-RESTART (~1-3 min outage), not seamless failover. Real # gateway HA would be a second VyOS with VRRP. - sid: "vm:100" state: started max_restart: 3 max_relocate: 2 # --- Watchdog --------------------------------------------------------------- # ⚠ WHY THIS MATTERS: adding ANY HA resource arms fencing cluster-wide. PVE's # default is `softdog`, a SOFTWARE watchdog — a kernel timer, which CANNOT fire if # the kernel itself is frozen. That is precisely the failure this cluster has # actually seen (pve2's Raven Ridge idle freeze). A hardware watchdog is # independent silicon and fires regardless. # # Verified available 2026-07-26: # pve1 (Intel i3-6100U) -> iTCO_wdt (timeout 30s) # pve2/pve3 (Ryzen 2400GE) -> sp5100_tco (timeout 60s) # Set per-host in inventory host_vars; empty string = leave PVE's softdog default. pve_ha_watchdog_module: ""