5 October 2026
Maintenance day
Last week the rack’s own sensors made the case for a maintenance day. The server’s CPU was running in the eighties at light load, and the big GPU had dropped off the PCIe bus three times in one day. So today I switched things off on purpose and opened them up.
The plan was hardware. Most of what I learned came from the things that had been quietly broken for a while.
The GPU that cooked itself
The RTX 3090 came out last week with dried-out paste and tired thermal pads. It idled at around 80°C, and under load it went to 97°C and fell off the bus. Today it got fresh paste and new pads.
It now idles at 35 to 42°C. The real test was fifteen minutes of nonstop image upscaling: 81 jobs, no errors, and the temperature climbed slowly from 75 to 80°C. The card hit its power limit, never its thermal limit. A 4096 × 4096 background removal takes 9.5 seconds. It is back to doing every GPU job in the house.
Swapping the server and the desktop
The server and my desktop traded places. The desktop’s motherboard became the server, and the server’s became my desktop. The new server no longer sits wedged into the top bay of the rack; it stands next to it, where it can breathe.
There was a second, smaller GPU meant to go in alongside the 3090. It is installed, and it does nothing, because the power supply has exactly one PCIe power cable and the card refused the adapter. A new power supply is on the list.
The first boot
The server came up and nothing on it started.
Moving to a different motherboard gave the network card a different name. The bridge that every virtual machine hangs off was configured for the old name, so it never came up, and nothing that needed the network could start. One fix at the console and it came up. The lesson is the useful part: a name the system derives from the hardware changes when the hardware does. Anything that matters should point at something that doesn’t.
Two cables instead of one
The server used to hang off a single cable that had a habit of dropping to 100 Mbps. It now has two links in an active-backup pair. I tested it the boring way: unplug one, watch, plug it back, unplug the other. Every direction lost at most one ping.
That evening the backup link dropped to 100 Mbps anyway. The main link didn’t notice, which is the point of having two. The cable is on the list.
The power
The old UPS’s batteries were not nearly dead, as I thought last week. They were completely dead. The rack is now on a smaller UPS whose batteries actually hold a charge, it restarts on its own when the power returns, and its beeper is switched off.
A small helper box that is supposed to wake the servers after a power cut had, it turns out, never done that once. A tool it needed was missing and it was aiming at an old address. It works now, and it sends the wake-up call at boot too.
I also wrote a staged shutdown plan for this UPS: first the work that can wait, then the storage, and the house’s own services last. Testing it means pulling the plug for real, so that is another day.
My own plugin turned itself off
The firewall box got opened too, for new paste. When it came back, my DPI bypass plugin was not running.
At boot the internet wasn’t up yet, the plugin’s health check failed three times in a row, and its watchdog switched it off to keep the household online. That is exactly what I built the watchdog to do. It is also slightly embarrassing to be caught by it. I started it again.
Alarms that had stopped looking
Moving to a new server exposed two alarms that only looked like they worked.
- CPU temperature. It was reading the sensor of the old processor’s brand. On the new one there was nothing to read, so it would have stayed silent however hot things got. It now reads the right sensor, with a lower threshold.
- Port speed. Renaming switch ports for the new cabling made the alarm see two histories for the same port. It didn’t fire wrongly. It stopped working altogether. It now takes one value per port.
Nearly all the other alarms of the day were me: every port I unplugged on purpose reported, faithfully, that it had been unplugged.
Everything thought it was the year 2000
The switch had restarted during the work, had no time source, and was confidently living in 2002. A few other devices go back to 1 January 2000 after every power cut, because a firewall rule never let them reach the time server. I added the rule and tested it by setting one of them an hour back. It corrected itself within seconds.
Still on the list
- A new power supply, so the second GPU can work.
- Point the GPU settings and the network config at stable IDs instead of names.
- The cable on the backup link.
- Pull the plug and test the shutdown plan from start to finish.
- Firewall updates, which want a reboot, so they wait for another day.
None of today’s fixes were dramatic. Most were things that had been reporting fine while doing nothing. That is the case for a maintenance day: not that something is on fire, but that you only find out what has been lying to you when you go and look.