Prove it first — one real server day where doing nothing was the fix
One day on a company server: a web page answering with an error, a request that timed out, and drives labelled as something they were not.
All three looked like faults, and all three were fine — the move that saved the data was the move nobody made.
The habit that told them apart — six steps you can run on your own machine.
My rule for the whole day — the same six steps you tick below, in my own capitals:
OBSERVE -> PROVE -> PASS
DOCUMENT -> CHANGE -> VERIFY AGAIN
The day in one glance
What the wrong move would have cost
| What it looked like | What the wrong move costs | What the machine proved |
|---|---|---|
| Two cables that should work as one | internet drops at the next restart | both came back as one link |
| Hard drives carrying old labels from a past job | healthy storage erased for good | nothing was touched |
| A folder called backup | a restore with nothing in it | gap found, plan fixed |
| A path between two machines | the reason hidden, not solved | left as it is, reason written down |
| A page that did not exist yet | eighteen live pages edited with no way back | page built, eighteen pages updated |
The six steps, and why each one exists
- Look first — note what you see before you touch it; that note is your way back
- Prove it — the machine’s own report settles what a guess cannot
- Let it pass — anything a restart forgets is unfinished work
- Write it down — the reason goes on the record before the change
- Change one thing — two changes hide which one mattered
- Check again — that is the proof it worked
📦 The day's own toolkit — 33 free tools, sorted by the six steps
rare find ·
nothing to install, it is already on the box · take what you need and skip the rest
Pick one, run it once, and you have done that step for real.
1 · OBSERVE — look first: see what is really there. (that day: drives wearing someone else’s labels)
See what every network card is really doing — nic-xray · github.com/ciroiriarte/nic-xray- Ask the switch what that cable is plugged into — lldpd · github.com/lldpd/lldpd
Find out what an old RAID label really is — mdadm --examine · man7.org/linux/man-pages/man8/mdadm.8
Read the storage’s own record without touching it — zdb -l · openzfs.github.io — zdb.8
Take a ZFS disk apart on paper, trusting no pool — zfs-forensic · github.com/SecurityRonin/zfs-forensic- Watch a drive’s health over months, not minutes — Scrutiny · github.com/AnalogJ/scrutiny
- See the honest free-space numbers at a glance — duf · github.com/muesli/duf
2 · PROVE — ask the machine, not a guess. (that day: a backup path that answered nothing)
- Watch one hop alone, with loss per hop — trippy · github.com/fujiapple852/trippy
See the whole route drawn on a map, in a window — OpenTrace · github.com/Archeb/opentrace- Read the bond’s own negotiation frames — termshark · github.com/gcla/termshark
- Name the process sending packets that go unanswered — bandwhich · github.com/imsnif/bandwhich
- Get a shell that already holds every network tool — netshoot · github.com/nicolaka/netshoot
3 · PASS — let it prove itself. (that day: a reboot that brought everything back)
- Check a backup archive without restoring it — vma verify · pve.proxmox.com/wiki/VMA
See what is really mounted, not what the config says — findmnt --verify · man7.org/linux/man-pages/man8/findmnt.8- Test a drive before it holds anything you care about — disk-burnin-and-testing · github.com/Spearfoot/disk-burnin-and-testing
4 · DOCUMENT — write it down before you change it. (that day: a reason written down instead of a route typed)
Put your whole settings folder in git — etckeeper · etckeeper.branchable.com
Freeze a machine’s entire configuration into one file — cfg2html · github.com/cfg2html/cfg2html- Record the session so you can replay what you did — asciinema · github.com/asciinema/asciinema
- Know the moment a backup stops reporting — Healthchecks · github.com/healthchecks/healthchecks with
runitor · github.com/bdd/runitor
5 · CHANGE — one thing at a time. (that day: eighteen pages edited with a copy kept)
- Rehearse a whole change and read what it would do — ansible --check --diff · docs.ansible.com — check mode
Test a web-server config before you reload it — nginx -t · nginx.org/en/docs/switches- Snapshot and replicate ZFS on a schedule — sanoid with syncoid · github.com/jimsalterjrs/sanoid
- Make that replication continuous — zrepl · github.com/zrepl/zrepl
- Browse a snapshot and restore one file in a browser — Backrest · github.com/garethgeorge/backrest
6 · VERIFY AGAIN — check it after. (that day: the same check, run again)
List the services still running an old library — needrestart · github.com/liske/needrestart- Undo exactly the changes between two snapshots — snapper undochange · manpages.opensuse.org — snapper.8
- Check the redirect and the certificate honestly — testssl.sh · github.com/testssl/testssl.sh
- Run real HTTP checks from a text file — hurl · github.com/Orange-OpenSource/hurl
- Prove every link on a site still works — lychee · github.com/lycheeverse/lychee
Read a web-server config and catch what eyes slide past — gixy-next · github.com/MegaManSec/Gixy-Next- Test one hostname against one address, DNS untouched — xh --resolve · github.com/ducaale/xh
See the redirect chain and handshake phase by phase — httptap · github.com/ozeranskii/httptap
Same method, already on this board, written by me: onehack.st/t/325503 — my two-minute self-undo timer for a live network change · onehack.st/t/325527 — my updater that finds every copy and undoes a failed update · onehack.st/t/326190 — my router rebuild on this same server
The answer came from the machine, not a guess — and the day’s best work was touching nothing.
Your problem, worked live
Send a machine, a script, a route that goes somewhere impossible to Ask Us Live.
Picked for what they teach, then solved live on 2026-10-17T18:30:00Z — nothing staged, nothing rehearsed.
Open ones stay up with the notes.
Bring it in onehack.st/tag/help.
🔍 The day, problem by problem
Two network cards bonded into one link. I wanted two physical cards working as one link, the 802.3ad/LACP setup in the kernel’s own bonding document. The file said the bond existed. Reality had both cards answering on their own while the bond sat with no working members. I read it, rebuilt it, then rebooted: the bond, both members, the address, the route and the internet all came back by themselves. A configuration that survives a restart is the one worth keeping.
A tunnel with its own identity. A remote host needed management access through an existing WireGuard tunnel into the OPNsense edge. Copying an existing client configuration was the fast move, and the wrong one: every client needs its own cryptographic identity. So the host got a unique private key, a unique public key, a unique tunnel address and its own peer entry — and the private key stayed on the machine, out of the conversation. I brought the tunnel up by hand first, not at boot. Prove it, then make it permanent.
The remote host that runs production. This host runs production infrastructure on pve.proxmox.com, so before a single backup command I inventoried everything: the backup filesystem held 1.1 TB total, about 250 GB used and 793 GB free; the root filesystem sat healthy at 94 GB total with 76 GB free; and the thin pool held far less real data than the sum of all provisioned disk sizes suggested. Provisioned capacity is not allocated blocks. Running on it: one OPNsense virtual machine and fifteen containers.
Drives wearing someone else’s labels. Several drives carried old Linux RAID labels from an earlier life, so I left those labels alone and asked the storage itself: two OpenZFS mirrors, both pools ONLINE, zero read errors, zero write errors, zero checksum errors, healthy SMART status, so nothing was touched.
The path that needed proving. The remote host needed to reach the independent backup server at home. Ten packets sent, none received, 100% loss. The easy move was to say “it just needs a static route” and type one. I did not.
Following the traffic to the next hop. The remote host has an internal transit network to its OPNsense VM. I identified the OPNsense address on that network and tested that address alone — which explained the path. The route stayed unwritten on purpose, with the reason written down instead: a route typed to quieten a symptom hides why the symptom came.
A folder named backup, one machine away from being real. A directory by that name held copies on the same machine: useful staging, and one failure away from being the only copy. Following where the data would actually travel after a failure showed the gap, in writing. The return path is scheduled now, which is worth more than a green tick on a status page.
The page the day built. The day’s other half was a public page where people send in a problem — I called it Ask Us Live. Deploying it immediately demonstrated the method: one deploy, then the same check again. Then the production navigation: I searched the live document root first and found 18 pages carrying the sequence, updated all 18, verified all 18 links, and kept the previous version for a rollback.
Nine things I deliberately did not touch: the disks with strange labels, the WireGuard tunnel at boot, the bond before its reboot, backups written onto the root filesystem, a static route, NGINX answering 404 and then 301, the firewall during a timed-out test, and eighteen production pages without a way back. In several cases the best engineering decision of the day was: leave it alone until the proof says otherwise.
What ChatGPT actually did. It held the context, questioned my assumptions and built the safe step-by-step checks.
It administered nothing: I had the consoles and the commands.
Where a suggestion met a system that proved otherwise, the suggestion changed — the reading did not.[1]
Tomorrow: prove the backup return path end to end, then install it. Something else will break soon enough, and I would rather prove why before I fix it.
The useful question about an assistant is not was it right the first time — it is did the process end up agreeing with the machine. ↩︎
!