Fixing a Postgres Backup Failure: Region Mismatch and a Missing NAS Copy
A LogQL finding led to a one-line AWS CLI region fix for postgres housekeeping — and uncovered a second gap: backups never synced to the NAS.
A LogQL finding led to a one-line AWS CLI region fix for postgres housekeeping — and uncovered a second gap: backups never synced to the NAS.
Replacing misleading volume-level Longhorn alerts with disk-level rules, recalibrating snapshot overhead to 50%, and how an unrelated incident accidentally produced the baseline data needed.
Empty Longhorn panels led to metrics-server, a Talos cert quirk, a reboot that broke Garage, and 16 silent hours of failed backups.
prometheus-server at 90% allocated, 5.6 GiB of real data. The alerts fired — but were they firing on the right thing? Snapshots, unreclaimed blocks, trim, and a PromQL join.
Longhorn allocation never coming down? Here's how to enable automated fstrim via a RecurringJob on Talos Linux — and what it can't fix.
Fan compound rule, seven new Prometheus rules, a six-group restructure, one Loki rule from the backup incident, and what's still on hold.
A stale Grafana metric, eight days of silent failure, and five root causes that each made sense in isolation. How a backup stopped without anyone noticing.
Four months of running alerting on a homelab cluster. One false positive, one correct-but-broken underlying issue, eight days of silent backup failure, and what a rule-by-rule review actually looks like.
Rolling NVMe upgrade on a live Talos cluster: phase=failed, the extraMounts unlock sequence, and why storageReserved needs an explicit value.
Three new RK1 worker nodes, three different Longhorn disk problems. The right way to add a node — and what happened when I didn't follow it.
Loki ruler setup, four log-based alert rules from real Part 10 findings, a silent config gotcha, end-to-end test, and the complete dual alerting architecture.
Six investigations across a Kubernetes cluster using LogQL: bootstrap artefacts, a silent 3-week backup failure, an Authelia crash sequence, and what high error volume actually means.