homelab-journey
Fixing a Postgres Backup Failure: Region Mismatch and a Missing NAS Copy
A LogQL finding led to a one-line AWS CLI region fix for postgres housekeeping — and uncovered a second gap: backups never synced to the NAS.
Documentation of my home lab evolution, from hardware builds to software experiments.
homelab-journey
A LogQL finding led to a one-line AWS CLI region fix for postgres housekeeping — and uncovered a second gap: backups never synced to the NAS.
homelab-journey
Replacing misleading volume-level Longhorn alerts with disk-level rules, recalibrating snapshot overhead to 50%, and how an unrelated incident accidentally produced the baseline data needed.
homelab-journey
Empty Longhorn panels led to metrics-server, a Talos cert quirk, a reboot that broke Garage, and 16 silent hours of failed backups.
homelab-journey
prometheus-server at 90% allocated, 5.6 GiB of real data. The alerts fired — but were they firing on the right thing? Snapshots, unreclaimed blocks, trim, and a PromQL join.
homelab-journey
Longhorn allocation never coming down? Here's how to enable automated fstrim via a RecurringJob on Talos Linux — and what it can't fix.
homelab-journey
A stale Grafana metric, eight days of silent failure, and five root causes that each made sense in isolation. How a backup stopped without anyone noticing.
homelab-journey
Rolling NVMe upgrade on a live Talos cluster: phase=failed, the extraMounts unlock sequence, and why storageReserved needs an explicit value.
homelab-journey
Three new RK1 worker nodes, three different Longhorn disk problems. The right way to add a node — and what happened when I didn't follow it.
homelab-journey
Loki ruler setup, four log-based alert rules from real Part 10 findings, a silent config gotcha, end-to-end test, and the complete dual alerting architecture.
homelab-journey
Six investigations across a Kubernetes cluster using LogQL: bootstrap artefacts, a silent 3-week backup failure, an Authelia crash sequence, and what high error volume actually means.
homelab-journey
Before hunting errors with LogQL, you need to know your log formats. Format discovery, the Garage INFO trap, klog envelopes, cardinality constraints, and fixing missing kube-system static pod logs.
homelab-journey
Talos system logs via Vector: why loki.source.tcp doesn't exist, how Vector fills the gap, and fixing Alloy's node-local filter in the same pass.