CloudNativePG Part 4: Breaking a Throwaway Database on Purpose
Extending Prometheus and Alertmanager to cover CloudNativePG, then chaos-testing failover, node loss, and restores against a throwaway database.
Extending Prometheus and Alertmanager to cover CloudNativePG, then chaos-testing failover, node loss, and restores against a throwaway database.
Installing CloudNativePG on Bletchley: the Cluster, MetalLB exposure, automated backups, and every real bug hit getting there.
Labeling Bletchley's two boards for CloudNativePG, auditing existing workloads for the same partition risk, and confirming there's room for it.
Same partition, opposite outcome, depending only on which board the primary happens to be on.
Garage's node ID reverted, Longhorn backups failed silently, and OpenBao turned out to have the same unmonitored risk — three alerts, two reactive and one proactive.
A LogQL finding led to a one-line AWS CLI region fix for postgres housekeeping — and uncovered a second gap: backups never synced to the NAS.
Replacing misleading volume-level Longhorn alerts with disk-level rules, recalibrating snapshot overhead to 50%, and how an unrelated incident accidentally produced the baseline data needed.
Empty Longhorn panels led to metrics-server, a Talos cert quirk, a reboot that broke Garage, and 16 silent hours of failed backups.
prometheus-server at 90% allocated, 5.6 GiB of real data. The alerts fired — but were they firing on the right thing? Snapshots, unreclaimed blocks, trim, and a PromQL join.
Longhorn allocation never coming down? Here's how to enable automated fstrim via a RecurringJob on Talos Linux — and what it can't fix.
Fan compound rule, seven new Prometheus rules, a six-group restructure, one Loki rule from the backup incident, and what's still on hold.
A stale Grafana metric, eight days of silent failure, and five root causes that each made sense in isolation. How a backup stopped without anyone noticing.