Prometheus
Fixing the Alerts: New Rules, Better Groups, Less Noise
Fan compound rule, seven new Prometheus rules, a six-group restructure, one Loki rule from the backup incident, and what's still on hold.
Guides on metrics collection, alerting, and monitoring infrastructure using Prometheus.
Prometheus
Fan compound rule, seven new Prometheus rules, a six-group restructure, one Loki rule from the backup incident, and what's still on hold.
Observability
Four months of running alerting on a homelab cluster. One false positive, one correct-but-broken underlying issue, eight days of silent backup failure, and what a rule-by-rule review actually looks like.
homelab-journey
Per-database pg_dump to Garage S3, a touch file state machine, daily/weekly rotation, a Go Prometheus exporter, and Alertmanager rules for the PostgreSQL backup system on Bletchley.
homelab-journey
Prometheus filling too fast — found a duplication and a 55% storage hog hiding in chart defaults. Fixed both, cut ingestion by 60%.
homelab-journey
Two files close the Prometheus ingress gap — then a latent cert-manager bug surfaces. The ESO key name mismatch that renewal tests don't catch.
homelab-journey
Grafana dashboards and 19 Prometheus alert rules for ZFS, SMART, disk temperature, fan speed, and Garage health — with manufacturer-sourced thresholds.
homelab-journey
Adding ZFS, SMART, and thermal collectors to the Bletchley cluster — and three ARM64 image attempts before finding one that actually works.
homelab-journey
Enabling Alertmanager on the Bletchley cluster: alerting rules, SMTP delivery, end-to-end testing, and two Talos-specific surprises.
homelab-journey
Importing Longhorn, Kubernetes, and resource dashboards onto the Bletchley cluster — and fixing the scrape config that was only collecting from one of four nodes.
homelab-journey
Adding Grafana to the Bletchley cluster: Longhorn-backed storage, pre-configured Prometheus datasource, and the Node Exporter Full dashboard showing live node metrics.
homelab-journey
Installing Prometheus and node exporter on Talos Linux: namespace labelling, values files, the duplicate pod gotcha, and confirming all four nodes are being scraped.