From Single Point of Failure to Highly Available: Re-architecting PMM for Production

Your monitoring system is the one thing that can’t go down when everything else does, yet for years, running PMM meant running a single point of failure to watch your fleet of highly-available databases. That changes with PMM HA.
In this talk we walk through how we redesigned Percona Monitoring and Management for high availability on Kubernetes: distributing state across three very different datastores (VictoriaMetrics for metrics, ClickHouse for query analytics, and PostgreSQL for inventory and configuration), and making each of them survive node loss, rescheduling, and full-cluster disaster recovery.
We’ll cover the hard parts honestly. You’ll leave with a concrete picture of what HA monitoring actually requires, and how to deploy a resilient PMM in your own environment.
Speaker

Tibi joined Percona in 2015 as a Consultant and has since transitioned into various roles, eventually becoming the Observability Tech Lead. Before Percona, he worked as a Senior Database Engineer at the world’s largest …









