MySQL + Vitess @ Slack: Managing Databases to Support 1+ Billion Daily Messages

Every message, reaction, and search in Slack ultimately hits a database. Serving over a billion messages a day means running MySQL and Vitess at a scale where nothing can be operated by hand; and where the database can never be the reason Slack is down.
Vitess gives you sharded MySQL. It does not give you a way to run that fleet at Slack’s scale; that part we had to build. This talk is a tour of the operational layer around Vitess, in three parts.
First, fleet automation: a detect-and-solve control loop and workflow-driven operations that manage thousands of instances: failovers, resharding, backups, and replacements; without a human in the loop. Second, connecting the app to the database fleet safely: Slack’s MySQL to Vitess proxy and overload monitoring & feedback loop back to client - so that the application degrades gracefully and problematic workloads don’t cascade into outages
If you’re adopting Vitess, you’ll leave knowing what to budget for beyond the database itself.
Speaker

Eduardo J. Ortega U. is a Physicist-turned Computer & Database Enthusiast who has spent the better part of the last decade operating fleets of tens of thousands of database instances - previously at Booking.com, and …









