Field Notes: hunting performance issues
Do you like to read about performance investigations? Me too!
I am fanatical about performance. Computers and software must be fast! Performance investigations are one of my favorite parts of how systems work. I have always liked problems of this type because they rarely turn out to be what you expect at first.
Just imagine. You are working on another feature in your project. You finally squeezed the last details out of your project manager. Tickets and design documents are full of details. You have a full mental model of the new feature in your head. You thought it through from every possible angle and wrote the initial code without O(n!) algorithms, having carefully thought through each detail. You deploy your changes to a staging cluster, run your first tests, open Grafana expecting to see the graphs you anticipated, and...
Does it sound familiar? You know how it ends, right? Yes, we both know. I have been in such situations dozens, if not hundreds, of times. We have all been there.
Although such situations are sometimes stressful, this is still one of my favorite parts of the profession. I have always liked this type of investigation. I have always had my nose buried in performance posts and have always been very happy to see them in my RSS feed. I always knew what to expect when I saw another post by Brendan Gregg, debugging stories from the ClickHouse Engineering Blog, Percona Database Performance Blog, and many, many more.
At some point, I thought, "hey, I also have a big collection of such stories. They might be interesting to other people too." So I have decided to start writing performance hunting stories from my own experience. It is a bit of a pity to leave them only in my memory, so I want to keep some notes about those investigations here - what I saw, what I tried, what didn't work, and what it actually turned out to be in the end.
Interested? Scroll below and let's dive together.