Field notes · 2 September 2025
Reading crash clusters without panicking the release train
A calm method for sorting crash volume into fix-now, watch, and ignore buckets before your next freeze.
A spike in crash-free session rate after a release is noisy. Volume alone does not tell you whether checkout, login, or a rarely opened settings screen is broken.
We group stack traces by the screen or journey they interrupt, then by how many unique users are affected—not only how many events fired. One user looping a broken retry can inflate totals without matching real harm.
Fix-now covers crashes on payment, authentication, and the home feed for more than a small fraction of active users. Watch covers new clusters on secondary screens. Ignore covers known vendor library noise that has a vendor ticket and no user complaints.
Share the three buckets with product and support in the same readout. Support then knows which reports to escalate, and engineering avoids thrashing on low-impact noise during a freeze.