SRE Weekly Issue #332

View on sreweekly.com

Articles

How Razorpay’s Notification Service Handles Increasing Load

Their notification service had complex load characteristics that made scaling up a tricky proposition.

Anand Prakash — Razorpay

How we improved on-call life by reducing pager noise

Coalescing alerts and adding dependencies in AlertManager were the key to reducing this team’s excessive pager load.

steveazz — GitLab

What’s allowed to count as a cause: ALERRT edition

Lorin Hochstein has started a series of blog posts on what we can learn about incident response from the Uvalde school shooting tragedy in the US. This article looks at how an organization’s perspective can affect their retrospective incident analysis.

Lorin Hochstein

The fog of war in Uvalde

My claim here is that we should assume the officer is telling the truth and was acting reasonably if we want to understand how these types of failure modes can happen.

Every retrospective ever:

We must assume that a person can act reasonably and still come to the wrong conclusion in order to make progress.

Lorin Hochstein

User settings, Lamport clocks and lightweight formal methods

How do you synchronize state between multiple browsers and a backend, and ensure that everyone’s state will eventually converge? These folks explain how they did it, and a bug they found through testing.

Jakub Mikians — Airspace Intelligence

MTTR: lower isn’t always better

MTTR is a mean, so it doesn’t tell you anything about the number of incidents, among other potential pitfalls.

Dan Slimmon

Google Cloud Platform outage report: europe-west2 cooling failure

Last week, I included a GCP outage in europe-west2. This week, Google posted this report about what went wrong, and it’s got layers.

Bonus: another GCP outage report

Google

It’s time to leave the leap second in the past

Meta wants to do away with leap seconds, because they make it especially difficult to create reliable systems.

Oleg Obleukhov and Ahmad Byagowi — Meta

3 common pitfalls of post-mortems