Fred Hebert

@ferd.ca

Principal SRE @ honeycomb.io, Tech Book Author, Resilience in Software Foundation board member, Erlang Ecosystem Foundation co-founder, Resilience Engineering fan. SRE-not-sorry. blog: https://ferd.ca notes: https://ferd.ca/notes/

Delegating pressure to the final individual, who is now tasked with continuously fixing the entire system’s misalignments through their personal choices.

One of the ironies about AI agents in ops tasks is that it feels like there has never been as much interest in creating a forgiving environment with proper structural support than through promising to remove humans from it, finally forcing a less individualistic and blameful approach to design.

Even after years, one of the weirdest parts of gardening to me is needing to harden the seedlings before transplanting them. Like “yes hold on a minute I gotta take the plants out so they can play outdoors for a while, but they gotta be in before streetlights turn on” is a real and necessary thing.

The ongoing stream of software engineering pieces that mention that the future is in writing spec but never bother to define what a specification is or at what abstraction levels it should be is appalling; arguably, tickets are a spec, the code is a spec, and work between both is connecting dots.

“[…] we reached for Recon, an amazing tool for diagnosing issues […] (the related Erlang in Anger is more-or-less required reading as all on-call engineers end up scouring its pages eventually)” Wild! I wrote these 10 years ago to help coworkers, and they’re still useful for real world issues now!

You’ve Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice Outage

On March 25th, voice and video on Discord suffered major degradation beginning at 12:13 PDT, lasting a little over three hours. Learn how the issue originated, how it affected systems across Discord, ...

discord.com

Infinite love to my SRE coworkers, one of whom casually dropped this single line in a retro: "the industrial hourly deploy train and its consequences have been a disaster for society"

I wrote for @resilienceinsoftware.org on "Superficial Blamelessness", where under the label of "blamelessness", we avoid punishing people, yet still focus fixes and interventions based on the same individualistic framing rather than a broader systemic stance. resilienceinsoftware.org/news/11502437

Superficial Blamelessness

In 2012, John Allspaw (then CTO of Etsy) wrote a seminal blog post on the need for what he called Blameless Postmortems. Built off the notion of a “Just Culture” from the research of Sidney Dekker, he...

resilienceinsoftware.org

I've gotten an early copy of Crisis Engineering by Marina Nitze, Matthew Weaver, and Mikey Dickerson, and just in time for its release today, here's my review of it: ferd.ca/notes/on-cri... TL:DR; I like it, good perspectives and an interesting mix of approaches.

One of my assignments asks whether a given set of psychological constructs amount to “good science” and it’s very weird as someone who’s written zero papers and done no real science to just attempt to go “sure dawg it’s okay science I guess but these hundreds of researchers could pick better models”

Every time I am faced with this Swedish login form, I have to say 'Logga in' out loud with a terminator voice. It is one of the few rules we can't change and must be respected.

swedish login from with a button that says 'Logga in'

hell yeah @resilienceinsoftware.org swag is in! the law of requisite variety states that only variety in the regulator can destroy variety in the system being regulated. so if you need to deal with complexity you know you gotta join the club & begrudgingly increase complexity to keep things simple

Black hoodie showing 'anti complexity complexity club' in a red explosion based on the logo of the Resilience in Software Foundation.

would you rather have many incidents of various small to moderate size for the foreseeable future or just one very big incident and then be done with them for good?

As usual @grimalkina.bsky.social is worth listening to. These principles are also worth considering and applying in all sorts of contexts. Here’s a sample from safety research I happened to read just yesterday (Dekker - Reconstructing human contributions to accidents, 2002) that aligns with it!

What is striking about many accidents in complex systems is that people were doing exactly the sorts of things they would usually be doing the things that usually lead to success and safety. Mishaps are more typically the result of everyday influences on everyday decision making than they are isolated cases of erratic individuals behaving unrepresentatively.
People are doing what makes sense given the situational indications, operational pressures, and organizational norms existing at the time. Accidents are seldom preceded by bizarre behavior. People's errors and mistakes (such as there are in any objective sense) are systematically coupled to their circumstances and tools and tasks. [...] What people do makes sense to them at the time—it has to, otherwise they would not do it. People do not come to work to do a bad job; they are not out to crash cars or airplanes or ground ships. The local rationality principle says that people do things that are reasonable, or rational, based on their limited knowledge, goals, and understanding of the situation and their limited resources at the time. […] failures are baked into the nature of people's work and organization; that they are symptoms of deeper trouble or by-products of systemic brittleness in the way business is done. It means having to find out what people did back there and then actually make sense of it given the organization and operation that surrounded them.
To explain outcome failure, it is necessary to convert the search for human failures into a search for human sensemaking. The question is not "where did people go wrong?" but "why did this assessment or action make sense to them at the time?" Such real insight is derived not from judging people from the position of retrospective outsider, but from seeing the world through the eyes of the protagonists at the time. When looking at the sequence of events from this perspective, a very different story often struggles into view.
Cat Hicks@grimalkina.bsky.social · 5mo ago

Look, it's a mess out there and you can react to that mess by deciding everyone else is a moron or you could react to it by deciding most people are trying to get by with a different context than yours and start working the problem. Those are your choices pretty much, can't choose "no mess"

When I joined the program I'm currently in, I told myself I would not read more papers and technical books outside of it because I'd need to balance about my energy levels—not spending it all on this. I was right, but also there's lots of other cool nerd shit I want to read through now and welp.

I really need to do this year’s garden planning. It’s gonna be time for seedlings in a few weeks and it’s gonna be nice to once again get going on one of them hobbies where you can’t really obsessively dictate the pace nor feel pressured in going faster. Just watch the plants grow and see what goes.

Back in December, we had a large outage at work. The internal investigation took a while and the internal report was roughly 40 pages long. For the public, we managed to try and condense it to a much shorter format that we think can still offer useful insights to other organizations:

Incident Report: Exercises, Cleanups, and Evacuations

Every year, Honeycomb runs disaster recovery scenarios in multiple environments, including in production. Although each of our instances runs in a single region, on at least three Availability Zones ...

honeycomb.io

I believe I had a relatively intuitive sense of how much Swiss cheese I could consume in one sitting before I can take no more. I also believe I now have a fairly empirical sense of how much literature about the Swiss Cheese Model I can consume in one sitting before I can take no more.

I also think this discussion about how SRE work is being devalued by these products is at least parallel to the discussion about how people are willing to write clear documentation for *AI* consumption, but never put any value on it when it was for *actual people*.

Chastity Blackwell@blackisis.bsky.social · 5mo ago

I really appreciate Fred's thinking around this sort of thing -- it's very frustrating to see this dichotomy playing out again for the umpteenth time. One of the things I notice in all the AI SRE products Fred talks about is again, most of them claim to be able to do incident analysis *for* you...

this thread is touching on something i find very important: efficiency is sometimes opposed to (or in tension with) other goals efficiency can preclude generalist systems that create flexibility, decentralization and surplus that add resiliency, or non-expert participation that gets people involved

Courtney Milan@courtneymilan.com · 6mo ago

BTW the best parts of this are "inefficient" in an economic sense. Someone says "hey I want trans-themed whistles, can you do that" and someone else is like "I love that, I wanna do it!" or someone says "can I have some to pass out at my wedding" and we're like "sure, what are your colors?"