This is Fine! A podcast about resilience engineering and software

Colette Alexander and Clint Byrum

33 episodes listed below

Listen to the show on

PRESS PLAY

Find your next episode

All episodes

In their own words

A podcast about resilience engineering and software. Ever wondered why things on the internet break? Do you work in software and wish that you could have a Dear-Abby-Like call-in show that could answer your deepest questions about how to make your workplace suck less? We're here to help! Write us anonymously at our open question form Email us at: [email protected] Call us and leave a voicemail, or text us at: ‪(401) 592-7574‬

As heard by us

Based on 6 episodes we listened to · June 2026

A practitioner-facing resilience engineering show about incidents, AI tooling, and the organizational patterns behind engineering judgment.

This is Fine! works best as a show about what happens after clean engineering diagrams meet real operating pressure. In the sampled episodes, resilience engineering and software operations keep widening into human systems work: incident response, root cause analysis, adaptive…

Read our full review in PlayNext →

Why you'd press play

You want incident advice that knows panic laughs at a multi-page runbook.

Press play if you want

  • incident guidance that says get help before it says anything fancier
  • AI coding talk that treats the model as a supervision problem, not a magic intern
Read the full recommendation in PlayNext →

Talks about

Best episodes of This is Fine! A podcast about resilience engineering and software

Short reviews from the PlayNext desk, based on the episodes we processed.

Outsourcing and Resilience

A grounded look at reliability across company and geography boundaries.

Clint and Colette answer a listener's question about reliability across company, geography, and consultant-client boundaries, and they do it without pretending there is an easy fix.

All the things about Incident Command

Incident response as real organizational work, with clear cross org impact.

This episode keeps its focus on incident response and the people who end up carrying it. It asks who really runs the room, why some incidents move more cleanly than others, and what it takes to make that job count.

How long should you wait after an incident to do your retro?

A practical case for keeping incident learning close to the event.

How long to wait after an incident before holding the retro becomes the central question, and the episode favors moving quickly enough that the value does not drain away with memory.

Incident Status: On Hold w/special guest Will Gallego

A practical case for treating incident status as an operational signal, not just a label.

This episode treats incident status as more than admin language. The best stretch is the one that separates being blocked from simply being delayed, since that distinction changes what the rest of the team does next.

Root Cause Analysis vs. Resilience Engineering

A reminder that failure analysis has to look past the incident and at the patterns around it.

The discussion starts with a hockey wrist injury and uses it to ask what actually caused a failure. That frame works because the hosts keep widening the view: root cause analysis, recurring patterns, and the risk of missing system problems stay in play without losing the thread.

How (Not) to Introduce Resilience Engineering at Work with special guest Michelle Casey

A wary, practical conversation about how resilience ideas meet real listeners.

The conversation works best when it stays grounded in the practical problem of explaining resilience engineering to busy, skeptical people in their own language.

Lund University - Academic Theory and Practice

A grounded look at how Lund's resilience engineering program moves from theory into practice.

Lund University comes through as a practical place to learn resilience engineering, not just a name on a diploma. What lands best is the program's structure and the in-person labs, which make the time commitment feel concrete.

The Messy 9 and Coding with AI - A Panel Discussion

A loose panel that treats AI coding as a supervision challenge, not a magic trick.

A panel spends most of its time on the practical shape of coding with AI: where LLMs help, where they turn into a supervision problem, and when the better move is to leave them out.

Podcasts like This is Fine! A podcast about resilience engineering and software

  • Software Engineering Dailysoftwareengineeringdaily.com

    Long-form interviews on AI coding agents, infrastructure, and security, plus a recurring SED News roundup.

  • The Changelog: Software Development, Open SourceChangelog Media

    Founders and engineers on AI coding agents like Claude Code, and supply chain security.

  • The Stack Overflow PodcastThe Stack Overflow Podcast

    Stack Overflow's Ryan Donovan interviews founders and engineers on AI agents, protocols, and dev tools.

  • Talk Python To MeMichael Kennedy

    Michael Kennedy interviews Python maintainers on packaging, security, web tools, and AI.

  • Python BytesMichael Kennedy and Brian Okken

    Michael Kennedy and Brian Okken run down the week's Python tooling, typing, and packaging news.

  • Arrested DevOpsMatt Stratton, Trevor Hess, Jessica Kerr, and Bridget Kromhout

    Matt Stratton interviews DevOps practitioners on security, platform engineering, and trusting AI-written code.

Episodes

  1. 1

    Complex Systems and the Messy Nine w/special guests Dave Woods and John Allspaw

    The hosts introduce a special episode featuring John Allspaugh and Dr.

    ·1h 8m
  2. 2

    The Messy 9 and Coding with AI - A Panel Discussion

    The episode opens in familiar life mode, with weather, kids, skiing, hockey, and the kind of Monday chaos that makes everyone sound a little more human.

    ·1h 43m·4 clips
  3. 3

    What's the ROI on Reliability and Resilience Work?

    The hosts discuss a listener question about calculating the ROI of reliability and resilience work in software, exploring why it's hard to measure, the role of product management, and models like Rasmussen's safety boundaries.

    ·58m·5 clips
  4. 4

    Going Solid

    Two software reliability engineers do a deep dive into the academic paper 'Going Solid: A Model of System Dynamics and Consequences for Patient Safety' by Richard Cook and Jens Rasmussen.

    ·1h 2m·1 clip
  5. 5

    How long should you wait after an incident to do your retro?

    Hidden gem

    Clint and Colette open from a cabin in northern Michigan.

    ·44m·2 clips
  6. 6

    Outsourcing and Resilience

    Hidden gem

    Clint and Colette open with a bit of life turbulence and the familiar sense that the brain only has so much room.

    ·42m·2 clips
  7. 7

    How (Not) to Introduce Resilience Engineering at Work with special guest Michelle Casey

    Hidden gem

    Clint and Colette open with travel banter about Australia and Fiji, including kangaroos, wallabies, and a short detour into the odd force of Australian wildlife.

    ·53m·3 clips
  8. 8

    The Year in Resilience w/special guest John Allspaw

    Hidden gem

    Clint, Colette, and John Allspaw mark the turn of the year with teasing, holiday logistics, and a fair bit of exasperation about how ready everyone is to leave this year behind.

    ·53m·3 clips
  9. 9

    First Stories/Second Stories

    Hidden gem

    Clint and Colette start with a San Francisco hotel, a conference, and a 4.3 earthquake.

    ·53m·3 clips
  10. 10

    Incident Status: On Hold w/special guest Will Gallego

    Hidden gem

    The hosts open on Thanksgiving-week chaos and the familiar trick of trying to cram five days of work into two.

    ·43m·2 clips
  11. 11

    All the things about Incident Command

    Hidden gem

    Clint and Colette start with Halloween, family logistics, and the usual busy life before drifting into incident work.

    ·37m·2 clips
  12. 12

    The 2025 DORA Report w/special guest Fred Hebert

    Hidden gem

    Clint and Colette open on the time change, and neither of them sounds especially pleased about it.

    ·59m·3 clips
  13. 13

    Lund University - Academic Theory and Practice

    Hidden gem

    The conversation starts with holiday weekend banter, kids running around, and a quick nod to what a house sounds like when it is on the verge of becoming childless-ish.

    ·1h 5m·4 clips
  14. 14

    Building and Revising Adaptive Capacity Sharing for Technical Incident Response with Beth Adele Long

    Hidden gem

    Clint opens on a long ski weekend, fresh powder, and the sun-bronzed aftermath of too many hours outside.

    ·1h 9m·4 clips
  15. 15

    Root Cause Analysis vs. Resilience Engineering

    Hidden gem

    Clint starts with a hockey wrist injury.

    ·1h·3 clips
  16. 16

    SRECon Americas 2026 Recap

    The hosts discuss their recent experiences at SRECon, including a spicy talk that sparked an impromptu AMA session.

    Apr 14, 2026·55m
  17. 17

    Runbooks: the Good, Bad and Ugly w/special guest Andrew Hatch

    The hosts discuss fixing a microphone issue and reflect on their troubleshooting process, relating it to incident management in software.

    Jun 3, 2025·54m
  18. 18

    What is an incident? How come no one declares them?

    The hosts discuss resilience engineering, on-call experiences, and personal anecdotes about floating tanks.

    May 21, 2025·55m
  19. 19

    Chaos Engineering w/special guest Casey Rosenthal

    The hosts discuss personal anecdotes and a metaphor about stretching tape, relating it to their work in resilience engineering.

    May 7, 2025·48m
  20. 20

    Burnout on Aisle 3

    The hosts discuss a listener question about why SRE leaders in resilience engineering are burnt out and overloaded due to incident-related anti-patterns.

    Apr 26, 2025·47m
  21. 21

    Resilience, Complexity, and Your Boss a collab w/Punk Rock Safety

    Apr 9, 2025·55m
  22. 22

    Live from SRECon Americas 2025

    Mar 28, 2025·49m
  23. 23

    Teaser Episode - Season 2

    Mar 12, 2025·24m
  24. 24

    Episode 10 - When They go Full ITIL on You w/special guest John Allspaw

    Feb 20, 2025·52m
  25. 25

    Episode 9 - Learning from Incidents with special guest Alex Elman Video episode •

    The hosts discuss personal tech mishaps and introduce the show's theme of resilience engineering in software.

    Feb 12, 2025·49m
  26. 26

    Episiode 8 - Why Human Factors and Not Technical Ones

    Jan 29, 2025·37m
  27. 27

    Episode 7 - AI and Resilience with special guest Courtney Nash

    The hosts discuss recent incidents and the importance of resilience engineering in software, sharing personal experiences and reflections.

    Jan 22, 2025·54m
  28. 28

    Episode 6 - Can You Buy Resilience? With Special Guest Steve McGhee

    The hosts catch up after the holidays, discussing illness, a canceled trip, and a chicken-related incident.

    Jan 8, 2025·56m
  29. 29

    Episode 5 - Curating Your Resilience Engineering 101

    Dec 22, 2024
  30. 30

    Episode 4 - A Look at the 2024 DORA Report

    Dec 11, 2024·47m
  31. 31

    Episode 3 - Lions, Tigers and Metrics oh my!

    Dec 4, 2024·31m
  32. 32

    Episode 2 - Does Software Need Safety?

    Nov 21, 2024
  33. 33

    Episode 1 - Every Second Counts

    The hosts introduce their new podcast about resilience engineering in software, inviting listeners to share their work struggles and successes.

    Nov 7, 2024·35m