<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Posts on Coffee or Blog</title>
    <link>https://blog.coffeeordeath.dev/posts/</link>
    <description>Recent content in Posts on Coffee or Blog</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.coffeeordeath.dev/posts/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Fumbling Through Loss</title>
      <link>https://blog.coffeeordeath.dev/posts/fumbling-through-loss/</link>
      <pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/fumbling-through-loss/</guid>
      <description>I got it wrong the first time and probably the second time too. I&amp;rsquo;ll likely get it wrong again, so here I am writing a reminder to future me.
A beloved colleague died unexpectedly. The email came the same day, evening for me, direct, no euphemisms. The video call was set for the next afternoon, late enough that people had a night to sit with it before they had to be on camera together.</description>
      <content>&lt;p&gt;I got it wrong the first time and probably the second time too. I&amp;rsquo;ll likely get it wrong again, so here I am writing a reminder to future me.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;A beloved colleague died unexpectedly.
The email came the same day, evening for me, direct, no euphemisms. The video call was set for the next afternoon, late enough that people had a night to sit with it before they had to be on camera together. Someone senior said her name out loud and then let the silence sit instead of filling it. By every measure I&amp;rsquo;d use later to judge how this should go, it was done right.&lt;/p&gt;
&lt;p&gt;Then, the following week, we sold part of the business, and a dozen people I&amp;rsquo;d worked closely with - people I looked forward to working with - were suddenly on someone else&amp;rsquo;s org chart. There was no email like the first one. Just an announcement calling it a strategic move and a Slack channel that went quiet within hours. Nobody had died. But it landed on top of someone who just had, and I remember thinking: we know how to do this well, we did it a week ago, so why does this one feel like an afterthought. I didn&amp;rsquo;t say that to anyone. It felt like complaining about a paper cut in a room where someone had just been in surgery.&lt;/p&gt;
&lt;p&gt;Nobody hands you a runbook for this. We have runbooks for a full disk at 3 AM. We have severity levels, escalation paths, comms templates, and a blameless review afterward. When a person is gone, most of us walk into the room with nothing, and the team watches to see what we do with it.&lt;/p&gt;
&lt;p&gt;This is what I&amp;rsquo;ve pieced together since. None of it is expertise, believe me. It&amp;rsquo;s fumbling, written down.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Kenneth Doka has a name for grief that isn&amp;rsquo;t openly acknowledged or socially allowed: &lt;a href=&#34;https://en.wikipedia.org/wiki/Disenfranchised_grief&#34;&gt;disenfranchised grief&lt;/a&gt;. Most workplaces create it without meaning to, because nothing at work makes room for grief. There&amp;rsquo;s no casserole, no service you&amp;rsquo;re expected at, no week off. There&amp;rsquo;s a Slack announcement, a reassigned Jira board, a cold email, and a standup the next morning where everyone is waiting to see if it&amp;rsquo;s okay to talk about it.&lt;/p&gt;
&lt;p&gt;Usually, it isn&amp;rsquo;t, because nobody says it is. So people don&amp;rsquo;t. They process it alone or not at all, and it shows up later as the engineer who stopped arguing in design reviews, or the one who started interviewing.&lt;/p&gt;
&lt;p&gt;It doesn&amp;rsquo;t have to be a death to land this way. I think about three kinds:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Someone dies.&lt;/strong&gt; The obvious one, and the one we&amp;rsquo;re least equipped for. The shock is real, and so is the operational hole.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Someone leaves without leaving.&lt;/strong&gt; A divestiture, a carve-out, a team transferred to another company. Pauline Boss calls this &lt;a href=&#34;https://www.mayoclinichealthsystem.org/hometown-health/speaking-of-health/coping-with-ambiguous-grief&#34;&gt;ambiguous loss&lt;/a&gt;: the person is gone from your world but not gone. They&amp;rsquo;re still on LinkedIn. You still see their name in &lt;code&gt;git blame&lt;/code&gt;, their avatar in an open pull request. Their Slack account is deactivated and their knowledge is still the only copy of how half the system works. It gets announced as a strategic win, so there&amp;rsquo;s nothing to mourn, officially. People mourn anyway, quietly, with no end point.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A third of the room is gone.&lt;/strong&gt; Layoffs. David Noer wrote about this in the early 90s and called it &lt;a href=&#34;https://psycnet.apa.org/record/1997-36609-008&#34;&gt;layoff survivor syndrome&lt;/a&gt;. The people who stay carry guilt about why it wasn&amp;rsquo;t them, scan every all-hands for the next signal, and pick up the work of the people who left without anyone adjusting a single deadline. A &lt;a href=&#34;https://leadershipiq.com/blogs/leadershipiq/29062401-dont-expect-layoff-survivors-to-be-grateful&#34;&gt;Leadership IQ survey&lt;/a&gt; of about 4,000 layoff survivors found 74% said their own productivity dropped and 69% said the quality of the company&amp;rsquo;s product did too.&lt;/p&gt;
&lt;p&gt;That tracks with everything I&amp;rsquo;ve seen. People under threat stop taking risks. They pick the safe option, ship less, flag less, propose nothing. Stevan Hobfoll&amp;rsquo;s &lt;a href=&#34;https://en.wikipedia.org/wiki/Conservation_of_resources_theory&#34;&gt;Conservation of Resources theory&lt;/a&gt; says it plainly: under stress, people protect what they have left. Your team is conserving.&lt;/p&gt;
&lt;p&gt;Three different losses, one failure mode. The company treats it as a transaction and moves on. The people can&amp;rsquo;t, or can&amp;rsquo;t without feeling like they&amp;rsquo;re doing it wrong.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The systems keep going. This is the part I&amp;rsquo;d want someone to have told me.&lt;/p&gt;
&lt;p&gt;When someone dies, the automation doesn&amp;rsquo;t know. Their name is still in the PagerDuty rotation. They&amp;rsquo;re still in CODEOWNERS, so PRs auto-request their review. The calendar invite for their 1:1 still fires every Tuesday. A bot pings the channel asking them to update a stale ticket. A welcome-back reminder from the HR tool arrives the day they would have returned from PTO.&lt;/p&gt;
&lt;p&gt;Each one of those is a small wound delivered by a system someone on your team built. They&amp;rsquo;ll notice every one. I notice every one.&lt;/p&gt;
&lt;p&gt;The first week is still the best version of this I&amp;rsquo;ve seen, so the checklist starts there. Treat it like an incident, because it is one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Confirm the facts, then say it the same day. Plainly, no euphemisms. Rumour spreads faster than a Slack post.&lt;/li&gt;
&lt;li&gt;Tell the people closest to them first, privately, before the broad announcement. Not in a channel. Not by email if you can help it.&lt;/li&gt;
&lt;li&gt;Give people a night before you put them on camera together.&lt;/li&gt;
&lt;li&gt;When you do get everyone together, say their name and let the silence sit. Don&amp;rsquo;t fill it.&lt;/li&gt;
&lt;li&gt;Pause the automation before it fires: on-call schedules, code review assignment, recurring meetings, ticket bots, HR and payroll workflows, anything that sends a message with their name on it. Don&amp;rsquo;t deprovision the account in a way that erases their history. Their commits and docs are part of the team&amp;rsquo;s memory.&lt;/li&gt;
&lt;li&gt;Give people the EAP number and say out loud that using it is normal. Then use it yourself if you need to.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t reassign their work in the same breath as the announcement. It will need to happen. It doesn&amp;rsquo;t need to happen that day.&lt;/li&gt;
&lt;li&gt;Let the team decide how they want to remember them. Some teams want a channel, some want a donation, some want to name something after them. It&amp;rsquo;s not yours to pick.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Divestitures and layoffs need the same playbook. We had it. We didn&amp;rsquo;t use it the second time. Say what happened plainly. Name who&amp;rsquo;s gone. Don&amp;rsquo;t let the reorg announcement be the only acknowledgement. How the company treats the people who left is the clearest signal the people who stayed will ever get about how they&amp;rsquo;ll be treated.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;As a leader, you lost them too.&lt;/p&gt;
&lt;p&gt;Arlie Hochschild&amp;rsquo;s term is &lt;a href=&#34;https://en.wikipedia.org/wiki/Emotional_labor&#34;&gt;emotional labor&lt;/a&gt;, managing your own feelings to meet what the job expects. For a manager during a loss, the job expects calm. So you do the work of being calm, all day, for everyone, and then you get to your own feelings sometime after the last meeting, if at all.&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;re also in the middle. Leadership above you wants the roadmap back on track. Your team below you needs time and a lighter load. Both are reasonable. They can&amp;rsquo;t both be fully met, and you&amp;rsquo;re the one who has to decide how much of each to give, usually without guidance, usually while absorbing some of the dropped work yourself.&lt;/p&gt;
&lt;p&gt;I don&amp;rsquo;t have a clean answer for this. What&amp;rsquo;s helped:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Telling my team that I&amp;rsquo;m finding it hard too. Not performing it, just saying it. It gives them permission they won&amp;rsquo;t otherwise take.&lt;/li&gt;
&lt;li&gt;Saying what I expect, out loud. &amp;ldquo;I don&amp;rsquo;t expect to see you online until the next steps meeting.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Having one peer manager I can be honest with. Not my boss, not my team.&lt;/li&gt;
&lt;li&gt;Pushing the roadmap conversation upward explicitly instead of quietly eating the gap. &amp;ldquo;We lost a person and a third of our context. Here&amp;rsquo;s what moves.&amp;rdquo; That&amp;rsquo;s a capacity conversation that you&amp;rsquo;re allowed to have.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you&amp;rsquo;re the person above the managers: they need somewhere to put this. If you don&amp;rsquo;t give it to them, you&amp;rsquo;ll find out where they put it when they resign.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;What I&amp;rsquo;d do now is mostly small and unglamorous.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Say the thing.&lt;/strong&gt; Name the person. Name what happened. Staying quiet can feel like you&amp;rsquo;re giving people space. From their side, it looks like nobody noticed. That Leadership IQ survey had one bright spot: survivors who rated their manager high on visibility, approachability, and candor were 72% less likely to report a productivity drop. Being present and straight with people moves the number.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Don&amp;rsquo;t aim for closure.&lt;/strong&gt; Especially with layoffs and divestitures, there&amp;rsquo;s no event that ends it. People will keep running into the gap for months, every time they need the person who knew how the billing pipeline worked. Expect that. It&amp;rsquo;s not a performance problem, but it is a visibility problem: the time lost working around that gap won&amp;rsquo;t show up anywhere unless you put it in front of leadership.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cut scope on purpose.&lt;/strong&gt; The team is smaller and sadder. Pretending it has the same capacity is how you turn grief into burnout. Stop the low-value work explicitly and let the team help decide what goes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Make it safe to not be fine.&lt;/strong&gt; Amy Edmondson&amp;rsquo;s &lt;a href=&#34;https://psychsafety.com/about-psychological-safety/&#34;&gt;psychological safety&lt;/a&gt; isn&amp;rsquo;t a poster. After a loss, it&amp;rsquo;s whether someone can say &amp;ldquo;I can&amp;rsquo;t take that on right now&amp;rdquo; in standup and have it be okay. It doesn&amp;rsquo;t come back on its own, and you rebuild it by going first.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Keep checking in after everyone else stops.&lt;/strong&gt; Week one, everyone is kind. Week six, the calendar has moved on and the person who&amp;rsquo;s still struggling is alone with it. Put a reminder in your calendar six weeks out. Then ask. Let them brush it off if they&amp;rsquo;re fine. Be the one who remembered.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Accept that you&amp;rsquo;ll fumble.&lt;/strong&gt; You&amp;rsquo;ll say something clumsy. You&amp;rsquo;ll forget to pause a bot. You&amp;rsquo;ll check in with the wrong person and miss the right one. Show up anyway.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Those two weeks sit side by side for me - one loss handled so well it became the standard I measure everything else against, the other that never got a fraction of that care because nobody called it a loss out loud.&lt;/p&gt;
&lt;p&gt;If I could go back, I&amp;rsquo;d tell myself the paper cut is still a cut. It doesn&amp;rsquo;t need a comparison to earn the right to be noticed, and neither does anyone on your team who&amp;rsquo;s still running into the gap six weeks from now.&lt;/p&gt;
&lt;p&gt;I bring them up now, usually when someone I&amp;rsquo;m talking to is in their own version of it and doesn&amp;rsquo;t quite believe it&amp;rsquo;s allowed to hurt this much. It&amp;rsquo;s not advice at that point, just proof that someone else has been there. I still get a pit in my stomach telling it, longer after it happened than I&amp;rsquo;d have guessed. I&amp;rsquo;ve stopped expecting that to go away.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Release Discipline in the Age of AI-Accelerated Development</title>
      <link>https://blog.coffeeordeath.dev/posts/release-discipline-in-the-age-of-ai-accelerated-development/</link>
      <pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/release-discipline-in-the-age-of-ai-accelerated-development/</guid>
      <description>Photo by Patrick Konior on Unsplash
A maturity model for shipping fast without breaking customer trust.
I made an earlier argument that deploying code and releasing a feature are not the same event, and that treating them as one causes unnecessary risk. That argument still holds. But the environment it was written for has changed.
AI-assisted development has made writing and shipping code dramatically cheaper. A change that used to take a sprint now takes an afternoon.</description>
      <content>&lt;p&gt;&lt;img alt=&#34;Control room overview&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/hero-control-room.jpg&#34;&gt;
&lt;em&gt;Photo by &lt;a href=&#34;https://unsplash.com/@patrickkonior&#34;&gt;Patrick Konior&lt;/a&gt; on &lt;a href=&#34;https://unsplash.com/photos/a-view-of-a-control-room-from-above-FwHJ6mJjVd4&#34;&gt;Unsplash&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;A maturity model for shipping fast without breaking customer trust.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I made &lt;a href=&#34;https://blog.coffeeordeath.dev/posts/decoupling-deployment-from-release/&#34;&gt;an earlier argument&lt;/a&gt; that deploying code and releasing a feature are not the same event, and that treating them as one causes unnecessary risk. That argument still holds. But the environment it was written for has changed.&lt;/p&gt;
&lt;p&gt;AI-assisted development has made &lt;em&gt;writing&lt;/em&gt; and &lt;em&gt;shipping&lt;/em&gt; code dramatically cheaper. A change that used to take a sprint now takes an afternoon. That&amp;rsquo;s a genuine gain, but it doesn&amp;rsquo;t automatically make the &lt;em&gt;judgment&lt;/em&gt; around releasing that change any faster or better. When the cost of producing change drops and the discipline around exposing change doesn&amp;rsquo;t rise to match it, the result is exactly what customers are describing: things move, break, or reappear differently, without warning, more often than they can absorb.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Here&amp;rsquo;s the core claim: in 2026, the bottleneck in software delivery is not how fast we can write code. It&amp;rsquo;s how deliberately we control who sees a change, when, and with what warning. AI removes the first constraint. It makes the second one more important, not less.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t a call to slow engineering down. It&amp;rsquo;s a call to stop routing all of that new speed straight at the customer, unfiltered.&lt;/p&gt;
&lt;h2 id=&#34;the-symptom-were-solving-for&#34;&gt;The symptom we&amp;rsquo;re solving for&lt;/h2&gt;
&lt;p&gt;Customers aren&amp;rsquo;t complaining that we ship too much. They&amp;rsquo;re complaining about three specific things, and it&amp;rsquo;s worth being precise because each has a different fix:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Rate of change&lt;/strong&gt; — things they learned last month look or behave differently now, with no signal it was coming.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Forced change&lt;/strong&gt; — a workflow they depend on changed or disappeared, and they had no way to stay on the old behavior while they adjusted.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unpredictability&lt;/strong&gt; — they can&amp;rsquo;t tell the difference between &amp;ldquo;this is a bug&amp;rdquo; and &amp;ldquo;this is intentional,&amp;rdquo; because both arrive the same way: silently.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;None of these are solved by shipping less. All three are solved by &lt;strong&gt;separating four things we currently bundle into one event&lt;/strong&gt;: code lands in production, a feature becomes visible, a customer is told about it, and an old behavior is retired. Each deserves its own timeline, its own owner, and its own decision.&lt;/p&gt;
&lt;p&gt;&lt;img alt=&#34;Deployment vs. Release Timeline diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/deployment-vs-release-timeline.svg&#34;&gt;&lt;/p&gt;
&lt;h2 id=&#34;the-four-question-test-for-every-user-visible-change&#34;&gt;The four-question test for every user-visible change&lt;/h2&gt;
&lt;p&gt;Before a change reaches a real customer, someone should be able to answer:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Answer determines&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Is it visible?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Whether it needs a flag at all&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Is it disruptive?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Whether it needs staged rollout or can go straight to 100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Is it reversible?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Whether we need a dual code path, or a simple toggle is enough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Does it remove something?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Whether it needs a deprecation notice and a minimum notice period&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A change that&amp;rsquo;s invisible (backend refactor, performance work) skips all of this, deploy it and move on. Everything else routes through the toolkit below. The mistake teams make under AI-accelerated velocity is treating &lt;em&gt;all&lt;/em&gt; changes like the first category because the code was cheap to produce.&lt;/p&gt;
&lt;h2 id=&#34;blast-radius-classification&#34;&gt;Blast-radius classification&lt;/h2&gt;
&lt;p&gt;Use this before merge, not after a customer complains. It should take seconds, not a meeting.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Class&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Rollout pattern&lt;/th&gt;
&lt;th&gt;Notice required&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Invisible&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Refactor, perf work, dependency bump&lt;/td&gt;
&lt;td&gt;Deploy directly&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Visible, additive&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New optional feature, new button&lt;/td&gt;
&lt;td&gt;Flag → rings → GA&lt;/td&gt;
&lt;td&gt;Changelog entry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Visible, behavioral&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Changed default, changed workflow&lt;/td&gt;
&lt;td&gt;Flag → opt-in beta → staged % → GA&lt;/td&gt;
&lt;td&gt;Advance in-app notice + changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Disruptive / breaking&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Removed capability, forced migration&lt;/td&gt;
&lt;td&gt;Dual path, minimum notice window, opt-out during transition&lt;/td&gt;
&lt;td&gt;Direct communication, not just changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;img alt=&#34;Blast-Radius Classification diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/release-discipline/blast-radius-classification.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;The AI-specific risk is that classes 2–4 get produced at the same velocity as class 1, and without this checkpoint they get &lt;em&gt;shipped&lt;/em&gt; at that velocity too. The checkpoint is cheap. Skipping it is what customers are feeling.&lt;/p&gt;
&lt;h2 id=&#34;the-toolkit&#34;&gt;The toolkit&lt;/h2&gt;
&lt;h3 id=&#34;1-feature-flags-but-with-a-lifecycle-not-just-an-onoff-switch&#34;&gt;1. Feature flags, but with a lifecycle, not just an on/off switch&lt;/h3&gt;
&lt;p&gt;Flags fail long-term not because teams don&amp;rsquo;t use them, but because nobody owns their &lt;em&gt;end state&lt;/em&gt;. A flag that&amp;rsquo;s still in the codebase eighteen months after full rollout is a liability, not a safety net. Every flag needs a stated life stage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Rollout flag (temporary, by default):&lt;/strong&gt; exists to ramp a change safely. Has an owner and an expected removal date at creation time. Options at end of life:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Graduate&lt;/strong&gt; — feature is fully rolled out and stable → flag is deleted, new behavior becomes the only behavior.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Revert&lt;/strong&gt; — didn&amp;rsquo;t work → old path stays, new path is removed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extend deliberately&lt;/strong&gt; — a real reason exists to keep ramping slowly (enterprise contracts, regulatory cohorts) → re-approve with a new date, don&amp;rsquo;t let it drift by default.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Operational flag (long-lived, by design):&lt;/strong&gt; kill switches, ops-only toggles. These are allowed to live indefinitely, but should be inventoried separately from rollout flags so the two don&amp;rsquo;t get confused in review.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Permission / entitlement flag:&lt;/strong&gt; controls who gets a capability (plan tier, beta cohort, region). Long-lived by design, owned by product.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&#34;Flag Lifecycle diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/release-discipline/flag-lifecycle.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ownership split that matters:&lt;/strong&gt; engineering owns &lt;em&gt;whether the flag exists and works&lt;/em&gt;; product/support own &lt;em&gt;when it flips for whom&lt;/em&gt;. That split is still correct, it just now needs a &lt;strong&gt;flag registry with an expiry date on every rollout flag&lt;/strong&gt;, reviewed monthly, or the flag count grows faster than the org&amp;rsquo;s ability to reason about it. AI-generated code makes it trivially easy to wrap a new flag around everything; that&amp;rsquo;s a reason to enforce the registry harder, not skip it.&lt;/p&gt;
&lt;h3 id=&#34;2-progressive-rollout-rings&#34;&gt;2. Progressive rollout rings&lt;/h3&gt;
&lt;p&gt;Don&amp;rsquo;t choose between &amp;ldquo;ship to everyone&amp;rdquo; and &amp;ldquo;ship to no one.&amp;rdquo; Use rings, and pick the entry ring based on the blast-radius class above:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Internal&lt;/strong&gt; — team, then company-wide dogfooding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Design partners / opt-in beta&lt;/strong&gt; — customers who explicitly asked to try new things early. This is where self-service enablement lives (see below).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Staged percentage&lt;/strong&gt; — 5% → 25% → 100%, gated on real usage signals and support ticket volume, not a calendar date.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;General availability&lt;/strong&gt; — announced, documented, supported.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img alt=&#34;Progressive Rollout diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/progressive-rollout.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;A change only needs to pass through every ring if it&amp;rsquo;s disruptive. Additive, low-risk changes can compress rings 3–4. The point isn&amp;rsquo;t ceremony, it&amp;rsquo;s that &lt;em&gt;someone decided&lt;/em&gt; how much exposure this change gets before it got any, instead of exposure being a side effect of when the deploy happened to land.&lt;/p&gt;
&lt;h3 id=&#34;3-self-service-enablement-beta-opt-in&#34;&gt;3. Self-service enablement (beta opt-in)&lt;/h3&gt;
&lt;p&gt;Give customers a way to &lt;em&gt;choose&lt;/em&gt; to be early, rather than &lt;em&gt;discovering&lt;/em&gt; they&amp;rsquo;re early. Concretely: an in-product &amp;ldquo;early access&amp;rdquo; or &amp;ldquo;labs&amp;rdquo; area where customers can turn on features ahead of GA, with a visible way to turn them back off. This does two things at once:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It converts &amp;ldquo;why did this change on me&amp;rdquo; into &amp;ldquo;I opted into this.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;It gives you a self-selected, motivated feedback cohort before wide rollout, which is a better signal than support tickets after the fact.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The precondition for this to work: opted-in features must be genuinely reversible by the customer. If turning it off doesn&amp;rsquo;t actually turn it off, don&amp;rsquo;t offer it as opt-in, that&amp;rsquo;s a forced change wearing an opt-in costume, and customers notice the difference immediately.&lt;/p&gt;
&lt;h3 id=&#34;4-change-communication-tiered-to-disruption-level&#34;&gt;4. Change communication, tiered to disruption level&lt;/h3&gt;
&lt;p&gt;Not every change deserves the same announcement weight. Matching the tier from the blast-radius table:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Additive:&lt;/strong&gt; changelog / release notes. Low ceremony, always shipped.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Behavioral:&lt;/strong&gt; in-app notice &lt;em&gt;before&lt;/em&gt; the change reaches a user&amp;rsquo;s account, plus changelog. Should say what&amp;rsquo;s changing and, if relevant, how to preview it early via opt-in beta.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Disruptive / breaking:&lt;/strong&gt; direct communication (email, CSM, in-app banner with acknowledgment) with a stated timeline, not just a mention in release notes. If there&amp;rsquo;s a deadline for an old behavior going away, that deadline should be visible to the customer well before it arrives, not just to us.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A useful gut check: if a customer would reasonably say &amp;ldquo;I wish someone had told me,&amp;rdquo; the tier was too low.&lt;/p&gt;
&lt;h3 id=&#34;5-deprecation-without-forced-change&#34;&gt;5. Deprecation without forced change&lt;/h3&gt;
&lt;p&gt;The single biggest driver of the &amp;ldquo;forced change&amp;rdquo; complaint is removing something before the data says it&amp;rsquo;s safe to. The pattern that avoids it:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;New path ships behind a flag, old path stays live.&lt;/li&gt;
&lt;li&gt;Both paths run in parallel long enough to gather real usage data, not a fixed arbitrary window, but until usage of the old path is actually low or zero.&lt;/li&gt;
&lt;li&gt;A minimum notice period is announced &lt;em&gt;before&lt;/em&gt; the old path is scheduled for removal, sized to the disruption class (days for a minor UI tweak, months for a workflow customers have built process around).&lt;/li&gt;
&lt;li&gt;Only once notice has elapsed and usage has dropped does the old path get deleted, as its own change, separately reviewed from the feature that replaced it.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img alt=&#34;Dual code paths diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/dual-code-paths.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;This costs engineering effort (two paths, temporarily) in exchange for the thing customers are actually asking for: predictability. That trade is almost always worth making for anything customer-facing.&lt;/p&gt;
&lt;h2 id=&#34;what-ai-accelerated-development-changes-specifically&#34;&gt;What AI-accelerated development changes, specifically&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The proposal rate goes up, the review/rollout rate doesn&amp;rsquo;t automatically follow.&lt;/strong&gt; The gap between them is where uncontrolled change leaks out. Treat rollout classification as a required step in the definition of done, not an optional nicety, it&amp;rsquo;s the part of the pipeline that didn&amp;rsquo;t get faster, so it needs to be protected, not skipped under pressure to keep pace.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flag sprawl accelerates.&lt;/strong&gt; If it&amp;rsquo;s cheap to generate a flag-wrapped change, flags will be created faster than they&amp;rsquo;re retired unless the registry-and-expiry habit from the toolkit above is enforced, not just recommended.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;It&amp;rsquo;s just a small change&amp;rdquo; stops being a reliable signal.&lt;/strong&gt; AI can produce a large, behaviorally significant change with the same apparent effort as a small one. Blast-radius classification should be based on what the change &lt;em&gt;does&lt;/em&gt;, not how much manual effort it took to write, that correlation has broken.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Volume of change is a communications problem now, not just an engineering one.&lt;/strong&gt; If ship velocity increases, either communication cadence scales with it, or customers experience the increase as noise. Product/support capacity to write good change notices becomes a real constraint on release pace, plan for it explicitly rather than discovering it during a rollout.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;a-maturity-model-to-self-assess-against&#34;&gt;A maturity model to self-assess against&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Deploy vs release&lt;/th&gt;
&lt;th&gt;Flags&lt;/th&gt;
&lt;th&gt;Communication&lt;/th&gt;
&lt;th&gt;Deprecation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;0 — Coupled&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Deploy = release, always&lt;/td&gt;
&lt;td&gt;None, or ad hoc&lt;/td&gt;
&lt;td&gt;Release notes after the fact, if at all&lt;/td&gt;
&lt;td&gt;Rip and replace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1 — Flags exist&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Some features flagged&lt;/td&gt;
&lt;td&gt;Used inconsistently, no registry&lt;/td&gt;
&lt;td&gt;Changelog exists&lt;/td&gt;
&lt;td&gt;Removal decided by engineering alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2 — Rings defined&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standard for risky changes&lt;/td&gt;
&lt;td&gt;Registry exists, no expiry discipline&lt;/td&gt;
&lt;td&gt;Tiered by change type&lt;/td&gt;
&lt;td&gt;Notice periods exist but informal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3 — Customer-facing control&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Default practice&lt;/td&gt;
&lt;td&gt;Expiry dates enforced, monthly review&lt;/td&gt;
&lt;td&gt;Self-serve beta / early access live&lt;/td&gt;
&lt;td&gt;Data-gated removal, formal notice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4 — Governed change contract&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Assumed, invisible as a decision&lt;/td&gt;
&lt;td&gt;Flags treated as inventory with owners&lt;/td&gt;
&lt;td&gt;Change comms scale with ship velocity by design&lt;/td&gt;
&lt;td&gt;Deprecation SLA is a documented customer commitment&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Most orgs feeling the pain described at the top of this post are sitting between Level 0 and 1, while their &lt;em&gt;code output&lt;/em&gt; has jumped to what used to require a Level 3 org&amp;rsquo;s engineering throughput. That mismatch, velocity outrunning governance, is the actual root cause, not &amp;ldquo;we ship too much&amp;rdquo; or &amp;ldquo;AI is risky.&amp;rdquo; The fix is closing the governance gap, not throttling the code.&lt;/p&gt;
&lt;h2 id=&#34;getting-started-in-order&#34;&gt;Getting started, in order&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Adopt the four-question test and blast-radius table&lt;/strong&gt; as a required checklist step before merge for anything customer-visible. This alone catches most of the &amp;ldquo;forced change&amp;rdquo; and &amp;ldquo;no warning&amp;rdquo; complaints.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stand up a flag registry with expiry dates.&lt;/strong&gt; Even a spreadsheet beats nothing. Review monthly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Define your rings&lt;/strong&gt; and who has authority to move a change from one ring to the next.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ship a self-serve early-access surface&lt;/strong&gt;, even a minimal one. It reframes the relationship from &amp;ldquo;this happened to me&amp;rdquo; to &amp;ldquo;I chose this.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Write a deprecation SLA&lt;/strong&gt;, minimum notice periods by disruption class, and hold to it publicly. This is the fastest way to rebuild trust with customers who&amp;rsquo;ve been burned by forced changes before.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;the-bottom-line&#34;&gt;The bottom line&lt;/h2&gt;
&lt;p&gt;AI didn&amp;rsquo;t create the need for release discipline, it just removed the natural speed limit that used to give teams cover for not having it. The teams that get faster &lt;em&gt;and&lt;/em&gt; keep customer trust in 2026 aren&amp;rsquo;t the ones writing more code carefully. They&amp;rsquo;re the ones who separated &amp;ldquo;we built it&amp;rdquo; from &amp;ldquo;you see it&amp;rdquo; a long time ago, and are now deliberately re-tuning that separation for a world where the first half of that sentence happens ten times faster than it used to.&lt;/p&gt;
&lt;p&gt;Deploy fearlessly. Release deliberately. That was true before AI. It&amp;rsquo;s the whole game now.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>The Pass-Through Problem</title>
      <link>https://blog.coffeeordeath.dev/posts/2026-04-10---the-pass-through-problem/</link>
      <pubDate>Fri, 10 Apr 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/2026-04-10---the-pass-through-problem/</guid>
      <description>Someone asks you a question. You don&amp;rsquo;t know the answer off the top of your head, so you paste it into Claude, copy the response, and send it back.
That&amp;rsquo;s not a human interaction. That&amp;rsquo;s a very slow API call with extra steps.
I keep seeing this pattern, at work, on forums, in social communities, and it makes me wonder if we&amp;rsquo;ve completely missed the point. Not of AI. Of ourselves.</description>
      <content>&lt;p&gt;Someone asks you a question. You don&amp;rsquo;t know the answer off the top of your head, so you paste it into Claude, copy the response, and send it back.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s not a human interaction. That&amp;rsquo;s a very slow API call with extra steps.&lt;/p&gt;
&lt;p&gt;I keep seeing this pattern, at work, on forums, in social communities, and it makes me wonder if we&amp;rsquo;ve completely missed the point. Not of AI. Of ourselves.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s what the pattern tells the person on the other end: &lt;em&gt;I could not be bothered to engage with your question.&lt;/em&gt; The response might be accurate. It might even be helpful. But it carries a clear signal: you were not worth the effort of actual thought. The sender probably didn&amp;rsquo;t mean it that way. It doesn&amp;rsquo;t matter. That&amp;rsquo;s what arrived.&lt;/p&gt;
&lt;p&gt;If I wanted an LLM answer, I would have asked the LLM. I asked &lt;strong&gt;you&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The useful version of this technology isn&amp;rsquo;t a faster copy-paste. It&amp;rsquo;s a forcing function. The interactions that &lt;em&gt;don&amp;rsquo;t&lt;/em&gt; require a human (the FAQ, the status update, the &amp;ldquo;what does this acronym mean&amp;rdquo;) should go away entirely. Build a better doc. Point to a bot. Remove the friction. That&amp;rsquo;s the job.&lt;/p&gt;
&lt;p&gt;What should be left over is the stuff that actually requires a person. Judgment calls. Messy context. The question behind the question. The moments that only work if someone actually gives a shit.&lt;/p&gt;
&lt;p&gt;Those interactions deserve more, not less. If AI is buying back any of your time, that&amp;rsquo;s where it goes. And if you&amp;rsquo;re one of the people who already gets that — start talking about it. The people around you are probably already losing the thread, and they don&amp;rsquo;t know it yet.&lt;/p&gt;
&lt;p&gt;What&amp;rsquo;s happening instead is the opposite. The low-effort questions get a low-effort LLM response dressed up as a human answer. The hard questions get the same treatment. And slowly, the expectation of genuine engagement just&amp;hellip; lowers.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s a word for what this looks like at scale: &lt;a href=&#34;https://pluralistic.net/2023/01/21/potemkin-ai/#hey-ho-lets-go&#34;&gt;enshittification&lt;/a&gt;. And I don&amp;rsquo;t think the people doing it are cynical. I believe they&amp;rsquo;re trying. They picked up - or were forced to use - a powerful tool and pointed it at a real problem. Nobody handed them a manual for which problems it should and shouldn&amp;rsquo;t touch.&lt;/p&gt;
&lt;p&gt;But good intentions don&amp;rsquo;t change what lands on the other end. The person who asked you something real still got a hollow answer. The gap between meaning well and doing well is exactly where the work is.&lt;/p&gt;
&lt;p&gt;The technology isn&amp;rsquo;t the problem. Mistaking the shortcut for the improvement is. And that mistake doesn&amp;rsquo;t land evenly. For some people, a hollow answer isn&amp;rsquo;t an inconvenience - it&amp;rsquo;s confirmation of something they were already afraid was true. The impact isn&amp;rsquo;t equally distributed. Neither is the responsibility to fix it.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Why Doesn&#39;t Anyone Own That?</title>
      <link>https://blog.coffeeordeath.dev/posts/why-doesnt-anyone-own-that/</link>
      <pubDate>Mon, 23 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/why-doesnt-anyone-own-that/</guid>
      <description>It&amp;rsquo;s 5 PM on a Thursday. Something broke mid-morning and your team has been grinding on it for hours. Good engineers, working hard, coming up empty because the system is poorly documented and the people who built it have either moved on or moved teams. You&amp;rsquo;ve burned most of the day and you&amp;rsquo;re no closer to resolution.
So you escalate. You ping the leads. You ask in the channel where someone, surely, knows this service well enough to point you in the right direction.</description>
      <content>&lt;p&gt;It&amp;rsquo;s 5 PM on a Thursday. Something broke mid-morning and your team has been grinding on it for hours. Good engineers, working hard, coming up empty because the system is poorly documented and the people who built it have either moved on or moved teams. You&amp;rsquo;ve burned most of the day and you&amp;rsquo;re no closer to resolution.&lt;/p&gt;
&lt;p&gt;So you escalate. You ping the leads. You ask in the channel where someone, surely, knows this service well enough to point you in the right direction.&lt;/p&gt;
&lt;p&gt;Silence.&lt;/p&gt;
&lt;p&gt;Not because nobody cares. Because nobody actually owns it. There&amp;rsquo;s a name in the CMDB, but that person will tell you they haven&amp;rsquo;t touched it since a reorg reshuffled their priorities. There&amp;rsquo;s a runbook, but it describes a version of the system that no longer exists. There&amp;rsquo;s a Slack channel with forty members and no clear authority.&lt;/p&gt;
&lt;p&gt;This is where the real cost of diffuse ownership lives. Not in the architecture review, not in the postmortem writeup. Right here, in a live incident, with a team that has already spent most of a day spinning their wheels and a business that is losing patience.&lt;/p&gt;
&lt;p&gt;The easy diagnosis is that engineers don&amp;rsquo;t want ownership. Nobody wants the accountability when something critical falls over. That read is intuitive, and it&amp;rsquo;s mostly wrong.&lt;/p&gt;
&lt;p&gt;Most engineers I&amp;rsquo;ve worked with want to own something. They want the stakes, the identity, the sense that they know a system cold and it shows. Ownership is a fundamental motivator in this work. It&amp;rsquo;s one of the few things that makes the job feel like more than ticket throughput.&lt;/p&gt;
&lt;p&gt;The problem isn&amp;rsquo;t appetite. It&amp;rsquo;s that we&amp;rsquo;ve built organizations that punish people for taking it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;we-say-we-want-owners-then-we-punish-them&#34;&gt;We say we want owners. Then we punish them.&lt;/h2&gt;
&lt;p&gt;Most performance systems weren&amp;rsquo;t designed to reward operational excellence. They were designed to measure visible output — features shipped, projects delivered, things a manager can point to in a calibration meeting and defend. Keeping a critical service healthy produces none of that. No launch. No demo. No ribbon cutting.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s what it does produce: a rating of &amp;ldquo;meets expectations.&amp;rdquo; Which, in most four-point HR scales, means no raise. Sounds fair until you consider that the engineer just spent a quarter instrumented to the gills on a service that didn&amp;rsquo;t go down, pushed for refactors the team had been deferring for two years, and responded to every incident that came their way. They did their job flawlessly. The organization scored it as ordinary because ordinary is the only category it fits in.&lt;/p&gt;
&lt;p&gt;That engineer is not doing that again next year. Neither is anyone who watched it happen.&lt;/p&gt;
&lt;p&gt;The performance review is just the most visible mechanism. The same signal gets sent in sprint planning when operational work gets bumped for roadmap items, in reorgs when the &amp;ldquo;boring&amp;rdquo; platform work gets deprioritized, in how leaders talk about the team that keeps things running versus the team shipping the new thing. It accumulates. People are paying attention even when you think they aren&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;Rational engineers respond rationally. Ownership becomes a liability, ambiguity becomes protection, and the next time someone asks who owns the broken thing, there genuinely isn&amp;rsquo;t an answer.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;accountability-without-authority-is-just-blame&#34;&gt;Accountability without authority is just blame.&lt;/h2&gt;
&lt;p&gt;There&amp;rsquo;s a second trap that even well-intentioned leaders fall into, and it&amp;rsquo;s more insidious than the performance review problem because it often comes wrapped in the language of empowerment.&lt;/p&gt;
&lt;p&gt;We assign ownership without giving people the actual power to act.&lt;/p&gt;
&lt;p&gt;Think about what that looks like in practice. An engineer is named owner of a critical platform service. They identify that the deployment process is bypassing validation steps and creating instability. They flag it. They write it up. They ask for a two-week pause on releases to address it. The response from leadership: &amp;ldquo;We can&amp;rsquo;t stop the roadmap, let&amp;rsquo;s find another way.&amp;rdquo; There is no other way. The service degrades, falls over, and in the postmortem the question on the table is why the service owner didn&amp;rsquo;t catch this sooner. It plays out constantly, in slightly different costumes, and the person in the owner role learns the same lesson every time.&lt;/p&gt;
&lt;p&gt;Ownership without the authority to say no to a risky deployment, to pull capacity from feature work when a system is genuinely at risk, to set a standard and enforce it — that isn&amp;rsquo;t ownership. It&amp;rsquo;s a title with no teeth. The person holding it gets the pager, gets the postmortem questions, gets the performance conversation, and had none of the leverage required to prevent any of it. Smart engineers recognize that setup fast. Word travels.&lt;/p&gt;
&lt;p&gt;The organizations that do this aren&amp;rsquo;t malicious. Most of them genuinely believe they&amp;rsquo;ve created clear ownership. They put a name in a doc and called it done. What they actually created is a designated scapegoat with a service catalog entry.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;the-cost-is-higher-than-you-think&#34;&gt;The cost is higher than you think.&lt;/h2&gt;
&lt;p&gt;Diffuse ownership doesn&amp;rsquo;t just create operational risk, though it absolutely does that. It degrades the institutional knowledge your organization depends on.&lt;/p&gt;
&lt;p&gt;When nobody owns a system, nobody deeply understands it. Documentation becomes aspirational fiction. Runbooks describe what someone &lt;em&gt;intended&lt;/em&gt; the system to do, not what it actually does. Incidents take longer, cost more, and teach less, because there&amp;rsquo;s no through-line of accountability that connects cause to effect to learning.&lt;/p&gt;
&lt;p&gt;And the talent impact is brutal. The engineers who want to own things, the ones you most want to retain, will eventually get tired of operating in that vacuum. They&amp;rsquo;ll go somewhere that lets them. The engineers who remain will have self-selected for comfort with ambiguity and low accountability. That is a culture you cannot easily reverse.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;what-actually-changes-this&#34;&gt;What actually changes this.&lt;/h2&gt;
&lt;p&gt;This isn&amp;rsquo;t a process problem. You can&amp;rsquo;t RACI-chart your way out of a culture that punishes ownership. You have to change what you reward.&lt;/p&gt;
&lt;p&gt;Start with how you evaluate performance. If your engineering managers cannot point to explicit, visible credit given to engineers for operational ownership — in promotion cases, in calibration sessions, in public recognition — then you are implicitly telling your team that it doesn&amp;rsquo;t count. Fix that first.&lt;/p&gt;
&lt;p&gt;Next, make ownership legible. A service with a named owner, a clear charter, and explicit authority boundaries is a service people will actually step up to own. Ambiguity is where accountability goes to die. Define the thing. Name the owner. Give them real power over their domain.&lt;/p&gt;
&lt;p&gt;Toyota figured this out on the factory floor decades ago. &lt;a href=&#34;https://en.wikipedia.org/wiki/Andon_(manufacturing)&#34;&gt;Their Andon&lt;/a&gt; system gives any line worker the authority to stop production the moment they identify a problem — not escalate it, not flag it for later, stop it. The entire premise is that the person closest to the issue has both the standing and the backing to act. It works because leadership built a culture where pulling that cord is the right call, not a career risk. We&amp;rsquo;re supposedly more sophisticated in software, yet we routinely put engineers in positions where they can see the problem clearly, have no authority to stop anything, and get blamed when it blows up anyway.&lt;/p&gt;
&lt;p&gt;Back your owners when it&amp;rsquo;s hard. When someone pulls the cord on a deployment, or pushes back on a timeline because the system genuinely can&amp;rsquo;t absorb it, support that call publicly. If you quietly override it three times, you&amp;rsquo;ve taught everyone watching that ownership is theater.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;the-question-worth-asking&#34;&gt;The question worth asking.&lt;/h2&gt;
&lt;p&gt;The next time you find yourself in that room, the incident is running, the silence is thick, and nobody&amp;rsquo;s hand goes up, resist the urge to blame your people.&lt;/p&gt;
&lt;p&gt;Ask instead what you, and the organization above you, built that made this feel like the right call.&lt;/p&gt;
&lt;p&gt;Because I promise you: somewhere, at some point, someone tried to own that thing. And something happened that taught them not to.&lt;/p&gt;
&lt;p&gt;Your job is to find out what that was, and fix it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Good ownership isn&amp;rsquo;t found. It&amp;rsquo;s cultivated, or it&amp;rsquo;s extinguished. Pick one.&lt;/strong&gt;&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>A Centimeter or Two</title>
      <link>https://blog.coffeeordeath.dev/posts/a-centimeter-or-two/</link>
      <pubDate>Fri, 13 Feb 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/a-centimeter-or-two/</guid>
      <description>The blood on my hands didn&amp;rsquo;t register as blood at first.
I&amp;rsquo;d thrown my mask off after the shot. Reflex. Pain management. Give the jaw some air. The pain was sharp and then immediately dull, the kind that makes you think stinger before it makes you think anything else. So I knelt on the ice, head down, eyes on the white below me, and stayed there. A couple of guys skated over.</description>
      <content>&lt;p&gt;The blood on my hands didn&amp;rsquo;t register as blood at first.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d thrown my mask off after the shot. Reflex. Pain management. Give the jaw some air. The pain was sharp and then immediately dull, the kind that makes you think &lt;em&gt;stinger&lt;/em&gt; before it makes you think anything else. So I knelt on the ice, head down, eyes on the white below me, and stayed there. A couple of guys skated over. Told them I was fine. Just a stinger. Give me a minute.&lt;/p&gt;
&lt;p&gt;I believed that.&lt;/p&gt;
&lt;p&gt;I stood up. Reached around to check on things. Brought my hands to my eyes.&lt;/p&gt;
&lt;p&gt;They were covered in blood.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;This is a shinny game. Eleven at night in Toronto, which is the only time you can get ice when you&amp;rsquo;re thirty-five and have kids and a job and a long list of reasons why you should probably be in bed. The guy who shot it was younger than the rest of us, wearing red, came in from the blue line on my right side and unloaded a hard wrist shot. I was dropping to make the save. The puck hit my shoulder on the way down, deflected hard, and found the gap between the bottom edge of my mask and the top of my neck guard.&lt;/p&gt;
&lt;p&gt;I didn&amp;rsquo;t know that gap existed.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve thought about that a lot since.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Once the blood started it didn&amp;rsquo;t stop. Someone handed me a towel. My friend swapped it for a wad of paper towels and I held it to my chin, to my neck, trying to figure out where exactly the problem was. He offered another wad a few minutes later. I replaced it and glanced briefly at the soaked one as it went in the trash.&lt;/p&gt;
&lt;p&gt;Then the nausea hit, and someone said &lt;em&gt;we should call 911&lt;/em&gt;, and I heard myself say that I was starting to feel light headed, and someone guided me to a bench.&lt;/p&gt;
&lt;p&gt;What happened next is the part I&amp;rsquo;ve been glossing over for two years.&lt;/p&gt;
&lt;p&gt;My hands went numb. Then my feet. Then they started to clench, fingers curling inward like something was pulling them closed. I slumped right. The feeling left my face. Then the right side of my body. I remember someone to my left trying to make sure I stayed conscious, talking at me, keeping me present. I remember paramedics removing my pads and skates while I sat there. I remember my vision blurring. I remember the questions getting harder to answer and then, for a while, not a lot else.&lt;/p&gt;
&lt;p&gt;I know I got onto a stretcher. I don&amp;rsquo;t remember getting onto it. I know it was cold outside, Toronto in winter, and I remember the bumps of the path to the ambulance because it had snowed. I remember the lights inside the ambulance and the sound of the sirens and the medics figuring out which hospital could take me fastest. I remember one of them going through my wallet for my health card and either not finding it or deciding I was declining too fast to stop. Both seemed possible at the time.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The ER doctor stitched me up and told me, while her hands were still in my neck, that whatever hit me was a centimeter or two from being fatal. Then she said she was going to call her son when she was done. He was around my age. Played hockey. She wanted to make sure he was ok.&lt;/p&gt;
&lt;p&gt;In the moment that was nothing. I was wearing hockey pants and a cartoon mummy wrap of tape and gauze, paying a forty-five dollar ambulance transfer fee, trying to find the cut with a selfie while waiting for a nurse. The cut was gaping — wound edges pulled wide apart, the kind that doesn&amp;rsquo;t close on its own — and even then it was hard to find in my beard. My brain was somewhere between shock and an unusual species of relief. I got stitched. Got discharged. My in-law picked me up outside a hospital I&amp;rsquo;d never been to, on a street I didn&amp;rsquo;t recognize.&lt;/p&gt;
&lt;p&gt;The part that landed later, once I&amp;rsquo;d told the story enough times to hear it out loud, was this: she finished her shift, went home, and called her son. Not because anything happened to him. Because something almost happened to me, a stranger, and she needed to hear his voice. That&amp;rsquo;s how close it was. Close enough that a doctor who sees bad nights for a living went home and made that call.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The helmet is still in my bag. Still perfectly usable. The puck found a gap the equipment didn&amp;rsquo;t account for, and the equipment came through fine.&lt;/p&gt;
&lt;p&gt;I added a dangler after. A promise to my wife. &lt;em&gt;Safer.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;It took me three weeks to get back on the ice and when I did, I was afraid of the right circle. Every time someone loaded up to shoot, my mouth hung open like it was waiting for something to happen. I couldn&amp;rsquo;t stop it. It took a while for that to go away.&lt;/p&gt;
&lt;p&gt;My friend waited up that night. I texted him from the parking lot at 1:30 that I was at my car. He was relieved. I drove to his place and we stood outside in the cold and dark and he told me about what happened after I left. Apparently nobody wanted to shoot on the other goalie for the rest of the game. They cut it short and went home. We stood there laughing about that, me in my skate shoes and stitches, him shaking his head at how quickly I&amp;rsquo;d gone from fine to gone and then somehow back to standing in his driveway at two in the morning looking more or less like a normal person.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the part that&amp;rsquo;s genuinely hard to sit with.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;I was holding my own neck on that bench. Applying pressure to stop the bleeding. There is a version of that night where I apply too much pressure, stop the blood to my own brain, and become a hockey fatality. A freak accident statistic at an eleven o&amp;rsquo;clock shinny game. A number in a report somewhere. The young guy in red spends years carrying a shot he didn&amp;rsquo;t do anything wrong on. My wife gets a phone call instead of a sleepy conversation in the dark.&lt;/p&gt;
&lt;p&gt;That version isn&amp;rsquo;t hypothetical. It was the same night, the same bench, the same hands.&lt;/p&gt;
&lt;p&gt;I don&amp;rsquo;t live with pain from this. No lasting damage, no real reminder except a scar that hides in my beard. I got out clean and drove myself home at two in the morning like it was nothing. Then I walked in the door and woke up my wife.&lt;/p&gt;
&lt;p&gt;She wanted to know I was ok. Then she wanted to know why I forgot to wear the right equipment, why I&amp;rsquo;d taken the car and left her stranded, how we were going to make sure this never happened again. All fair questions. The last one doesn&amp;rsquo;t really have an answer. When she&amp;rsquo;d run out of things to be furious about, she told me to wash up and come to bed. She&amp;rsquo;d look at the stitches in the morning.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve told the gory part of this story plenty. &lt;em&gt;Got cut playing hockey, here&amp;rsquo;s the bloody equipment.&lt;/em&gt; I&amp;rsquo;ve told it like it&amp;rsquo;s a war story, something to pull out for shock value.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s not what it is.&lt;/p&gt;
&lt;p&gt;I was the lucky one. The gap was there the whole time. I just didn&amp;rsquo;t know it until the puck found it.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Some Nights You Just Show Up</title>
      <link>https://blog.coffeeordeath.dev/posts/2026-01-18---some-nights-you-just-show-up/</link>
      <pubDate>Sun, 18 Jan 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/2026-01-18---some-nights-you-just-show-up/</guid>
      <description>Easy save. I could smother it, kill the play. Instead I kick it out to their other forward.
Not a mistake. A choice. I&amp;rsquo;m bored and my team&amp;rsquo;s up four goals and I want to make another save. So I manufacture chaos, create my own work, turn an easy night into something that feels like hockey.
That&amp;rsquo;s the first thirty-six minutes.
The second half, I&amp;rsquo;m making four saves in a row and watching the fifth one trickle past my pad anyway.</description>
      <content>&lt;p&gt;Easy save. I could smother it, kill the play. Instead I kick it out to their other forward.&lt;/p&gt;
&lt;p&gt;Not a mistake. A choice. I&amp;rsquo;m bored and my team&amp;rsquo;s up four goals and I want to make another save. So I manufacture chaos, create my own work, turn an easy night into something that feels like hockey.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the first thirty-six minutes.&lt;/p&gt;
&lt;p&gt;The second half, I&amp;rsquo;m making four saves in a row and watching the fifth one trickle past my pad anyway. My defense is skating like they&amp;rsquo;re underwater. Passes are dying on sticks. I&amp;rsquo;m doing everything right and we&amp;rsquo;re getting shelled.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the thing about being on-call: some weeks, everything you touch works. You&amp;rsquo;re rolling back that bad deploy before anyone notices. You&amp;rsquo;re catching the disk space issue at 87% instead of 99%. You&amp;rsquo;re so dialed in you start getting cocky, maybe let a minor alert sit an extra minute just to see if it self-heals, because you know you can fix it anyway.&lt;/p&gt;
&lt;p&gt;Other weeks, you&amp;rsquo;re making all the right calls and it doesn&amp;rsquo;t matter. The primary fails over cleanly and the secondary&amp;rsquo;s already degraded. You catch the memory leak but the restart triggers a cascade. You did the runbook perfectly and the system finds a new way to fail anyway.&lt;/p&gt;
&lt;p&gt;Both are the job.&lt;/p&gt;
&lt;p&gt;The dangerous part is thinking the first kind of week means you&amp;rsquo;re great at this, or the second kind means you&amp;rsquo;re not. Some nights the puck just goes in. Some nights your coverage is perfect and it still finds the gap. The only consistent thing is that you have to show up.&lt;/p&gt;
&lt;p&gt;You can&amp;rsquo;t manufacture chaos in production the way I do on ice - that&amp;rsquo;s how you end up in a post-mortem explaining why you thought it would be &amp;ldquo;fun&amp;rdquo; to test failover during peak traffic. But you also can&amp;rsquo;t let the brutal weeks convince you that you&amp;rsquo;re not doing it right.&lt;/p&gt;
&lt;p&gt;The shinny game is rec league hockey at eleven PM with a bunch of tired old men. Nobody&amp;rsquo;s keeping score except me. The second half, we lost 8-2.&lt;/p&gt;
&lt;p&gt;In production, someone&amp;rsquo;s always keeping score. But the principle holds: some nights you just show up, make the saves you can make, and skate off when it&amp;rsquo;s done.&lt;/p&gt;
&lt;p&gt;The job is to show up anyway.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>The Lie That Taught Me About Blameless Culture</title>
      <link>https://blog.coffeeordeath.dev/posts/the-lie-that-taught-me-about-blameless-culture/</link>
      <pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/the-lie-that-taught-me-about-blameless-culture/</guid>
      <description>I was standing outside a conference room watching my team lie to another team about a database outage. It was day four.
Through the glass door, I could see the engineer on the call, explaining with impressive confidence that our cloud provider was having issues. Any minute now, they said, the vendor would resolve it and services would come back up.
I pulled up the provider&amp;rsquo;s status page. Green across the board.</description>
      <content>&lt;p&gt;I was standing outside a conference room watching my team lie to another team about a database outage. It was day four.&lt;/p&gt;
&lt;p&gt;Through the glass door, I could see the engineer on the call, explaining with impressive confidence that our cloud provider was having issues. Any minute now, they said, the vendor would resolve it and services would come back up.&lt;/p&gt;
&lt;p&gt;I pulled up the provider&amp;rsquo;s status page. Green across the board. Checked their Twitter. Nothing.
I caught the eye of the engineer leading the response and motioned them outside.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;What actually happened?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The way they looked at the floor told me everything. Then: &amp;ldquo;Accidental deletion. We&amp;rsquo;re working on recovery.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the thing about inheriting a team: You don&amp;rsquo;t just inherit their code and their systems. You inherit their patterns. Their fears. The lessons they learned from whoever led them before you.&lt;/p&gt;
&lt;p&gt;My team had learned that mistakes meant consequences. That admitting fault meant exposure. That the safest play was to redirect blame to something faceless—a cloud provider, a vendor, a system &amp;ldquo;acting weird.&amp;rdquo; They&amp;rsquo;d learned this so well that lying to the team we supported felt safer than telling their new boss the truth.&lt;/p&gt;
&lt;p&gt;I had about thirty seconds to decide what kind of leader I was going to be.&lt;/p&gt;
&lt;h2 id=&#34;the-conversation&#34;&gt;The Conversation&lt;/h2&gt;
&lt;p&gt;We found an empty office.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I&amp;rsquo;m not mad about the database,&amp;rdquo; I said. &amp;ldquo;I&amp;rsquo;m concerned you felt you couldn&amp;rsquo;t tell me.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They stared at their hands.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I need you to understand something. I&amp;rsquo;m four days into this role. I don&amp;rsquo;t know your systems. I don&amp;rsquo;t know the history here. But I know this: I can&amp;rsquo;t help you fix problems I don&amp;rsquo;t know about. And the other team can&amp;rsquo;t tell us how severe this actually is if we&amp;rsquo;re feeding them fiction.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Previous leadership didn&amp;rsquo;t react well to mistakes.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;There it was. The inherited culture, delivered in one sentence.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Okay. Here&amp;rsquo;s the new standard: You tell me the truth, always. I&amp;rsquo;ll handle the hard conversations. You focus on the fix. Deal?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They looked up. &amp;ldquo;Deal.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;falling-on-the-sword&#34;&gt;Falling on the Sword&lt;/h2&gt;
&lt;p&gt;I walked back into that conference room.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I need to correct something,&amp;rdquo; I said. &amp;ldquo;This isn&amp;rsquo;t a cloud outage. One of our engineers accidentally deleted the pre-production database. We&amp;rsquo;re working on recovery now, but I need your help: How critical is this to your operations? What&amp;rsquo;s the actual impact?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The silence lasted maybe three seconds. Felt like an hour.&lt;/p&gt;
&lt;p&gt;Then: &amp;ldquo;Thank you for being straight with us. Here&amp;rsquo;s what we need&amp;hellip;&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They walked us through their priorities. Which services mattered most. What data they could recreate versus what needed restoration. The conversation became a collaboration instead of a performance.&lt;/p&gt;
&lt;p&gt;We got the database back. It took six hours and some creative recovery work, but we got it back.&lt;/p&gt;
&lt;h2 id=&#34;the-post-mortem&#34;&gt;The Post-Mortem&lt;/h2&gt;
&lt;p&gt;The next day, I gathered the team in the conference room for our first post-mortem. I&amp;rsquo;d never run one with them before, so I started by explaining how we were going to do them from now on.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;When we name people in post-mortems,&amp;rdquo; I said, &amp;ldquo;it&amp;rsquo;s only for successes. Who responded quickly. Who had the knowledge to guide recovery. Who communicated clearly under pressure.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I looked at the engineer who&amp;rsquo;d led the recovery effort. &amp;ldquo;You did excellent work getting us back online.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&amp;ldquo;For failures, the buck stops with me. I&amp;rsquo;m the lead. Anything that goes wrong happened on my watch. My job is to figure out what systemic failures allowed this to happen and fix them.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I watched the team process this. Some looked skeptical. Some looked relieved. All of them looked like they were waiting for the other shoe to drop.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Here&amp;rsquo;s what we&amp;rsquo;re changing. Permissions on production-like environments. Backup verification procedures. Documentation on recovery processes. And most importantly: the expectation that you can always tell me the truth, especially when things break.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;the-timing&#34;&gt;The Timing&lt;/h2&gt;
&lt;p&gt;Looking back, week one was actually the perfect time for this to happen.&lt;/p&gt;
&lt;p&gt;I had no history with this team. No pattern of reactions for them to predict. No baggage about &amp;ldquo;how we&amp;rsquo;ve always done things.&amp;rdquo; When I said &amp;ldquo;this is the new standard,&amp;rdquo; there was no previous version of me to contradict it.&lt;/p&gt;
&lt;p&gt;Being new meant I had nothing to lose and everything to prove. I couldn&amp;rsquo;t lean on established trust or authority. I had to show them, in real time, what kind of leader I was going to be.&lt;/p&gt;
&lt;p&gt;Culture isn&amp;rsquo;t built during the smooth times. It&amp;rsquo;s built in how you respond when things break. And things always break.&lt;/p&gt;
&lt;h2 id=&#34;what-stuck&#34;&gt;What Stuck&lt;/h2&gt;
&lt;p&gt;That engineer never lied to me again. Neither did anyone else on the team. Not because I&amp;rsquo;d scared them straight, but because I&amp;rsquo;d shown them the alternative: honesty gets you help, not punishment.&lt;/p&gt;
&lt;p&gt;The other team? They became one of our strongest advocates. Years later, they still reference that incident as the moment they knew they could trust us.&lt;/p&gt;
&lt;p&gt;The post-mortem approach stuck too. Name people for successes, take responsibility for failures. It sounds simple, but it changes everything about how teams respond to pressure.&lt;/p&gt;
&lt;p&gt;I couldn&amp;rsquo;t change what my team had learned before I arrived. I couldn&amp;rsquo;t undo whatever had taught them that self-preservation trumped transparency. But I could show them a different future, starting with one deleted database and one honest conversation.&lt;/p&gt;
&lt;p&gt;The bad news: People will test your principles during crisis. The good news: Crisis is when principles matter most.&lt;/p&gt;
&lt;p&gt;The job is to be ready anyway.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Everything I learned about Ops, I learned playing goal</title>
      <link>https://blog.coffeeordeath.dev/posts/everything-i-learned-about-ops-i-learned-playing-goal/</link>
      <pubDate>Thu, 06 Nov 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/everything-i-learned-about-ops-i-learned-playing-goal/</guid>
      <description>Well, maybe not everything.
Look, being a goalie is objectively ridiculous. You strap on forty pounds of equipment designed to protect you from frozen rubber traveling at speeds that would make physicists frown. Then you stand in front of a net and dare people to shoot at you. It&amp;rsquo;s a strange job.
But here&amp;rsquo;s the thing: being a goalie is basically the same job as running ops, or security, or honestly, any part of software development where you&amp;rsquo;re the one who has to keep the thing from breaking.</description>
      <content>&lt;p&gt;Well, maybe &lt;strong&gt;not everything&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Look, being a goalie is objectively ridiculous. You strap on forty pounds of equipment designed to protect you from frozen rubber traveling at speeds that would make physicists frown. Then you stand in front of a net and dare people to shoot at you. It&amp;rsquo;s a strange job.&lt;/p&gt;
&lt;p&gt;But here&amp;rsquo;s the thing: being a goalie is basically the same job as running ops, or security, or honestly, any part of software development where you&amp;rsquo;re the one who has to keep the thing from breaking.&lt;/p&gt;
&lt;p&gt;Both jobs share the same core truth that nobody tells you in the job description: when everything&amp;rsquo;s working, no one notices you&amp;rsquo;re there. But the &lt;em&gt;second&lt;/em&gt; something gets past you? Everyone knows exactly where you were and what you should have done differently.&lt;/p&gt;
&lt;p&gt;You spend most of your time reading patterns, anticipating problems before they fully materialize, and communicating with people who may or may not be listening. You&amp;rsquo;re simultaneously trying to follow the systems that keep you consistent while also staying loose enough to react when something weird happens. Because something weird &lt;em&gt;always&lt;/em&gt; happens.&lt;/p&gt;
&lt;p&gt;And whether it&amp;rsquo;s a breakaway in overtime or a production incident at 2 AM, you learn pretty quickly that freezing up isn&amp;rsquo;t an option. You make the save or you don&amp;rsquo;t. You stop the breach or you don&amp;rsquo;t. Then you&amp;rsquo;ve got about eight seconds to shake it off and get ready for the next shot, because there&amp;rsquo;s always a next shot.&lt;/p&gt;
&lt;p&gt;The bad news? You can&amp;rsquo;t stop everything. The good news? Neither can anyone else. The job is to be ready anyway.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Decoupling Deployment from Release: An Architectural Imperative</title>
      <link>https://blog.coffeeordeath.dev/posts/decoupling-deployment-from-release/</link>
      <pubDate>Tue, 30 Sep 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/decoupling-deployment-from-release/</guid>
      <description>Photo by Patrick Konior on Unsplash
We need to talk about how we ship software. Not because we&amp;rsquo;re doing it wrong, but because we can do it so much better, for our customers, for our colleagues, and for ourselves.
The Core Problem: Deployment Isn&amp;rsquo;t Release Here&amp;rsquo;s the shift we need to make: deploying code and releasing features are not the same thing, and they shouldn&amp;rsquo;t happen at the same moment.</description>
      <content>&lt;p&gt;&lt;img alt=&#34;Control room overview&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/hero-control-room.jpg&#34;&gt;
&lt;em&gt;Photo by &lt;a href=&#34;https://unsplash.com/@patrickkonior&#34;&gt;Patrick Konior&lt;/a&gt; on &lt;a href=&#34;https://unsplash.com/photos/a-view-of-a-control-room-from-above-FwHJ6mJjVd4&#34;&gt;Unsplash&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We need to talk about how we ship software. Not because we&amp;rsquo;re doing it wrong, but because we can do it so much better, for our customers, for our colleagues, and for ourselves.&lt;/p&gt;
&lt;h2 id=&#34;the-core-problem-deployment-isnt-release&#34;&gt;The Core Problem: Deployment Isn&amp;rsquo;t Release&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s the shift we need to make: &lt;strong&gt;deploying code and releasing features are not the same thing, and they shouldn&amp;rsquo;t happen at the same moment.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&#34;Deployment vs. Release Timeline diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/deployment-vs-release-timeline.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;When we conflate deployment with release, we create unnecessary risk and friction. We deploy code to production, and suddenly a feature is live, ready or not. If something goes wrong, we scramble. We hotfix. We roll back entire deployments, potentially affecting unrelated changes. We make our product and support teams reactive instead of strategic.&lt;/p&gt;
&lt;p&gt;There is a better way.&lt;/p&gt;
&lt;h2 id=&#34;why-this-matters-safety-first&#34;&gt;Why This Matters: Safety First&lt;/h2&gt;
&lt;p&gt;Our customers trust us with their business. Our colleagues across product, support, and operations depend on stable, predictable releases. When we tie deployment directly to feature availability, we put both at risk.&lt;/p&gt;
&lt;p&gt;Safe software releases mean:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Our customers experience fewer disruptions&lt;/li&gt;
&lt;li&gt;Our support team can prepare for changes before they go live&lt;/li&gt;
&lt;li&gt;Our product team can control the narrative and timing of new features&lt;/li&gt;
&lt;li&gt;Our engineers can deploy confidently, knowing they have safety nets in place&lt;/li&gt;
&lt;li&gt;We can see the impact of our changes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This isn&amp;rsquo;t about adding bureaucracy. It&amp;rsquo;s about adding control.&lt;/p&gt;
&lt;h2 id=&#34;the-mindset-shift-for-engineers&#34;&gt;The Mindset Shift for Engineers&lt;/h2&gt;
&lt;p&gt;I know what some of you are thinking: &amp;ldquo;This sounds like more process. More overhead. More barriers between my code and production.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Let me reframe it: What if deployment was &lt;strong&gt;easier&lt;/strong&gt; because it was &lt;strong&gt;safer&lt;/strong&gt;? What if you could push code to production multiple times a day without the anxiety of &amp;ldquo;did we just break something for customers?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;This requires us to think differently:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Code goes to production &lt;strong&gt;dark&lt;/strong&gt;, deployed but not activated&lt;/li&gt;
&lt;li&gt;Features are controlled by flags, not by deployment timing&lt;/li&gt;
&lt;li&gt;Your code can live in production for days or weeks before anyone uses it&lt;/li&gt;
&lt;li&gt;When issues arise, you flip a switch instead of rolling back a deployment&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&#34;Dark launch diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/dark-launch.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;This is actually &lt;strong&gt;more&lt;/strong&gt; engineering control, not less. You own the technical deployment. Product and support teams own the feature release timing. Everyone wins.&lt;/p&gt;
&lt;h2 id=&#34;empowering-product-and-support&#34;&gt;Empowering Product and Support&lt;/h2&gt;
&lt;p&gt;Our product managers understand customer needs, market timing, and strategic rollout. Our support team knows when they&amp;rsquo;re ready to handle inquiries about new features. Our documentation should be created in production.&lt;/p&gt;
&lt;p&gt;When we give them control over feature flags, we put release decisions in the hands of the people best positioned to make them. Engineering provides the capability; product and support control the activation.&lt;/p&gt;
&lt;p&gt;This isn&amp;rsquo;t about taking power away from engineering, it&amp;rsquo;s about giving everyone the right kind of power.&lt;/p&gt;
&lt;h2 id=&#34;stop-removing-what-people-still-use&#34;&gt;Stop Removing What People Still Use&lt;/h2&gt;
&lt;p&gt;We&amp;rsquo;ve all been there: a feature is marked for deprecation, so we rip it out. Then we discover customers were still using it, or they needed more transition time.&lt;/p&gt;
&lt;p&gt;With feature flags and dual code paths, we can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Run old and new implementations side by side&lt;/li&gt;
&lt;li&gt;Give customers time to migrate naturally&lt;/li&gt;
&lt;li&gt;Gather actual usage data before removal&lt;/li&gt;
&lt;li&gt;Roll back feature changes without rolling back deployments&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We keep both paths alive until we &lt;strong&gt;know&lt;/strong&gt; the old one isn&amp;rsquo;t needed. Data drives decisions, not assumptions.&lt;/p&gt;
&lt;h2 id=&#34;creating-dual-code-paths&#34;&gt;Creating Dual Code Paths&lt;/h2&gt;
&lt;p&gt;The technical pattern is straightforward:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;if (featureFlag.isEnabled(&amp;#39;new-checkout-flow&amp;#39;)) {
  // New implementation
  return enhancedCheckout();
} else {
  // Current implementation
  return traditionalCheckout();
}
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;img alt=&#34;Dual code paths diagram&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/dual-code-paths.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;This enables:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Rapid feature releases, toggle a flag, feature goes live&lt;/li&gt;
&lt;li&gt;Instant rollbacks, toggle it back if issues arise&lt;/li&gt;
&lt;li&gt;Gradual rollouts, enable for 10% of users, then 50%, then 100%&lt;/li&gt;
&lt;li&gt;A/B testing, compare new vs. old with real data&lt;/li&gt;
&lt;li&gt;Safe deprecation, maintain the old path until usage drops to zero&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&#34;Rollback Comparison&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/rollback-comparison.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img alt=&#34;Progressive Rollout&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/progressive-rollout.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;Yes, you maintain two paths temporarily. But the flexibility and safety are worth it.&lt;/p&gt;
&lt;h2 id=&#34;we-already-test-in-production-lets-do-it-right&#34;&gt;We Already Test in Production, Let&amp;rsquo;s Do It Right&lt;/h2&gt;
&lt;p&gt;Here&amp;rsquo;s an uncomfortable truth: we all test in production. Even with the best staging environments, production has data patterns, load characteristics, and edge cases we can&amp;rsquo;t fully replicate.&lt;/p&gt;
&lt;p&gt;Feature flags let us test in production &lt;strong&gt;safely&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Enable new features for internal users first&lt;/li&gt;
&lt;li&gt;Roll out to beta customers who opt in&lt;/li&gt;
&lt;li&gt;Gradually expand to broader audiences&lt;/li&gt;
&lt;li&gt;Monitor in real time and react instantly&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Instead of pretending we don&amp;rsquo;t test in production, let&amp;rsquo;s acknowledge it and build systems that make it safe and controlled.&lt;/p&gt;
&lt;h2 id=&#34;making-monitoring-and-observability-easier&#34;&gt;Making Monitoring and Observability Easier&lt;/h2&gt;
&lt;p&gt;When features are flag-controlled, monitoring becomes clearer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Dashboard metrics tagged by feature flag status&lt;/li&gt;
&lt;li&gt;Alerts that distinguish between feature issues and infrastructure issues&lt;/li&gt;
&lt;li&gt;A/B comparison of performance between flag states&lt;/li&gt;
&lt;li&gt;Clear correlation between flag changes and system behaviour&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&#34;Observability Dashboard&#34; src=&#34;https://blog.coffeeordeath.dev/images/decoupling-deployment/observability-dashboard.svg&#34;&gt;&lt;/p&gt;
&lt;p&gt;You&amp;rsquo;ll spend less time debugging &amp;ldquo;what changed?&amp;rdquo; because you&amp;rsquo;ll know exactly what changed and when.&lt;/p&gt;
&lt;h2 id=&#34;this-is-about-your-value&#34;&gt;This Is About Your Value&lt;/h2&gt;
&lt;p&gt;Every engineer brings immense value. You solve complex problems. You build systems that serve customers. You keep the platform running.&lt;/p&gt;
&lt;p&gt;This approach amplifies your impact:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deploy more frequently without stress&lt;/li&gt;
&lt;li&gt;Spend less time on hotfixes and emergency rollbacks&lt;/li&gt;
&lt;li&gt;Focus on building new capabilities instead of managing release logistics&lt;/li&gt;
&lt;li&gt;See your features succeed because they&amp;rsquo;re released at the right time, to the right users&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Your technical expertise deserves better tools. Feature-flagged releases are those better tools.&lt;/p&gt;
&lt;h2 id=&#34;getting-started&#34;&gt;Getting Started&lt;/h2&gt;
&lt;p&gt;We don&amp;rsquo;t need to transform everything overnight. Here&amp;rsquo;s how we can begin:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Start with your next major feature&lt;/strong&gt;, wrap it in a feature flag&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Integrate flags into your monitoring&lt;/strong&gt;, make flag status visible in dashboards&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coordinate with product and support&lt;/strong&gt;, let them control release timing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gather data before deprecation&lt;/strong&gt;, keep old code paths until usage proves they&amp;rsquo;re unnecessary&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Share what works&lt;/strong&gt;, as teams succeed with this approach, spread the practices&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is a journey, not a destination. Each team can adopt at their own pace.&lt;/p&gt;
&lt;h2 id=&#34;the-bottom-line&#34;&gt;The Bottom Line&lt;/h2&gt;
&lt;p&gt;Separating deployment from release isn&amp;rsquo;t just a technical practice, it is a philosophy of safety, control, and collaboration. It acknowledges that great software delivery involves engineering excellence &lt;strong&gt;and&lt;/strong&gt; thoughtful release management.&lt;/p&gt;
&lt;p&gt;We have the talent. We have the technical capability. Now let&amp;rsquo;s build the practices that let us ship with confidence.&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s deploy fearlessly and release deliberately. Our customers, our colleagues, and our code deserve nothing less.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>ZFS Raid Types</title>
      <link>https://blog.coffeeordeath.dev/posts/zfs-raid-types/</link>
      <pubDate>Mon, 21 Apr 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/zfs-raid-types/</guid>
      <description>If I didn&amp;rsquo;t have to spend a bunch of time moving virtual machines around, this probably wouldn&amp;rsquo;t matter very much. But, alas, I did find myself building zfs raid arrays a few times and couldn&amp;rsquo;t seem to remember what I wanted where.
🧂 Take with a conservative grain of salt.
RAID Type Min Disks Usable Capacity Fault Tolerance Performance Notes Notes Striped 1+ 100% of total None Fastest read/write, no redundancy Equivalent to RAID0 Mirror 2+ 50% of total 1 disk per mirror vdev Excellent read, good write, fast recovery Equivalent to RAID1 in traditional RAID RAIDZ1 3+ N - 1 1 disk failure Slower write, decent read Similar to RAID5 RAIDZ2 4+ N - 2 2 disk failures Slower write, decent read Similar to RAID6 RAIDZ3 5+ N - 3 3 disk failures Slowest write, decent read For high fault tolerance ZFS RAID10 4+ (even) 50% of total 1 disk per mirror vdev Best balance of performance + redundancy Stripe of mirrors (manual mirror vdevs) dRAID 3+ Varies (RAIDZ-like) Parity-based (configurable) Improved resilver vs.</description>
      <content>&lt;p&gt;If I didn&amp;rsquo;t have to spend a bunch of time moving virtual machines around, this probably wouldn&amp;rsquo;t matter very much. But, alas, I did find myself building zfs raid arrays a few times and couldn&amp;rsquo;t seem to remember what I wanted where.&lt;/p&gt;
&lt;p&gt;🧂 Take with a conservative grain of salt.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;RAID Type&lt;/th&gt;
&lt;th&gt;Min Disks&lt;/th&gt;
&lt;th&gt;Usable Capacity&lt;/th&gt;
&lt;th&gt;Fault Tolerance&lt;/th&gt;
&lt;th&gt;Performance Notes&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Striped&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1+&lt;/td&gt;
&lt;td&gt;100% of total&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Fastest read/write, no redundancy&lt;/td&gt;
&lt;td&gt;Equivalent to RAID0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mirror&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2+&lt;/td&gt;
&lt;td&gt;50% of total&lt;/td&gt;
&lt;td&gt;1 disk per mirror vdev&lt;/td&gt;
&lt;td&gt;Excellent read, good write, fast recovery&lt;/td&gt;
&lt;td&gt;Equivalent to RAID1 in traditional RAID&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RAIDZ1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3+&lt;/td&gt;
&lt;td&gt;N - 1&lt;/td&gt;
&lt;td&gt;1 disk failure&lt;/td&gt;
&lt;td&gt;Slower write, decent read&lt;/td&gt;
&lt;td&gt;Similar to RAID5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RAIDZ2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4+&lt;/td&gt;
&lt;td&gt;N - 2&lt;/td&gt;
&lt;td&gt;2 disk failures&lt;/td&gt;
&lt;td&gt;Slower write, decent read&lt;/td&gt;
&lt;td&gt;Similar to RAID6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;RAIDZ3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5+&lt;/td&gt;
&lt;td&gt;N - 3&lt;/td&gt;
&lt;td&gt;3 disk failures&lt;/td&gt;
&lt;td&gt;Slowest write, decent read&lt;/td&gt;
&lt;td&gt;For high fault tolerance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ZFS RAID10&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4+ (even)&lt;/td&gt;
&lt;td&gt;50% of total&lt;/td&gt;
&lt;td&gt;1 disk per mirror vdev&lt;/td&gt;
&lt;td&gt;Best balance of performance + redundancy&lt;/td&gt;
&lt;td&gt;Stripe of mirrors (manual mirror vdevs)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;dRAID&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3+&lt;/td&gt;
&lt;td&gt;Varies (RAIDZ-like)&lt;/td&gt;
&lt;td&gt;Parity-based (configurable)&lt;/td&gt;
&lt;td&gt;Improved resilver vs. RAIDZ, good balance&lt;/td&gt;
&lt;td&gt;Requires newer ZFS versions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Special VDEV&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1+&lt;/td&gt;
&lt;td&gt;Not for main storage&lt;/td&gt;
&lt;td&gt;Depends on layout&lt;/td&gt;
&lt;td&gt;High IOPS for metadata/small files&lt;/td&gt;
&lt;td&gt;Used for performance, not redundancy&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;blockquote&gt;
&lt;p&gt;N = Number of disks in the vdev&lt;/p&gt;
&lt;/blockquote&gt;
</content>
    </item>
    
    <item>
      <title>New Team, New Processes</title>
      <link>https://blog.coffeeordeath.dev/posts/new-team-new-processes/</link>
      <pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/new-team-new-processes/</guid>
      <description>Starting fresh with a new team—whether it&amp;rsquo;s stepping into the net as a goalie or joining a company as a Site Reliability Engineer—comes with a rush of excitement, uncertainty, and the need to quickly adapt. In both worlds, you&amp;rsquo;re expected to understand the system, earn trust fast, and make the right decisions under pressure. This piece explores the parallels between guarding the crease in ice hockey and taking on infrastructure responsibilities in a new engineering org: the importance of communication, learning team dynamics, managing risk, and building confidence through early wins.</description>
      <content>&lt;p&gt;Starting fresh with a new team—whether it&amp;rsquo;s stepping into the net as a goalie or joining a company as a Site Reliability Engineer—comes with a rush of excitement, uncertainty, and the need to quickly adapt. In both worlds, you&amp;rsquo;re expected to understand the system, earn trust fast, and make the right decisions under pressure. This piece explores the parallels between guarding the crease in ice hockey and taking on infrastructure responsibilities in a new engineering org: the importance of communication, learning team dynamics, managing risk, and building confidence through early wins.&lt;/p&gt;
&lt;p&gt;When you join a new hockey team as a goalie, you don’t just protect the net—you learn how the defensemen play, how aggressive the forwards are, and when the team tends to collapse or stretch the ice. Every team has its own rhythm and unspoken rules.
The same is true when joining a new company as an SRE. The architecture is only part of the story; the real learning curve is understanding how incidents are handled, who actually makes the calls in a crunch, and which “informal” channels carry the most context. Just like a goalie can’t force their style on a team that plays differently, a new SRE needs to observe, adapt, and gradually find ways to complement and strengthen the existing flow. It&amp;rsquo;s less about proving you’re technically sound (that’s assumed) and more about showing that you can read the play and make the team better as a result.&lt;/p&gt;
&lt;p&gt;In both hockey and SRE work, the moment things go sideways is when communication really counts. As a goalie, you see the whole ice—you’re the only one facing the play—and you need to call out shifts, loose players, or breakdowns in real time. If you hesitate, the puck’s in the net.
As an SRE, it’s similar during an incident: you often have a high-level view of what’s happening across systems, and the ability to stay calm and clearly relay what you’re seeing can make the difference between a contained issue and a full-blown outage. But communication isn’t just about shouting the loudest—it’s about knowing when to speak, who needs to hear what, and doing it in a way that builds confidence rather than panic. Like a goalie steadying the team during a penalty kill, an SRE who brings calm clarity in high-stakes moments becomes an anchor for everyone else to rally around.&lt;/p&gt;
&lt;h3 id=&#34;so-how-do-you-build-trust-in-a-new-team&#34;&gt;So, how do you build trust in a new team?&lt;/h3&gt;
&lt;p&gt;Trust isn’t earned by pretending to be perfect—it’s built by showing up consistently, being accountable, and owning your impact, good or bad. As a goalie, you’re going to let in goals. Some are unstoppable, some are a result of defensive breakdowns, and some are just flat-out on you.
What matters is how you handle it. Do you sulk? Blame? Or do you skate over, tap a teammate’s shin pads, and reset for the next faceoff? The same dynamic plays out in SRE. Early on, you’ll miss signals, ship a bad change, or overlook something you &lt;em&gt;should&lt;/em&gt; have caught. Heck, there are days where you&amp;rsquo;ll break production.
But by being transparent, acknowledging the miss, and showing what you’re doing to improve, you demonstrate humility and build credibility. Teams trust people who are honest about their fallibility and show a pattern of learning. Whether it’s a soft goal from the point or a mistuned alert that paged the team at 3 a.m., owning it builds more goodwill than perfection ever could. That openness also invites others to bring their guard down, leading to better collaboration and stronger long-term bonds.&lt;/p&gt;
&lt;p&gt;As both a goalie and a leader, you quickly learn that most mistakes aren’t rooted in malice or incompetence—they’re byproducts of unclear systems, missing context, or just the fast pace of the game. When someone misses an assignment on the ice or forgets to update a runbook, it’s rarely because they don’t care—it’s usually because the process failed them. As an SRE leader, it’s your job to zoom out after the fact and ask not just &lt;em&gt;who&lt;/em&gt; made the mistake, but &lt;em&gt;why&lt;/em&gt; it made sense in that moment—and how the system can be improved so the next person has a better shot. Maybe the alert was noisy, the handoff lacked key context, or a dependency wasn’t clearly documented. Whatever it is, you shift the focus from blame to learning. Like a goalie giving quiet guidance during a timeout, your job is to reset the team without undercutting their confidence. Owning your own mistakes gives you the credibility to address others with empathy—and turning those moments into durable process improvements is how teams get stronger over time.&lt;/p&gt;
&lt;p&gt;Whether you&amp;rsquo;re stepping onto the ice or into a new engineering org, the fundamentals are the same: learn the team, communicate clearly under pressure, take ownership, and lead with empathy. Both roles are high-trust, high-impact, and deeply human. You’re not just reacting—you’re anticipating, supporting, and constantly adjusting to help the people around you succeed. The best teams don’t expect perfection—they expect presence, accountability, and a shared commitment to get better together. In the end, it&amp;rsquo;s not about never letting a goal in or preventing every outage. It&amp;rsquo;s about being the kind of teammate who learns fast, shows up when it matters, and helps turn every setback into forward motion.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Goalie as a Service</title>
      <link>https://blog.coffeeordeath.dev/posts/goalie-as-a-service/</link>
      <pubDate>Wed, 19 Mar 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/goalie-as-a-service/</guid>
      <description>I remember the moment vividly—a hard shot struck the lower side of my mask. The impact was sharp, a bolt of pain radiating through my face. I dropped my head to the ice, letting the sting settle before pushing myself up onto my knees. It wasn’t until I stretched my neck, looking around, that the blood started flowing. When I glanced down, the realization hit—red pooling rapidly beneath me. Instinct took over.</description>
      <content>&lt;p&gt;I remember the moment vividly—a hard shot struck the lower side of my mask. The impact was sharp, a bolt of pain radiating through my face. I dropped my head to the ice, letting the sting settle before pushing myself up onto my knees. It wasn’t until I stretched my neck, looking around, that the blood started flowing. When I glanced down, the realization hit—red pooling rapidly beneath me. Instinct took over. I stripped off my blocker, glove, and helmet, heading straight for the dressing room. Players were already scrambling for medical supplies as I pressed a hand to my face, trying to assess the damage and scrambling to find the source of the blood.&lt;/p&gt;
&lt;p&gt;In that moment, I wasn’t thinking about the pain. I was thinking about getting back on my feet, assessing the damage, and figuring out my next move—just like in an on-call incident.&lt;/p&gt;
&lt;p&gt;Being an amateur hockey goalie mirrors the world of DevOps and Site Reliability Engineering (SRE) in ways that only those who’ve been through the chaos can truly appreciate. Both demand the ability to anticipate disasters, react with composure under pressure, and recover quickly when everything goes sideways. A goalie can’t afford to dwell on the last goal any more than an SRE can get stuck in post-mortem blame. The job is about resilience, adaptability, and learning from every incident—whether it’s a slapshot to the mask or a production outage at peak traffic.&lt;/p&gt;
</content>
    </item>
    
  </channel>
</rss>
