<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>SRE on Coffee or Blog</title>
    <link>https://blog.coffeeordeath.dev/tags/sre/</link>
    <description>Recent content in SRE on Coffee or Blog</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Mon, 23 Mar 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.coffeeordeath.dev/tags/sre/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Why Doesn&#39;t Anyone Own That?</title>
      <link>https://blog.coffeeordeath.dev/posts/why-doesnt-anyone-own-that/</link>
      <pubDate>Mon, 23 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/why-doesnt-anyone-own-that/</guid>
      <description>It&amp;rsquo;s 5 PM on a Thursday. Something broke mid-morning and your team has been grinding on it for hours. Good engineers, working hard, coming up empty because the system is poorly documented and the people who built it have either moved on or moved teams. You&amp;rsquo;ve burned most of the day and you&amp;rsquo;re no closer to resolution.
So you escalate. You ping the leads. You ask in the channel where someone, surely, knows this service well enough to point you in the right direction.</description>
      <content>&lt;p&gt;It&amp;rsquo;s 5 PM on a Thursday. Something broke mid-morning and your team has been grinding on it for hours. Good engineers, working hard, coming up empty because the system is poorly documented and the people who built it have either moved on or moved teams. You&amp;rsquo;ve burned most of the day and you&amp;rsquo;re no closer to resolution.&lt;/p&gt;
&lt;p&gt;So you escalate. You ping the leads. You ask in the channel where someone, surely, knows this service well enough to point you in the right direction.&lt;/p&gt;
&lt;p&gt;Silence.&lt;/p&gt;
&lt;p&gt;Not because nobody cares. Because nobody actually owns it. There&amp;rsquo;s a name in the CMDB, but that person will tell you they haven&amp;rsquo;t touched it since a reorg reshuffled their priorities. There&amp;rsquo;s a runbook, but it describes a version of the system that no longer exists. There&amp;rsquo;s a Slack channel with forty members and no clear authority.&lt;/p&gt;
&lt;p&gt;This is where the real cost of diffuse ownership lives. Not in the architecture review, not in the postmortem writeup. Right here, in a live incident, with a team that has already spent most of a day spinning their wheels and a business that is losing patience.&lt;/p&gt;
&lt;p&gt;The easy diagnosis is that engineers don&amp;rsquo;t want ownership. Nobody wants the accountability when something critical falls over. That read is intuitive, and it&amp;rsquo;s mostly wrong.&lt;/p&gt;
&lt;p&gt;Most engineers I&amp;rsquo;ve worked with want to own something. They want the stakes, the identity, the sense that they know a system cold and it shows. Ownership is a fundamental motivator in this work. It&amp;rsquo;s one of the few things that makes the job feel like more than ticket throughput.&lt;/p&gt;
&lt;p&gt;The problem isn&amp;rsquo;t appetite. It&amp;rsquo;s that we&amp;rsquo;ve built organizations that punish people for taking it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;we-say-we-want-owners-then-we-punish-them&#34;&gt;We say we want owners. Then we punish them.&lt;/h2&gt;
&lt;p&gt;Most performance systems weren&amp;rsquo;t designed to reward operational excellence. They were designed to measure visible output — features shipped, projects delivered, things a manager can point to in a calibration meeting and defend. Keeping a critical service healthy produces none of that. No launch. No demo. No ribbon cutting.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s what it does produce: a rating of &amp;ldquo;meets expectations.&amp;rdquo; Which, in most four-point HR scales, means no raise. Sounds fair until you consider that the engineer just spent a quarter instrumented to the gills on a service that didn&amp;rsquo;t go down, pushed for refactors the team had been deferring for two years, and responded to every incident that came their way. They did their job flawlessly. The organization scored it as ordinary because ordinary is the only category it fits in.&lt;/p&gt;
&lt;p&gt;That engineer is not doing that again next year. Neither is anyone who watched it happen.&lt;/p&gt;
&lt;p&gt;The performance review is just the most visible mechanism. The same signal gets sent in sprint planning when operational work gets bumped for roadmap items, in reorgs when the &amp;ldquo;boring&amp;rdquo; platform work gets deprioritized, in how leaders talk about the team that keeps things running versus the team shipping the new thing. It accumulates. People are paying attention even when you think they aren&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;Rational engineers respond rationally. Ownership becomes a liability, ambiguity becomes protection, and the next time someone asks who owns the broken thing, there genuinely isn&amp;rsquo;t an answer.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;accountability-without-authority-is-just-blame&#34;&gt;Accountability without authority is just blame.&lt;/h2&gt;
&lt;p&gt;There&amp;rsquo;s a second trap that even well-intentioned leaders fall into, and it&amp;rsquo;s more insidious than the performance review problem because it often comes wrapped in the language of empowerment.&lt;/p&gt;
&lt;p&gt;We assign ownership without giving people the actual power to act.&lt;/p&gt;
&lt;p&gt;Think about what that looks like in practice. An engineer is named owner of a critical platform service. They identify that the deployment process is bypassing validation steps and creating instability. They flag it. They write it up. They ask for a two-week pause on releases to address it. The response from leadership: &amp;ldquo;We can&amp;rsquo;t stop the roadmap, let&amp;rsquo;s find another way.&amp;rdquo; There is no other way. The service degrades, falls over, and in the postmortem the question on the table is why the service owner didn&amp;rsquo;t catch this sooner. It plays out constantly, in slightly different costumes, and the person in the owner role learns the same lesson every time.&lt;/p&gt;
&lt;p&gt;Ownership without the authority to say no to a risky deployment, to pull capacity from feature work when a system is genuinely at risk, to set a standard and enforce it — that isn&amp;rsquo;t ownership. It&amp;rsquo;s a title with no teeth. The person holding it gets the pager, gets the postmortem questions, gets the performance conversation, and had none of the leverage required to prevent any of it. Smart engineers recognize that setup fast. Word travels.&lt;/p&gt;
&lt;p&gt;The organizations that do this aren&amp;rsquo;t malicious. Most of them genuinely believe they&amp;rsquo;ve created clear ownership. They put a name in a doc and called it done. What they actually created is a designated scapegoat with a service catalog entry.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;the-cost-is-higher-than-you-think&#34;&gt;The cost is higher than you think.&lt;/h2&gt;
&lt;p&gt;Diffuse ownership doesn&amp;rsquo;t just create operational risk, though it absolutely does that. It degrades the institutional knowledge your organization depends on.&lt;/p&gt;
&lt;p&gt;When nobody owns a system, nobody deeply understands it. Documentation becomes aspirational fiction. Runbooks describe what someone &lt;em&gt;intended&lt;/em&gt; the system to do, not what it actually does. Incidents take longer, cost more, and teach less, because there&amp;rsquo;s no through-line of accountability that connects cause to effect to learning.&lt;/p&gt;
&lt;p&gt;And the talent impact is brutal. The engineers who want to own things, the ones you most want to retain, will eventually get tired of operating in that vacuum. They&amp;rsquo;ll go somewhere that lets them. The engineers who remain will have self-selected for comfort with ambiguity and low accountability. That is a culture you cannot easily reverse.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;what-actually-changes-this&#34;&gt;What actually changes this.&lt;/h2&gt;
&lt;p&gt;This isn&amp;rsquo;t a process problem. You can&amp;rsquo;t RACI-chart your way out of a culture that punishes ownership. You have to change what you reward.&lt;/p&gt;
&lt;p&gt;Start with how you evaluate performance. If your engineering managers cannot point to explicit, visible credit given to engineers for operational ownership — in promotion cases, in calibration sessions, in public recognition — then you are implicitly telling your team that it doesn&amp;rsquo;t count. Fix that first.&lt;/p&gt;
&lt;p&gt;Next, make ownership legible. A service with a named owner, a clear charter, and explicit authority boundaries is a service people will actually step up to own. Ambiguity is where accountability goes to die. Define the thing. Name the owner. Give them real power over their domain.&lt;/p&gt;
&lt;p&gt;Toyota figured this out on the factory floor decades ago. &lt;a href=&#34;https://en.wikipedia.org/wiki/Andon_(manufacturing)&#34;&gt;Their Andon&lt;/a&gt; system gives any line worker the authority to stop production the moment they identify a problem — not escalate it, not flag it for later, stop it. The entire premise is that the person closest to the issue has both the standing and the backing to act. It works because leadership built a culture where pulling that cord is the right call, not a career risk. We&amp;rsquo;re supposedly more sophisticated in software, yet we routinely put engineers in positions where they can see the problem clearly, have no authority to stop anything, and get blamed when it blows up anyway.&lt;/p&gt;
&lt;p&gt;Back your owners when it&amp;rsquo;s hard. When someone pulls the cord on a deployment, or pushes back on a timeline because the system genuinely can&amp;rsquo;t absorb it, support that call publicly. If you quietly override it three times, you&amp;rsquo;ve taught everyone watching that ownership is theater.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;the-question-worth-asking&#34;&gt;The question worth asking.&lt;/h2&gt;
&lt;p&gt;The next time you find yourself in that room, the incident is running, the silence is thick, and nobody&amp;rsquo;s hand goes up, resist the urge to blame your people.&lt;/p&gt;
&lt;p&gt;Ask instead what you, and the organization above you, built that made this feel like the right call.&lt;/p&gt;
&lt;p&gt;Because I promise you: somewhere, at some point, someone tried to own that thing. And something happened that taught them not to.&lt;/p&gt;
&lt;p&gt;Your job is to find out what that was, and fix it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Good ownership isn&amp;rsquo;t found. It&amp;rsquo;s cultivated, or it&amp;rsquo;s extinguished. Pick one.&lt;/strong&gt;&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>The Lie That Taught Me About Blameless Culture</title>
      <link>https://blog.coffeeordeath.dev/posts/the-lie-that-taught-me-about-blameless-culture/</link>
      <pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/the-lie-that-taught-me-about-blameless-culture/</guid>
      <description>I was standing outside a conference room watching my team lie to another team about a database outage. It was day four.
Through the glass door, I could see the engineer on the call, explaining with impressive confidence that our cloud provider was having issues. Any minute now, they said, the vendor would resolve it and services would come back up.
I pulled up the provider&amp;rsquo;s status page. Green across the board.</description>
      <content>&lt;p&gt;I was standing outside a conference room watching my team lie to another team about a database outage. It was day four.&lt;/p&gt;
&lt;p&gt;Through the glass door, I could see the engineer on the call, explaining with impressive confidence that our cloud provider was having issues. Any minute now, they said, the vendor would resolve it and services would come back up.&lt;/p&gt;
&lt;p&gt;I pulled up the provider&amp;rsquo;s status page. Green across the board. Checked their Twitter. Nothing.
I caught the eye of the engineer leading the response and motioned them outside.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;What actually happened?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The way they looked at the floor told me everything. Then: &amp;ldquo;Accidental deletion. We&amp;rsquo;re working on recovery.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the thing about inheriting a team: You don&amp;rsquo;t just inherit their code and their systems. You inherit their patterns. Their fears. The lessons they learned from whoever led them before you.&lt;/p&gt;
&lt;p&gt;My team had learned that mistakes meant consequences. That admitting fault meant exposure. That the safest play was to redirect blame to something faceless—a cloud provider, a vendor, a system &amp;ldquo;acting weird.&amp;rdquo; They&amp;rsquo;d learned this so well that lying to the team we supported felt safer than telling their new boss the truth.&lt;/p&gt;
&lt;p&gt;I had about thirty seconds to decide what kind of leader I was going to be.&lt;/p&gt;
&lt;h2 id=&#34;the-conversation&#34;&gt;The Conversation&lt;/h2&gt;
&lt;p&gt;We found an empty office.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I&amp;rsquo;m not mad about the database,&amp;rdquo; I said. &amp;ldquo;I&amp;rsquo;m concerned you felt you couldn&amp;rsquo;t tell me.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They stared at their hands.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I need you to understand something. I&amp;rsquo;m four days into this role. I don&amp;rsquo;t know your systems. I don&amp;rsquo;t know the history here. But I know this: I can&amp;rsquo;t help you fix problems I don&amp;rsquo;t know about. And the other team can&amp;rsquo;t tell us how severe this actually is if we&amp;rsquo;re feeding them fiction.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Previous leadership didn&amp;rsquo;t react well to mistakes.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;There it was. The inherited culture, delivered in one sentence.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Okay. Here&amp;rsquo;s the new standard: You tell me the truth, always. I&amp;rsquo;ll handle the hard conversations. You focus on the fix. Deal?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They looked up. &amp;ldquo;Deal.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;falling-on-the-sword&#34;&gt;Falling on the Sword&lt;/h2&gt;
&lt;p&gt;I walked back into that conference room.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;I need to correct something,&amp;rdquo; I said. &amp;ldquo;This isn&amp;rsquo;t a cloud outage. One of our engineers accidentally deleted the pre-production database. We&amp;rsquo;re working on recovery now, but I need your help: How critical is this to your operations? What&amp;rsquo;s the actual impact?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The silence lasted maybe three seconds. Felt like an hour.&lt;/p&gt;
&lt;p&gt;Then: &amp;ldquo;Thank you for being straight with us. Here&amp;rsquo;s what we need&amp;hellip;&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They walked us through their priorities. Which services mattered most. What data they could recreate versus what needed restoration. The conversation became a collaboration instead of a performance.&lt;/p&gt;
&lt;p&gt;We got the database back. It took six hours and some creative recovery work, but we got it back.&lt;/p&gt;
&lt;h2 id=&#34;the-post-mortem&#34;&gt;The Post-Mortem&lt;/h2&gt;
&lt;p&gt;The next day, I gathered the team in the conference room for our first post-mortem. I&amp;rsquo;d never run one with them before, so I started by explaining how we were going to do them from now on.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;When we name people in post-mortems,&amp;rdquo; I said, &amp;ldquo;it&amp;rsquo;s only for successes. Who responded quickly. Who had the knowledge to guide recovery. Who communicated clearly under pressure.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I looked at the engineer who&amp;rsquo;d led the recovery effort. &amp;ldquo;You did excellent work getting us back online.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&amp;ldquo;For failures, the buck stops with me. I&amp;rsquo;m the lead. Anything that goes wrong happened on my watch. My job is to figure out what systemic failures allowed this to happen and fix them.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I watched the team process this. Some looked skeptical. Some looked relieved. All of them looked like they were waiting for the other shoe to drop.&lt;/p&gt;
&lt;p&gt;&amp;ldquo;Here&amp;rsquo;s what we&amp;rsquo;re changing. Permissions on production-like environments. Backup verification procedures. Documentation on recovery processes. And most importantly: the expectation that you can always tell me the truth, especially when things break.&amp;rdquo;&lt;/p&gt;
&lt;h2 id=&#34;the-timing&#34;&gt;The Timing&lt;/h2&gt;
&lt;p&gt;Looking back, week one was actually the perfect time for this to happen.&lt;/p&gt;
&lt;p&gt;I had no history with this team. No pattern of reactions for them to predict. No baggage about &amp;ldquo;how we&amp;rsquo;ve always done things.&amp;rdquo; When I said &amp;ldquo;this is the new standard,&amp;rdquo; there was no previous version of me to contradict it.&lt;/p&gt;
&lt;p&gt;Being new meant I had nothing to lose and everything to prove. I couldn&amp;rsquo;t lean on established trust or authority. I had to show them, in real time, what kind of leader I was going to be.&lt;/p&gt;
&lt;p&gt;Culture isn&amp;rsquo;t built during the smooth times. It&amp;rsquo;s built in how you respond when things break. And things always break.&lt;/p&gt;
&lt;h2 id=&#34;what-stuck&#34;&gt;What Stuck&lt;/h2&gt;
&lt;p&gt;That engineer never lied to me again. Neither did anyone else on the team. Not because I&amp;rsquo;d scared them straight, but because I&amp;rsquo;d shown them the alternative: honesty gets you help, not punishment.&lt;/p&gt;
&lt;p&gt;The other team? They became one of our strongest advocates. Years later, they still reference that incident as the moment they knew they could trust us.&lt;/p&gt;
&lt;p&gt;The post-mortem approach stuck too. Name people for successes, take responsibility for failures. It sounds simple, but it changes everything about how teams respond to pressure.&lt;/p&gt;
&lt;p&gt;I couldn&amp;rsquo;t change what my team had learned before I arrived. I couldn&amp;rsquo;t undo whatever had taught them that self-preservation trumped transparency. But I could show them a different future, starting with one deleted database and one honest conversation.&lt;/p&gt;
&lt;p&gt;The bad news: People will test your principles during crisis. The good news: Crisis is when principles matter most.&lt;/p&gt;
&lt;p&gt;The job is to be ready anyway.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>Everything I learned about Ops, I learned playing goal</title>
      <link>https://blog.coffeeordeath.dev/posts/everything-i-learned-about-ops-i-learned-playing-goal/</link>
      <pubDate>Thu, 06 Nov 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/everything-i-learned-about-ops-i-learned-playing-goal/</guid>
      <description>Well, maybe not everything.
Look, being a goalie is objectively ridiculous. You strap on forty pounds of equipment designed to protect you from frozen rubber traveling at speeds that would make physicists frown. Then you stand in front of a net and dare people to shoot at you. It&amp;rsquo;s a strange job.
But here&amp;rsquo;s the thing: being a goalie is basically the same job as running ops, or security, or honestly, any part of software development where you&amp;rsquo;re the one who has to keep the thing from breaking.</description>
      <content>&lt;p&gt;Well, maybe &lt;strong&gt;not everything&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Look, being a goalie is objectively ridiculous. You strap on forty pounds of equipment designed to protect you from frozen rubber traveling at speeds that would make physicists frown. Then you stand in front of a net and dare people to shoot at you. It&amp;rsquo;s a strange job.&lt;/p&gt;
&lt;p&gt;But here&amp;rsquo;s the thing: being a goalie is basically the same job as running ops, or security, or honestly, any part of software development where you&amp;rsquo;re the one who has to keep the thing from breaking.&lt;/p&gt;
&lt;p&gt;Both jobs share the same core truth that nobody tells you in the job description: when everything&amp;rsquo;s working, no one notices you&amp;rsquo;re there. But the &lt;em&gt;second&lt;/em&gt; something gets past you? Everyone knows exactly where you were and what you should have done differently.&lt;/p&gt;
&lt;p&gt;You spend most of your time reading patterns, anticipating problems before they fully materialize, and communicating with people who may or may not be listening. You&amp;rsquo;re simultaneously trying to follow the systems that keep you consistent while also staying loose enough to react when something weird happens. Because something weird &lt;em&gt;always&lt;/em&gt; happens.&lt;/p&gt;
&lt;p&gt;And whether it&amp;rsquo;s a breakaway in overtime or a production incident at 2 AM, you learn pretty quickly that freezing up isn&amp;rsquo;t an option. You make the save or you don&amp;rsquo;t. You stop the breach or you don&amp;rsquo;t. Then you&amp;rsquo;ve got about eight seconds to shake it off and get ready for the next shot, because there&amp;rsquo;s always a next shot.&lt;/p&gt;
&lt;p&gt;The bad news? You can&amp;rsquo;t stop everything. The good news? Neither can anyone else. The job is to be ready anyway.&lt;/p&gt;
</content>
    </item>
    
    <item>
      <title>New Team, New Processes</title>
      <link>https://blog.coffeeordeath.dev/posts/new-team-new-processes/</link>
      <pubDate>Mon, 31 Mar 2025 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/new-team-new-processes/</guid>
      <description>Starting fresh with a new team—whether it&amp;rsquo;s stepping into the net as a goalie or joining a company as a Site Reliability Engineer—comes with a rush of excitement, uncertainty, and the need to quickly adapt. In both worlds, you&amp;rsquo;re expected to understand the system, earn trust fast, and make the right decisions under pressure. This piece explores the parallels between guarding the crease in ice hockey and taking on infrastructure responsibilities in a new engineering org: the importance of communication, learning team dynamics, managing risk, and building confidence through early wins.</description>
      <content>&lt;p&gt;Starting fresh with a new team—whether it&amp;rsquo;s stepping into the net as a goalie or joining a company as a Site Reliability Engineer—comes with a rush of excitement, uncertainty, and the need to quickly adapt. In both worlds, you&amp;rsquo;re expected to understand the system, earn trust fast, and make the right decisions under pressure. This piece explores the parallels between guarding the crease in ice hockey and taking on infrastructure responsibilities in a new engineering org: the importance of communication, learning team dynamics, managing risk, and building confidence through early wins.&lt;/p&gt;
&lt;p&gt;When you join a new hockey team as a goalie, you don’t just protect the net—you learn how the defensemen play, how aggressive the forwards are, and when the team tends to collapse or stretch the ice. Every team has its own rhythm and unspoken rules.
The same is true when joining a new company as an SRE. The architecture is only part of the story; the real learning curve is understanding how incidents are handled, who actually makes the calls in a crunch, and which “informal” channels carry the most context. Just like a goalie can’t force their style on a team that plays differently, a new SRE needs to observe, adapt, and gradually find ways to complement and strengthen the existing flow. It&amp;rsquo;s less about proving you’re technically sound (that’s assumed) and more about showing that you can read the play and make the team better as a result.&lt;/p&gt;
&lt;p&gt;In both hockey and SRE work, the moment things go sideways is when communication really counts. As a goalie, you see the whole ice—you’re the only one facing the play—and you need to call out shifts, loose players, or breakdowns in real time. If you hesitate, the puck’s in the net.
As an SRE, it’s similar during an incident: you often have a high-level view of what’s happening across systems, and the ability to stay calm and clearly relay what you’re seeing can make the difference between a contained issue and a full-blown outage. But communication isn’t just about shouting the loudest—it’s about knowing when to speak, who needs to hear what, and doing it in a way that builds confidence rather than panic. Like a goalie steadying the team during a penalty kill, an SRE who brings calm clarity in high-stakes moments becomes an anchor for everyone else to rally around.&lt;/p&gt;
&lt;h3 id=&#34;so-how-do-you-build-trust-in-a-new-team&#34;&gt;So, how do you build trust in a new team?&lt;/h3&gt;
&lt;p&gt;Trust isn’t earned by pretending to be perfect—it’s built by showing up consistently, being accountable, and owning your impact, good or bad. As a goalie, you’re going to let in goals. Some are unstoppable, some are a result of defensive breakdowns, and some are just flat-out on you.
What matters is how you handle it. Do you sulk? Blame? Or do you skate over, tap a teammate’s shin pads, and reset for the next faceoff? The same dynamic plays out in SRE. Early on, you’ll miss signals, ship a bad change, or overlook something you &lt;em&gt;should&lt;/em&gt; have caught. Heck, there are days where you&amp;rsquo;ll break production.
But by being transparent, acknowledging the miss, and showing what you’re doing to improve, you demonstrate humility and build credibility. Teams trust people who are honest about their fallibility and show a pattern of learning. Whether it’s a soft goal from the point or a mistuned alert that paged the team at 3 a.m., owning it builds more goodwill than perfection ever could. That openness also invites others to bring their guard down, leading to better collaboration and stronger long-term bonds.&lt;/p&gt;
&lt;p&gt;As both a goalie and a leader, you quickly learn that most mistakes aren’t rooted in malice or incompetence—they’re byproducts of unclear systems, missing context, or just the fast pace of the game. When someone misses an assignment on the ice or forgets to update a runbook, it’s rarely because they don’t care—it’s usually because the process failed them. As an SRE leader, it’s your job to zoom out after the fact and ask not just &lt;em&gt;who&lt;/em&gt; made the mistake, but &lt;em&gt;why&lt;/em&gt; it made sense in that moment—and how the system can be improved so the next person has a better shot. Maybe the alert was noisy, the handoff lacked key context, or a dependency wasn’t clearly documented. Whatever it is, you shift the focus from blame to learning. Like a goalie giving quiet guidance during a timeout, your job is to reset the team without undercutting their confidence. Owning your own mistakes gives you the credibility to address others with empathy—and turning those moments into durable process improvements is how teams get stronger over time.&lt;/p&gt;
&lt;p&gt;Whether you&amp;rsquo;re stepping onto the ice or into a new engineering org, the fundamentals are the same: learn the team, communicate clearly under pressure, take ownership, and lead with empathy. Both roles are high-trust, high-impact, and deeply human. You’re not just reacting—you’re anticipating, supporting, and constantly adjusting to help the people around you succeed. The best teams don’t expect perfection—they expect presence, accountability, and a shared commitment to get better together. In the end, it&amp;rsquo;s not about never letting a goal in or preventing every outage. It&amp;rsquo;s about being the kind of teammate who learns fast, shows up when it matters, and helps turn every setback into forward motion.&lt;/p&gt;
</content>
    </item>
    
  </channel>
</rss>
