<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Reliability on Coffee or Blog</title>
    <link>https://blog.coffeeordeath.dev/tags/reliability/</link>
    <description>Recent content in Reliability on Coffee or Blog</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Sun, 04 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.coffeeordeath.dev/tags/reliability/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Why Good SREs Feel Like Wall Builders</title>
      <link>https://blog.coffeeordeath.dev/posts/why-good-sres-feel-like-wall-builders/</link>
      <pubDate>Sun, 04 Oct 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.coffeeordeath.dev/posts/why-good-sres-feel-like-wall-builders/</guid>
      <description>Someone pinged me at 4:40 on a Friday: &amp;ldquo;need a new Postgres instance, going live Monday, should be quick.&amp;rdquo; I asked what the read/write ratio looked like and what happens to the checkout flow if the instance falls over for ten minutes. But, they hadn&amp;rsquo;t thought about either one, because nobody had ever asked them to.
Working with a bad SRE team is frustrating for all the obvious reasons: things break, deploys fail, and nobody answers the page.</description>
      <content>&lt;p&gt;Someone pinged me at 4:40 on a Friday: &amp;ldquo;need a new Postgres instance, going live Monday, should be quick.&amp;rdquo; I asked what the read/write ratio looked like and what happens to the checkout flow if the instance falls over for ten minutes. But, they hadn&amp;rsquo;t thought about either one, because nobody had ever asked them to.&lt;/p&gt;
&lt;p&gt;Working with a bad SRE team is frustrating for all the obvious reasons: things break, deploys fail, and nobody answers the page.&lt;/p&gt;
&lt;p&gt;Working with a good one is frustrating for a totally different reason, we ask the questions you weren&amp;rsquo;t ready to answer.&lt;/p&gt;
&lt;p&gt;How many reads per write? What happens when the cache is empty? Can service A deploy while service B is down? None of those questions are a no, they&amp;rsquo;re how we figure out how deep the foundation has to go before you put a building on top of it. But from the other side of the request, it looks a lot like someone building a wall.&lt;/p&gt;
&lt;p&gt;And that&amp;rsquo;s been true for about as long as there have been ops teams. What is rapidly changing is how fast the building goes up.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;One engineer with an agent can now produce in a weekend what used to take a small team a quarter. Ask for &amp;ldquo;a scalable user service&amp;rdquo; and you can have two services, a Postgres instance, a Redis cache, and the Terraform to stand it all up before lunch. The code usually works, but, the problem lies in everything that didn&amp;rsquo;t happen on the way to code completion.&lt;/p&gt;
&lt;p&gt;On a normal team, code passed through people. A senior engineer reviewing the PR asks why there are two services. A tech lead in a design review asks what growth looks like. Someone who has been paged for this kind of thing before asks what happens when it falls over. Those conversations were slow and, honestly, occasionally annoying, but they were where most of the operational questions actually got asked. Even then, few organizations I&amp;rsquo;ve worked in wanted to talk about users per hour, how much data they&amp;rsquo;d be moving, or latency targets before there were even mockups of the new UI.&lt;/p&gt;
&lt;p&gt;When one person and an agent build the whole thing, those conversations don&amp;rsquo;t happen. The agent was asked for something that works, and it delivered exactly that, but nobody asked it for something that survives.&lt;/p&gt;
&lt;p&gt;So the first time anyone asks those questions is when the request lands on the platform team, and the reaction is predictable:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;I built this in two days. Why is the infra review taking three?&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Look, I get it. From where they&amp;rsquo;re sitting, we&amp;rsquo;re the slowest step in the pipeline by a kilometre (mile). We&amp;rsquo;re also the first step where anyone has bothered to ask how the thing behaves under load.&lt;/p&gt;
&lt;p&gt;I wrote about the release side of this in &lt;a href=&#34;https://blog.coffeeordeath.dev/posts/release-discipline-in-the-age-of-ai-accelerated-development/&#34;&gt;Release Discipline in the Age of AI-Accelerated Development&lt;/a&gt;. This is the same gap, just one layer down.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;I see the same three patterns land on the platform team, over and over, and I&amp;rsquo;d bet most of you have seen at least one of them this year.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The default database.&lt;/strong&gt; The agent stands up Postgres with default settings and generated migrations. It works fine in dev. Nobody has asked about the read/write ratio, how big the largest table will be in a year, or how long the data needs to be kept. The one that usually bites first is connections. Every pod opens its own pool, autoscaling adds pods during a traffic spike, and Postgres runs out of connections at exactly the moment you need it most.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The cache that became a database.&lt;/strong&gt; Redis goes in front of a slow route to &amp;ldquo;make it faster.&amp;rdquo; There&amp;rsquo;s no eviction policy, no TTL, and no plan for what happens when Redis restarts or gets flushed. A cache is fine as long as the service can live without it. Once the service falls over on a cold cache, it isn&amp;rsquo;t a cache anymore, it&amp;rsquo;s a primary data store that nobody chose, nobody sized, and nobody is backing up (and don&amp;rsquo;t get me started on backups).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The distributed monolith.&lt;/strong&gt; The prompt said &amp;ldquo;microservices,&amp;rdquo; so you got two of them, in separate repos with separate pipelines (no monorepo huh?). But A can&amp;rsquo;t do anything without B, they have to deploy together, and there are no timeouts, retry budgets, or fallbacks between them (did we even connect the network segments?!). You&amp;rsquo;re paying for microservices (network hops, distributed tracing, two pipelines) and getting none of the benefits (independent deploys, failure isolation). It&amp;rsquo;s just a monolith that talks to itself over the network.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;It&amp;rsquo;s tempting to blame the developer(s) here, but, that&amp;rsquo;s not right, and honestly it doesn&amp;rsquo;t help anyway.&lt;/p&gt;
&lt;p&gt;An agent does what the prompt says(this part may age poorly). Ask for &amp;ldquo;a service that reads from Postgres and caches in Redis&amp;rdquo; and that&amp;rsquo;s what you&amp;rsquo;ll get, built competently. It probably won&amp;rsquo;t stop to ask what happens to database connections when autoscaling kicks in, how A behaves if B takes fifteen seconds to respond, or whether this needed to be two services at all. Most agents will happily work through those questions if you ask them to, very few will raise them on their own.&lt;/p&gt;
&lt;p&gt;I wouldn&amp;rsquo;t call that a flaw in the model, it&amp;rsquo;s more of a gap in experience. Operational instincts come from getting paged for something you built, and the model has never been paged. Plenty of the people prompting it haven&amp;rsquo;t been either, which is fine, that&amp;rsquo;s what the rest of the team used to be for.&lt;/p&gt;
&lt;p&gt;So the platform team ends up being the first place in the pipeline that carries that context. To someone moving that fast it feels like a wall, but from our side we&amp;rsquo;re just checking the ground before the concrete goes in.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;I&amp;rsquo;m not interested in slowing anyone down (most Platform Engineers aren&amp;rsquo;t either), the speed is real and it&amp;rsquo;s useful. What I want is for the operational questions to get asked earlier, by the developer and the agent, so they aren&amp;rsquo;t all landing on us at the end.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you&amp;rsquo;re building with agents, prompt for day two.&lt;/strong&gt; Put operational constraints in the prompt, not just functionality.&lt;/p&gt;
&lt;p&gt;Instead of:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Build a user service with Redis caching.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Try:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Build a user service with Redis caching. If Redis is unavailable, fall back to Postgres and keep serving. Set explicit TTLs and a memory cap with an eviction policy. Add timeouts and a circuit breaker on downstream calls.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It&amp;rsquo;s three more sentences, and it gets you something the review will mostly wave through.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you&amp;rsquo;re on a platform team, put your rules where the agent can read them.&lt;/strong&gt; A twenty-page review process is a good way to teach people to route around you. Write your operational defaults (connection limits, required metrics and alerts, cache TTLs, timeout budgets) into an &lt;code&gt;AGENTS.md&lt;/code&gt; in your repo templates. The agent picks it up, and the first draft lands closer to what you&amp;rsquo;d have asked for anyway. Then give people Terraform modules that are already sized sensibly and already have monitoring, scaling bounds, and backups wired in. The paved path has to be the easiest path for the agent too, otherwise nobody&amp;rsquo;s going to use it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;the-napkin-spec&#34;&gt;The napkin spec&lt;/h2&gt;
&lt;p&gt;Before the build starts, or at the latest before the infra request goes in, answer five questions. It fits on a napkin, and it covers most of what a review would ask anyway.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Failure:&lt;/strong&gt; If this service or its database goes down completely, what does the user see?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scale:&lt;/strong&gt; What do memory, storage, connections, and throughput look like at 3x today&amp;rsquo;s load?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coupling:&lt;/strong&gt; Can this deploy while its dependencies are deploying or down?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State:&lt;/strong&gt; Where does the data live, how long is it kept, and has anyone actually restored it from a backup?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ownership:&lt;/strong&gt; What tells you it&amp;rsquo;s broken, and who gets paged? (If the answer is &amp;ldquo;nobody,&amp;rdquo; &lt;a href=&#34;https://blog.coffeeordeath.dev/posts/why-doesnt-anyone-own-that/&#34;&gt;that&amp;rsquo;s its own problem&lt;/a&gt;.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&amp;ldquo;I don&amp;rsquo;t know&amp;rdquo; is a perfectly fine answer to any of these, and it&amp;rsquo;s a lot cheaper to find that out here than in an incident channel.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The service in the intro? It went live Monday, more or less on schedule. But, the base image was built in the wrong account, the prod environment was missing the database password, and nobody had put the new service&amp;rsquo;s URL into config management, so checkout was down for the first four hours while three of us traced it live on a call that was supposed to be someone&amp;rsquo;s lunch. Once the &amp;ldquo;oops, we forgot about those configuration items&amp;rdquo; were fixed, the application scaled out so far and so fast that it saturated our connection pool and completely locked up the database and took several more hours to stabilize, but, we still don&amp;rsquo;t have a clean answer on the read/write ratio. Sadly, we got our answer on what happens to checkout when the instance isn&amp;rsquo;t there: it breaks, loudly, and the week after launch wasn&amp;rsquo;t spent on anything in the quarterly plan. It was spent rightsizing the thing that should have been sized on Friday.&lt;/p&gt;
</content>
    </item>
    
  </channel>
</rss>
