<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Cloud on Luke Little</title><link>https://lukelittle.com/categories/cloud/</link><description>Recent content in Cloud on Luke Little</description><image><title>Luke Little</title><url>https://lukelittle.com/images/og-default.png</url><link>https://lukelittle.com/images/og-default.png</link></image><generator>Hugo -- 0.152.2</generator><language>en</language><lastBuildDate>Mon, 20 Oct 2025 09:00:00 -0500</lastBuildDate><atom:link href="https://lukelittle.com/categories/cloud/index.xml" rel="self" type="application/rss+xml"/><item><title>Learning from the October 20 AWS Outage: Questions Every Team Should Ask</title><link>https://lukelittle.com/posts/2025/10/learning-from-the-october-20-aws-outage-questions-every-team-should-ask/</link><pubDate>Mon, 20 Oct 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/10/learning-from-the-october-20-aws-outage-questions-every-team-should-ask/</guid><description>What happened in the October 20 AWS outage, how to explain the impact to leadership, and the resilience questions every team should ask.</description><content:encoded><![CDATA[<p>On the morning of October 20, AWS us-east-1 services were degraded—in particular, DNS services for DynamoDB. Most of us didn&rsquo;t find out from monitoring alerts or dashboards. We found out because the apps on our phones stopped working.</p>
<p>That&rsquo;s the reality of modern infrastructure incidents: they often surface as user-facing failures long before the official root cause analysis lands in your inbox.</p>
<p>I wrote about this outage for <a href="https://blog.ippon.tech/explaining-the-october-20-aws-outage-to-leadership/">Ippon Technologies</a>, focusing on three critical aspects that go beyond just understanding what broke:</p>
<h2 id="what-actually-happened">What Actually Happened</h2>
<p>This wasn&rsquo;t a full region going dark—it was a DNS problem at a foundational layer. DNS (Domain Name System) is the internet&rsquo;s address book. When DNS breaks, everything that depends on it breaks too.</p>
<p>It&rsquo;s like someone removing all the street signs in a city overnight—your services are still there, but nothing can find its way.</p>
<p>For many teams, it meant increased error rates, intermittent failures, and retry storms. Not catastrophic downtime, but the kind of disruption that floods support tickets and frustrates users.</p>
<h2 id="having-the-leadership-conversation">Having the Leadership Conversation</h2>
<p>This is where many engineers struggle. You know the issue was upstream. You know it&rsquo;s AWS&rsquo;s infrastructure. But leadership doesn&rsquo;t care about the cloud provider—they care about impact and what you&rsquo;re doing about it.</p>
<p>The key is framing it as a <strong>dependency visibility problem</strong>, not a blame game:</p>
<blockquote>
<p>&ldquo;We were affected by a regional outage in AWS&rsquo;s us-east-1 region due to DNS issues. This exposed areas where we&rsquo;re overly dependent on single-region infrastructure. We&rsquo;re using this to map our regional dependencies, prioritize applications by criticality, and identify where we need fallback logic and multi-region routing.&rdquo;</p>
</blockquote>
<p>That&rsquo;s ownership. That&rsquo;s a path forward.</p>
<h2 id="the-hard-questions-you-need-to-answer">The Hard Questions You Need to Answer</h2>
<p>The article explores four critical questions every team should be able to answer confidently:</p>
<ol>
<li><strong>Which of your apps run in us-east-1?</strong></li>
<li><strong>Which rely on DynamoDB?</strong></li>
<li><strong>Which of those are Tier 1 or customer-facing?</strong></li>
<li><strong>Which of those have active-active failover across regions?</strong></li>
</ol>
<p>Most teams can&rsquo;t answer these questions. Not because they&rsquo;re negligent—but because cloud estates grow organically. Services get deployed. Teams change. Documentation drifts. Before you know it, you&rsquo;re running critical workloads on infrastructure patterns that no one fully understands anymore.</p>
<h2 id="making-resilience-visible">Making Resilience Visible</h2>
<p>The full article dives into:</p>
<ul>
<li><strong>Structured risk assessment</strong> using tools like AWS Resilience Hub to define applications and assess risk against RTO/RPO targets</li>
<li><strong>Chaos engineering</strong> with AWS Fault Injection Service (FIS) to validate that resilience isn&rsquo;t just theoretical</li>
<li><strong>Cultural shifts</strong> to prioritize resilience alongside feature delivery</li>
<li><strong>Practical next steps</strong> for mapping dependencies and building observability around failure modes</li>
</ul>
<h2 id="why-this-matters">Why This Matters</h2>
<p>Today&rsquo;s outage wasn&rsquo;t catastrophic—but it was loud enough to get everyone&rsquo;s attention. It revealed real architectural risks that often go unnoticed until they become outages.</p>
<p>Outages like this are reminders, not just disruptions. They&rsquo;re opportunities to begin conversations across architecture, risk, and engineering teams about what resilience really means for your organization.</p>
<p>Not in terms of making everything indestructible, but in making risk visible, decisions intentional, and recovery predictable.</p>
<p>👉 <strong><a href="https://blog.ippon.tech/explaining-the-october-20-aws-outage-to-leadership/">Read the full article on the Ippon blog</a></strong> for detailed guidance on communicating with leadership, assessing your infrastructure, and building measurable resilience practices.</p>
<p>The goal isn&rsquo;t perfection—it&rsquo;s visibility, intention, and readiness for when the next incident inevitably arrives.</p>
]]></content:encoded></item></channel></rss>