<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Engineering on Luke Little</title><link>https://lukelittle.com/categories/engineering/</link><description>Recent content in Engineering on Luke Little</description><image><title>Luke Little</title><url>https://lukelittle.com/images/og-default.png</url><link>https://lukelittle.com/images/og-default.png</link></image><generator>Hugo -- 0.152.2</generator><language>en</language><lastBuildDate>Mon, 13 Apr 2026 14:00:00 -0400</lastBuildDate><atom:link href="https://lukelittle.com/categories/engineering/index.xml" rel="self" type="application/rss+xml"/><item><title>Scaling Resilience with AWS Resilience Hub: A Multi-Account Reality Check</title><link>https://lukelittle.com/posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/</link><pubDate>Mon, 13 Apr 2026 14:00:00 -0400</pubDate><guid>https://lukelittle.com/posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/</guid><description>Most resilience programs fail because they never move beyond isolated assessments. Here&amp;#39;s how to scale AWS Resilience Hub across accounts and build a real program.</description><content:encoded><![CDATA[<p>Everyone in financial services talks about resilience. We have DR plans, architecture diagrams, dashboards, and increasingly, tools like AWS Resilience Hub. On paper, it all looks good. In practice, most resilience programs don&rsquo;t fail because of missing tooling — they fail because they never move beyond isolated assessments.</p>
<p>A team runs a Resilience Hub assessment. They get a score. Maybe they even fix a few findings. And then nothing happens. No aggregation. No cadence. No program. Just a snapshot that lives in one account, attached to one application, reviewed once.</p>
<p>I&rsquo;ve seen this pattern enough times in financial services to know it&rsquo;s not a tooling problem. It&rsquo;s an organizational one. And if you want resilience that actually means something, you have to confront it directly.</p>
<h2 id="resilience-doesnt-live-in-one-account">Resilience Doesn&rsquo;t Live in One Account</h2>
<p>Modern financial systems don&rsquo;t live in one place. They span line-of-business accounts, shared platform layers, data environments, and third-party dependencies. The critical user journeys — payments, trades, claims — cut across all of them.</p>
<p>Yet Resilience Hub is typically deployed against one app, in one account, with one assessment. That&rsquo;s not resilience. That&rsquo;s local optimization pretending to be a program.</p>
<p>As outlined in the <a href="https://docs.aws.amazon.com/whitepapers/latest/building-resilient-financial-services/building-resilient-financial-services.html">Building Resilient Financial Services</a> whitepaper, resilience must be measurable and demonstrable across interconnected services, not evaluated in isolation. A payment flow that touches API Gateway in one account, a processing Lambda in a shared services account, and DynamoDB in a data account doesn&rsquo;t care that each of those components passed its individual resilience check. What matters is whether the flow as a whole can survive failure.</p>
<p>This is the gap Resilience Hub doesn&rsquo;t close on its own. It was designed to evaluate applications, not programs. If you want the latter, you have to build it.</p>
<h2 id="the-pattern-that-actually-works-hub-and-spoke">The Pattern That Actually Works: Hub-and-Spoke</h2>
<p>If you want to scale resilience, you need to treat it like a distributed measurement system. Each account owns its applications and runs its own assessments. But the program — the aggregation, the scoring, the trend analysis — lives centrally.</p>
<p><img src="/posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/Scaling%20Resilience%20with%20AWS%20Resilience%20Hub_hu_3f32f4eb9a84f4cb.webp" srcset="/posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/Scaling%20Resilience%20with%20AWS%20Resilience%20Hub_hu_9ef337f03144deac.webp 750w, /posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/Scaling%20Resilience%20with%20AWS%20Resilience%20Hub_hu_3f32f4eb9a84f4cb.webp 1231w" sizes="(max-width: 800px) 100vw, 750px"
       width="1231" height="1160" alt="Hub-and-Spoke Multi-Account Resilience Architecture" loading="lazy" decoding="async"></p>
<p>This isn&rsquo;t complicated architecture. It&rsquo;s a standard IAM role in each spoke account that a central account can assume, a scheduled Lambda or Step Functions workflow that pulls assessment data across accounts and normalizes the structure, a data store in S3 or DynamoDB that tracks history (not just current state), and a dashboard layer — QuickSight or whatever your organization already uses — tied to critical services and business impact.</p>
<p>The hard part isn&rsquo;t building it. The hard part is getting the organizational agreement that resilience data should flow centrally in the first place.</p>
<h2 id="resilience-hub-is-a-data-source-not-a-dashboard">Resilience Hub Is a Data Source, Not a Dashboard</h2>
<p>This is where I think most teams go wrong. They use Resilience Hub as a dashboard — open the console, look at the score, maybe screenshot it for a quarterly review. That&rsquo;s underutilizing it significantly.</p>
<p>Resilience Hub is a modeling engine, a policy evaluator, and a scoring system. It is not your enterprise view, your governance layer, or your program. The real value emerges when you stop looking at the UI and start treating it as a data source.</p>
<p>The API surface is small and that&rsquo;s a good thing:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span>client<span style="color:#f92672">.</span>list_apps()
</span></span><span style="display:flex;"><span>client<span style="color:#f92672">.</span>list_app_assessments()
</span></span><span style="display:flex;"><span>client<span style="color:#f92672">.</span>describe_app_assessment()
</span></span></code></pre></div><p>From those three calls, you can build enterprise-wide resilience scoring, trend analysis over time, drift detection across environments, and audit evidence tied to real systems. The UI tells you what happened once. The API lets you prove improvement over time — and that&rsquo;s what regulators actually care about.</p>
<h2 id="where-this-connects-to-program-cadence">Where This Connects to Program Cadence</h2>
<p>This is the part most organizations completely miss. Resilience is not a tool, a dashboard, or a quarterly checkbox. It&rsquo;s a cadence.</p>
<p>I keep coming back to this when working with financial services clients: the difference between organizations that have a resilience program and organizations that have a resilience report is whether they review on a regular cadence and actually act on what they find.</p>
<table>
  <thead>
      <tr>
          <th>Cadence</th>
          <th>What actually happens</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Weekly</td>
          <td>Pull Resilience Hub data, detect drift</td>
      </tr>
      <tr>
          <td>Monthly</td>
          <td>Review resilience posture across services</td>
      </tr>
      <tr>
          <td>Quarterly</td>
          <td>Validate against real scenarios</td>
      </tr>
      <tr>
          <td>Continuous</td>
          <td>Track improvement and regressions</td>
      </tr>
  </tbody>
</table>
<p>Monthly reviews. Quarterly scenario testing. Continuous improvement loops. If you&rsquo;re not doing these, you don&rsquo;t have a program — you have a document that says you do.</p>
<h2 id="from-assessment-to-evidence">From Assessment to Evidence</h2>
<p>Regulators don&rsquo;t care about your architecture diagrams. They care about whether services stay within tolerance, whether failures are tested, and whether you improve over time. These aren&rsquo;t unreasonable asks, but they require evidence that most organizations can&rsquo;t produce because they never built the infrastructure to collect it.</p>
<p>A multi-account Resilience Hub strategy gives you centralized visibility across distributed systems, consistent measurement aligned to business services, and — critically — evidence you can actually defend when someone asks how you know your systems are resilient.</p>
<p>Most organizations never make this leap. They run assessments, fix findings, and move on. But resilience doesn&rsquo;t scale that way. It scales when you treat it as a system — measured continuously, aggregated centrally, and reviewed on a cadence.</p>
<p>AWS Resilience Hub gives you the raw signal. What you build around it determines whether you have a tool or a capability.</p>
]]></content:encoded></item><item><title>The Missing Layer in Your Enterprise AI Stack: AWS Bedrock Guardrails</title><link>https://lukelittle.com/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/</link><pubDate>Sat, 28 Feb 2026 12:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/</guid><description>Bedrock Guardrails wrap every model call with PII redaction, denied topics, grounding checks, and an audit trail. Here&amp;#39;s how banks can use them.</description><content:encoded><![CDATA[<p>Everyone wants to ship AI into production. Almost no one wants to own what happens when it goes wrong.
I&rsquo;ve been in enough rooms with financial services clients to know how this plays out. A team builds something impressive on Bedrock — a RAG-powered knowledge assistant, an internal compliance copilot, a customer-facing chatbot. The demo looks great. Then someone in Legal raises their hand. What happens if it leaks a customer&rsquo;s SSN? What if it makes a recommendation that sounds like investment advice? What if a clever user tricks it into ignoring your system prompt?</p>
<p>The AI project stalls. Not because the technology isn&rsquo;t ready — because the governance layer isn&rsquo;t there.</p>
<p>AWS Bedrock Guardrails is that governance layer. And in regulated environments like banking, insurance, and healthcare, it&rsquo;s not optional — it&rsquo;s the prerequisite for going to production.
This post walks through what Guardrails actually does, how it works under the hood, why it matters specifically in financial services, and how to implement it with code you can actually deploy.</p>
<h2 id="what-guardrails-solves">What Guardrails Solves</h2>
<p>Let&rsquo;s be direct about the problem. Large language models have four failure modes that matter most in regulated industries:</p>
<p>Harmful content generation: Even well-prompted models can produce hate speech, violent content, or guidance on misconduct if pushed in the right direction — especially in customer-facing contexts where you can&rsquo;t predict every input.</p>
<p>Prompt injection and jailbreaks: Sophisticated users will attempt to override your system prompt, bypass your application logic, or extract information from your context window that they shouldn&rsquo;t have. This isn&rsquo;t theoretical — it&rsquo;s the first thing a red team tests.</p>
<p>PII leakage: In a RAG system where your model has access to customer records, there&rsquo;s a real risk of the model surfacing one customer&rsquo;s information in another customer&rsquo;s session, or including SSNs and account numbers in a response that gets logged, cached, or screenshotted.</p>
<p>Hallucination: For a general-purpose chatbot, hallucination is annoying. For a compliance assistant answering questions about regulatory requirements, it&rsquo;s a liability.</p>
<p>Bedrock Guardrails addresses all four — as a managed layer that sits between your application and the foundation model, evaluating both input and output independently.</p>
<h2 id="how-it-works">How It Works</h2>
<p>The core mental model is simple: guardrails wrap the model invocation, not the model itself. You define a set of policies once, attach them to your Bedrock calls, and every prompt and every response gets evaluated against those policies before anything reaches the end user.</p>
<p><img src="/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/bedrock-guardrails-evaluation-flow_hu_cc4debf7697de5c1.webp" srcset="/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/bedrock-guardrails-evaluation-flow_hu_4eb2d1f1ede614c9.webp 750w, /posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/bedrock-guardrails-evaluation-flow_hu_cc4debf7697de5c1.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="470" alt="Bedrock Guardrails Evaluation Flow" loading="lazy" decoding="async"></p>
<p><img src="/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/aws-bedrock-guardrails-architecture_hu_4936088e56b4de33.webp" srcset="/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/aws-bedrock-guardrails-architecture_hu_30aaaaa93372f4c1.webp 750w, /posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/aws-bedrock-guardrails-architecture_hu_4936088e56b4de33.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="882" alt="AWS Bedrock Guardrails Architecture" loading="lazy" decoding="async"></p>
<p>There are two evaluation passes:</p>
<ul>
<li>Input evaluation runs before the prompt reaches the foundation model. If the user&rsquo;s message violates a policy, the model never sees it — you get a blocked message back immediately.</li>
<li>Output evaluation runs after the model generates a response. The model might have produced something that passes input filters but fails on output — hallucinated content that contradicts your source documents, or a response that inadvertently includes PII from the retrieved context.</li>
</ul>
<p>If either pass blocks, you get back a configurable message. The model response is never surfaced to the user.</p>
<h2 id="the-six-policy-types">The Six Policy Types</h2>
<h3 id="1-content-filters">1. Content Filters</h3>
<p>Detect and filter harmful content across six categories: Hate, Insults, Sexual, Violence, Misconduct, and Prompt Attack. Each category has an adjustable filter strength — Low, Medium, or High — so you can calibrate based on your use case. A customer service chatbot for a brokerage doesn&rsquo;t need the same thresholds as an internal developer tool.</p>
<p>AWS extended content filtering to code-related content in 2025, which matters for any application where users can submit or request code. Harmful content in comments, variable names, and string literals is now caught at the same level as prose.</p>
<h3 id="2-prompt-attack-detection">2. Prompt Attack Detection</h3>
<p>This sits inside content filters but deserves its own callout. Jailbreaks and prompt injections are the most common adversarial inputs your application will face once it&rsquo;s live. Guardrails detects both and gives you the option to block or log them — useful for incident response when your security team wants to know who tried what.</p>
<h3 id="3-denied-topics">3. Denied Topics</h3>
<p>Define topics that are off-limits in the context of your application. For a retail banking chatbot, this might be investment advice (FINRA), cryptocurrency recommendations, or competitor product comparisons. You describe the topic in plain language; AWS uses that description to classify user inputs and model responses.</p>
<p>This is one of the more powerful policy types for financial services, because it lets you draw a hard line around regulatory risk without having to enumerate every possible phrasing of a question.</p>
<h3 id="4-sensitive-information-filters-pii-redaction">4. Sensitive Information Filters (PII Redaction)</h3>
<p>Bedrock Guardrails uses probabilistic ML detection to identify PII in both inputs and outputs. Predefined entity types include: SSN, Date of Birth, phone numbers, email addresses, credit card numbers, driver&rsquo;s license numbers, bank account numbers, and more.</p>
<p>For anything not on the predefined list — like account routing numbers in a proprietary format, or internal employee IDs — you can add custom regex patterns.</p>
<p>When PII is detected, you have two options: block the entire message, or mask the sensitive fields and allow the rest through. Masking is useful for logging and audit scenarios where you want to retain the conversation structure without storing raw PII.</p>
<h3 id="5-contextual-grounding-checks">5. Contextual Grounding Checks</h3>
<p>This is the hallucination filter, and it&rsquo;s the most technically interesting policy type for RAG applications.</p>
<p>Contextual grounding checks compare the model&rsquo;s response against two things: the source documents retrieved from your knowledge base, and the user&rsquo;s original query. It generates two scores:</p>
<ul>
<li>Grounding score: How factually consistent is the response with the source material?</li>
<li>Relevance score: Does the response actually answer what was asked?</li>
</ul>
<p>You set a threshold between 0 and 0.99 for each. A response below either threshold gets blocked. AWS recommends starting around 0.7 for both and adjusting based on testing.
In practice, this means if your compliance knowledge base says &ldquo;employees must complete annual AML training,&rdquo; and the model responds &ldquo;employees should complete AML training within 90 days of hire&rdquo; — that&rsquo;s a grounding failure. The content is plausible; it&rsquo;s just not what your source says. Contextual grounding catches it.</p>
<h3 id="6-automated-reasoning-checks">6. Automated Reasoning Checks</h3>
<p>This is the newest and most powerful capability for factual accuracy. Where contextual grounding uses ML scoring, Automated Reasoning uses formal logic — encoding your organization&rsquo;s policies as structured logical rules, then verifying model responses against those rules mathematically.</p>
<p>The practical implication: Automated Reasoning doesn&rsquo;t just score a response; it can explain why a response is incorrect and what correction would make it valid. For HR policy bots, compliance Q&amp;A systems, and any use case where you need to be able to show your work to an auditor, this is the capability that changes the conversation.</p>
<h2 id="implementation">Implementation</h2>
<p>Let&rsquo;s make this concrete. Here&rsquo;s a Terraform module that creates a guardrail configured for a financial services knowledge assistant:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-terraform" data-lang="terraform"><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_bedrock_guardrail&#34;</span> <span style="color:#e6db74">&#34;finserv_assistant&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span>                      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;finserv-knowledge-assistant&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">description</span>               <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;Guardrails for retail banking knowledge assistant&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">blocked_input_messaging</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;I&#39;m not able to help with that request. Please contact your relationship manager for assistance.&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">blocked_outputs_messaging</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;I wasn&#39;t able to generate a response that meets our quality standards. Please rephrase your question.&#34;</span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # PII: block sensitive data in inputs, mask in outputs
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">sensitive_information_policy_config</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">pii_entities_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;SSN&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">action</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;BLOCK&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">pii_entities_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;US_BANK_ACCOUNT_NUMBER&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">action</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;ANONYMIZE&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">pii_entities_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;CREDIT_DEBIT_CARD_NUMBER&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">action</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;ANONYMIZE&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">pii_entities_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;US_PASSPORT_NUMBER&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">action</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;BLOCK&#34;</span>
</span></span><span style="display:flex;"><span>    }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">    # Custom regex for internal account IDs
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>    <span style="color:#a6e22e">regexes_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">name</span>        <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;internal-account-id&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">description</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;Internal account reference numbers&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">pattern</span>     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;ACC-[0-9]{8}&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">action</span>      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;ANONYMIZE&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # Block investment advice topics — FINRA risk mitigation
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">topic_policy_config</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">topics_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">name</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;investment-advice&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">definition</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;Specific recommendations to buy, sell, or hold financial securities, stocks, bonds, mutual funds, ETFs, or other investment products.&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;DENY&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">examples</span>   <span style="color:#f92672">=</span> [
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;Should I buy more Apple stock?&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;What funds should I put my 401k into?&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;Is now a good time to sell my bonds?&#34;</span>
</span></span><span style="display:flex;"><span>      ]
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">topics_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">name</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;competitor-products&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">definition</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;Comparisons or recommendations involving competing financial institutions or their products.&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;DENY&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # Content filters — calibrated for customer-facing context
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">content_policy_config</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>            <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HATE&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">input_strength</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HIGH&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">output_strength</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HIGH&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>            <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;INSULTS&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">input_strength</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;MEDIUM&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">output_strength</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;MEDIUM&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>            <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;PROMPT_ATTACK&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">input_strength</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HIGH&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">output_strength</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;NONE&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>            <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;VIOLENCE&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">input_strength</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HIGH&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">output_strength</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HIGH&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # Contextual grounding — prevent hallucination in RAG responses
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">contextual_grounding_policy_config</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;GROUNDING&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">threshold</span> <span style="color:#f92672">=</span> <span style="color:#ae81ff">0</span>.<span style="color:#ae81ff">75</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">filters_config</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">type</span>      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;RELEVANCE&#34;</span>
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">threshold</span> <span style="color:#f92672">=</span> <span style="color:#ae81ff">0</span>.<span style="color:#ae81ff">70</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # Encrypt guardrail configuration with customer-managed key
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">kms_key_arn</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_kms_key</span>.<span style="color:#a6e22e">bedrock_guardrail_key</span>.<span style="color:#a6e22e">arn</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">tags</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">Environment</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;production&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">Application</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;finserv-knowledge-assistant&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">Compliance</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;FFIEC&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>Now attach it to your Bedrock invocation:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>bedrock <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#34;bedrock-runtime&#34;</span>, region_name<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;us-east-1&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">invoke_with_guardrails</span>(prompt: str, source_documents: list[str]) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Invoke a Bedrock model with Guardrails applied to both input and output.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    source_documents: list of text chunks retrieved from your knowledge base
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    guardrail_id <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;your-guardrail-id&#34;</span>
</span></span><span style="display:flex;"><span>    guardrail_version <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;DRAFT&#34;</span>  <span style="color:#75715e"># Use a pinned version in production</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Format the request with grounding source for contextual checks</span>
</span></span><span style="display:flex;"><span>    grounding_source <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;</span><span style="color:#ae81ff">\n\n</span><span style="color:#e6db74">&#34;</span><span style="color:#f92672">.</span>join(source_documents)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    response <span style="color:#f92672">=</span> bedrock<span style="color:#f92672">.</span>invoke_model(
</span></span><span style="display:flex;"><span>        modelId<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;anthropic.claude-3-5-sonnet-20241022-v2:0&#34;</span>,
</span></span><span style="display:flex;"><span>        guardrailIdentifier<span style="color:#f92672">=</span>guardrail_id,
</span></span><span style="display:flex;"><span>        guardrailVersion<span style="color:#f92672">=</span>guardrail_version,
</span></span><span style="display:flex;"><span>        body<span style="color:#f92672">=</span>json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;anthropic_version&#34;</span>: <span style="color:#e6db74">&#34;bedrock-2023-05-31&#34;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;max_tokens&#34;</span>: <span style="color:#ae81ff">1024</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;system&#34;</span>: <span style="color:#e6db74">f</span><span style="color:#e6db74">&#39;&#39;&#39;You are a helpful assistant for retail banking customers.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">Answer questions based only on the provided documentation.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">If the answer is not in the documentation, say so clearly.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">Documentation:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74"></span><span style="color:#e6db74">{</span>grounding_source<span style="color:#e6db74">}</span><span style="color:#e6db74">&#39;&#39;&#39;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;messages&#34;</span>: [
</span></span><span style="display:flex;"><span>                {
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#34;role&#34;</span>: <span style="color:#e6db74">&#34;user&#34;</span>,
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#34;content&#34;</span>: [
</span></span><span style="display:flex;"><span>                        {
</span></span><span style="display:flex;"><span>                            <span style="color:#e6db74">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;text&#34;</span>,
</span></span><span style="display:flex;"><span>                            <span style="color:#e6db74">&#34;text&#34;</span>: prompt,
</span></span><span style="display:flex;"><span>                            <span style="color:#e6db74">&#34;guardContent&#34;</span>: {
</span></span><span style="display:flex;"><span>                                <span style="color:#e6db74">&#34;text&#34;</span>: {
</span></span><span style="display:flex;"><span>                                    <span style="color:#e6db74">&#34;qualifiers&#34;</span>: [<span style="color:#e6db74">&#34;query&#34;</span>]
</span></span><span style="display:flex;"><span>                                }
</span></span><span style="display:flex;"><span>                            }
</span></span><span style="display:flex;"><span>                        }
</span></span><span style="display:flex;"><span>                    ]
</span></span><span style="display:flex;"><span>                }
</span></span><span style="display:flex;"><span>            ]
</span></span><span style="display:flex;"><span>        }),
</span></span><span style="display:flex;"><span>        contentType<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;application/json&#34;</span>,
</span></span><span style="display:flex;"><span>        accept<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;application/json&#34;</span>
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    result <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(response[<span style="color:#e6db74">&#34;body&#34;</span>]<span style="color:#f92672">.</span>read())
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Check if guardrails intervened</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;amazon-bedrock-guardrailAction&#34;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;INTERVENED&#34;</span>:
</span></span><span style="display:flex;"><span>        guardrail_trace <span style="color:#f92672">=</span> response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;amazon-bedrock-trace&#34;</span>, {})
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;blocked&#34;</span>: <span style="color:#66d9ef">True</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;reason&#34;</span>: guardrail_trace<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;guardrail&#34;</span>, {})<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;actionReason&#34;</span>, <span style="color:#e6db74">&#34;Policy violation&#34;</span>),
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;response&#34;</span>: <span style="color:#66d9ef">None</span>
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;blocked&#34;</span>: <span style="color:#66d9ef">False</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;reason&#34;</span>: <span style="color:#66d9ef">None</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;response&#34;</span>: result[<span style="color:#e6db74">&#34;content&#34;</span>][<span style="color:#ae81ff">0</span>][<span style="color:#e6db74">&#34;text&#34;</span>]
</span></span><span style="display:flex;"><span>    }
</span></span></code></pre></div><p>One thing worth calling out: the guardContent qualifier on the user message tells Guardrails which part of the prompt to evaluate for the relevance check. Without it, Guardrails would try to evaluate your entire system prompt (including the retrieved documents) as if it were the user query — which produces noisy results.</p>
<h2 id="the-audit-trail">The Audit Trail</h2>
<p>A guardrail that blocks requests is only half the picture. The other half is knowing what it blocked, when, and why.</p>
<p>Every Guardrail invocation emits metrics to Amazon CloudWatch:</p>
<ul>
<li><code>GuardrailInvocations</code> — total count</li>
<li><code>GuardrailInterventions</code> — how many were blocked</li>
<li><code>GuardrailIntervention[PolicyType]</code> — breakdowns by policy</li>
</ul>
<p>Set up a CloudWatch alarm on GuardrailInterventions spiking above your baseline and you have an early warning system for adversarial use or misconfigured prompts.</p>
<p>For a more complete audit trail — the kind that satisfies an FFIEC examiner or a SOC 2 auditor — route blocked events through EventBridge to an S3 bucket and query them with Athena. The pattern looks like this:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># Lambda function triggered by EventBridge rule on Bedrock Guardrail events</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> datetime <span style="color:#f92672">import</span> datetime
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>s3 <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#34;s3&#34;</span>)
</span></span><span style="display:flex;"><span>AUDIT_BUCKET <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;your-ai-audit-logs-bucket&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Log guardrail interventions to immutable S3 audit trail.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    audit_record <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;timestamp&#34;</span>: datetime<span style="color:#f92672">.</span>utcnow()<span style="color:#f92672">.</span>isoformat(),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;guardrail_id&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;guardrailId&#34;</span>),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;policy_triggered&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;policyType&#34;</span>),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;action&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;action&#34;</span>),  <span style="color:#75715e"># BLOCKED or ANONYMIZED</span>
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;session_id&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;sessionId&#34;</span>),
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Do NOT log the raw prompt — PII may not be fully redacted at this point</span>
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;prompt_token_count&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;inputTokenCount&#34;</span>),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;region&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;awsRegion&#34;</span>),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;model_id&#34;</span>: event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#34;modelId&#34;</span>)
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    key <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;guardrail-interventions/</span><span style="color:#e6db74">{</span>datetime<span style="color:#f92672">.</span>utcnow()<span style="color:#f92672">.</span>strftime(<span style="color:#e6db74">&#39;%Y/%m/</span><span style="color:#e6db74">%d</span><span style="color:#e6db74">&#39;</span>)<span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>context<span style="color:#f92672">.</span>aws_request_id<span style="color:#e6db74">}</span><span style="color:#e6db74">.json&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    s3<span style="color:#f92672">.</span>put_object(
</span></span><span style="display:flex;"><span>        Bucket<span style="color:#f92672">=</span>AUDIT_BUCKET,
</span></span><span style="display:flex;"><span>        Key<span style="color:#f92672">=</span>key,
</span></span><span style="display:flex;"><span>        Body<span style="color:#f92672">=</span>json<span style="color:#f92672">.</span>dumps(audit_record),
</span></span><span style="display:flex;"><span>        ServerSideEncryption<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;aws:kms&#34;</span>,
</span></span><span style="display:flex;"><span>        BucketKeyEnabled<span style="color:#f92672">=</span><span style="color:#66d9ef">True</span>
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#34;statusCode&#34;</span>: <span style="color:#ae81ff">200</span>}
</span></span></code></pre></div><p>This gives you an immutable, KMS-encrypted log of every guardrail intervention — queryable by date, policy type, model, or session ID without ever storing the raw prompt content.</p>
<h2 id="what-this-actually-changes-for-banks">What This Actually Changes for Banks</h2>
<p>I keep coming back to one question in these conversations: what does it take to get an AI project from a successful proof-of-concept to a production system a compliance officer will sign off on?</p>
<p>The answer usually involves four things: data isolation, access controls, auditability, and behavioral controls. The first three are solved problems on AWS — VPC endpoints, IAM, CloudTrail. The fourth one — actually constraining what the model says and does — has historically required custom application logic that&rsquo;s brittle, hard to test, and invisible to your governance team.</p>
<p>Bedrock Guardrails changes that. It gives you behavioral controls that are:</p>
<ul>
<li><strong>Centralized.</strong> One guardrail definition applied consistently across every invocation, every session, every user.</li>
<li><strong>Versioned.</strong> You can pin a guardrail version to your production deployment and test changes in a draft version before promoting.</li>
<li><strong>Auditable.</strong> Every intervention is observable through CloudWatch metrics and loggable through EventBridge.</li>
<li><strong>Model-agnostic.</strong> The ApplyGuardrail API works independently of the foundation model — you can apply your guardrail to Claude, Titan, Llama, and even third-party models outside of Bedrock through the standalone API.</li>
</ul>
<p>That last point matters more than it sounds. Most banks aren&rsquo;t going to standardize on a single foundation model. As the model landscape evolves, your safety policies shouldn&rsquo;t have to be rewritten every time you swap out the underlying model.</p>
<h2 id="getting-started">Getting Started</h2>
<p>The fastest way to get a guardrail running is through the AWS console — there&rsquo;s a test playground in the Guardrails UI where you can paste prompts and verify your policies before deploying. Start there, calibrate your contextual grounding thresholds against real examples from your knowledge base, then export the configuration to Terraform or CloudFormation for your production deployment.
A few things to validate before go-live:</p>
<ul>
<li>Test your PII detection against real data samples (anonymized). The predefined entity types work well for standard formats; you&rsquo;ll discover gaps quickly with actual data.</li>
<li>Set your contextual grounding thresholds conservatively at first (0.7/0.7) and monitor your block rate. Too many false positives means end users get frustrated; too few means you&rsquo;re letting hallucinations through.</li>
<li>Verify your denied topics by trying to phrase a restricted question a dozen different ways. The topic detection is robust, but your definition matters — vague definitions lead to both over-blocking and under-blocking.</li>
</ul>
<p>If you&rsquo;re building in a regulated environment and you&rsquo;re not running Guardrails, you&rsquo;re carrying a liability that grows every day your AI system is in production. The capability exists. The question is whether you implement it before something goes wrong, or after.</p>
]]></content:encoded></item><item><title>Designing Pre-Trade Risk Controls on AWS (SEC Rule 15c3-5)</title><link>https://lukelittle.com/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/</link><pubDate>Sat, 14 Feb 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/</guid><description>How to implement real-time trading risk controls using Apache Kafka and Spark to meet SEC Rule 15c3-5 requirements, with lessons from the Knight Capital incident</description><content:encoded><![CDATA[<h2 id="introduction">Introduction</h2>
<p>On August 1, 2012, Knight Capital Group—one of the largest market makers on the New York Stock Exchange—lost $440 million in 45 minutes due to a software deployment failure. The incident nearly bankrupted the firm and sent shockwaves through financial markets. While the technical details are fascinating, the real lesson lies in what wasn&rsquo;t there: an effective, centralized mechanism to stop runaway automation before catastrophic losses occurred.</p>
<p>This post explores how modern streaming architectures using Apache Kafka and Apache Spark can implement the kind of real-time risk controls that regulations now require—and that Knight Capital desperately needed. We&rsquo;ll connect the dots between a historic trading disaster, regulatory requirements, and a hands-on demo you can deploy yourself.</p>
<h2 id="the-knight-capital-incident-what-happened">The Knight Capital Incident: What Happened?</h2>
<p>On that August morning, Knight Capital deployed new trading software to eight servers. Due to an operational error, one server retained old code that had been repurposed. When the market opened, this server began executing a dormant algorithm called &ldquo;Power Peg&rdquo; that was never meant to run in production.</p>
<p>The result was catastrophic:</p>
<ul>
<li>The algorithm sent millions of unintended orders to the market</li>
<li>Knight accumulated massive, unwanted positions in 154 stocks</li>
<li>The firm lost $440 million in approximately 45 minutes</li>
<li>Knight Capital required a $400 million emergency bailout to survive</li>
</ul>
<h3 id="the-core-failure-modes">The Core Failure Modes</h3>
<p>Several factors contributed to the disaster:</p>
<ol>
<li><strong>Partial Deployment</strong>: Not all servers received the correct code update</li>
<li><strong>Lack of Centralized Control</strong>: No single point could halt all trading activity</li>
<li><strong>Insufficient Pre-Trade Controls</strong>: Orders weren&rsquo;t validated against risk limits before execution</li>
<li><strong>Delayed Detection</strong>: The problem wasn&rsquo;t identified and stopped quickly enough</li>
</ol>
<p>The Knight Capital incident wasn&rsquo;t just a software bug—it was a systems design failure. The firm lacked the architectural patterns needed to maintain &ldquo;direct and exclusive control&rdquo; over its market access, a concept that would soon become central to regulatory requirements.</p>
<h2 id="enter-sec-rule-15c3-5-the-market-access-rule">Enter SEC Rule 15c3-5: The Market Access Rule</h2>
<p>In response to concerns about the risks posed by direct market access and algorithmic trading, the SEC adopted Rule 15c3-5 in November 2010 (before the Knight incident, though Knight&rsquo;s failure validated the rule&rsquo;s necessity).</p>
<h3 id="what-is-market-access">What is Market Access?</h3>
<p>Market access refers to the ability to send orders directly to exchanges or alternative trading systems. Broker-dealers that provide market access—whether for their own trading or for customers—act as gatekeepers to the markets.</p>
<h3 id="what-the-rule-requires">What the Rule Requires</h3>
<p>SEC Rule 15c3-5, formally titled &ldquo;Risk Management Controls for Brokers or Dealers with Market Access,&rdquo; requires broker-dealers to:</p>
<ol>
<li>
<p><strong>Implement Risk Management Controls</strong>: Establish, document, and maintain a system of risk management controls and supervisory procedures reasonably designed to manage the financial, regulatory, and other risks of market access.</p>
</li>
<li>
<p><strong>Pre-Trade Controls</strong>: Implement controls that prevent the entry of orders that exceed appropriate pre-set credit or capital thresholds, or that appear to be erroneous.</p>
</li>
<li>
<p><strong>Direct and Exclusive Control</strong>: Broker-dealers must have &ldquo;direct and exclusive control&rdquo; over the technology that provides market access. This means they cannot delegate control to customers or third parties—they must retain the ability to stop trading immediately.</p>
</li>
<li>
<p><strong>Regular Review</strong>: Controls must be reviewed and tested regularly to ensure they&rsquo;re working as intended.</p>
</li>
</ol>
<h3 id="the-direct-and-exclusive-control-concept">The &ldquo;Direct and Exclusive Control&rdquo; Concept</h3>
<p>This phrase is critical. It means:</p>
<ul>
<li>The broker-dealer must be able to disable or limit market access immediately</li>
<li>Control cannot be delegated to customers or outsourced</li>
<li>There must be a centralized mechanism to enforce risk limits</li>
<li>The firm must maintain supervisory procedures over all market access</li>
</ul>
<p>The rule doesn&rsquo;t prescribe specific technologies (it doesn&rsquo;t say &ldquo;you must use Kafka&rdquo;), but it does mandate capabilities that modern streaming architectures are well-suited to provide.</p>
<h3 id="regulatory-text">Regulatory Text</h3>
<p>From the SEC&rsquo;s adopting release:</p>
<blockquote>
<p>&ldquo;The rule requires a broker-dealer with market access to establish, document, and maintain a system of risk management controls and supervisory procedures reasonably designed to manage the financial, regulatory, and other risks of this business activity.&rdquo;</p>
</blockquote>
<p>The rule specifically addresses:</p>
<ul>
<li>Financial risk management (credit and capital thresholds)</li>
<li>Regulatory risk management (compliance with regulatory requirements)</li>
<li>Erroneous order controls (preventing clearly erroneous orders from reaching the market)</li>
</ul>
<h2 id="connecting-regulation-to-architecture">Connecting Regulation to Architecture</h2>
<p>Let&rsquo;s translate regulatory requirements into architectural patterns:</p>
<table>
  <thead>
      <tr>
          <th>Regulatory Requirement</th>
          <th>Architectural Pattern</th>
          <th>Our Demo Implementation</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Direct and exclusive control</td>
          <td>Centralized kill switch with authoritative state</td>
          <td>Kafka compacted topic for kill state</td>
      </tr>
      <tr>
          <td>Pre-trade risk controls</td>
          <td>Real-time order validation before routing</td>
          <td>Order router checks kill state</td>
      </tr>
      <tr>
          <td>Prevent erroneous orders</td>
          <td>Automated detection of anomalous patterns</td>
          <td>Spark streaming detects threshold breaches</td>
      </tr>
      <tr>
          <td>Supervisory procedures</td>
          <td>Audit trail and manual override capability</td>
          <td>Audit topic + operator console API</td>
      </tr>
      <tr>
          <td>Regular review and testing</td>
          <td>Observable, testable system</td>
          <td>CloudWatch dashboards + demo scripts</td>
      </tr>
  </tbody>
</table>
<p>The key insight: <strong>Separation of concerns between detection and enforcement</strong>.</p>
<ul>
<li><strong>Detection</strong> (Spark): Analyzes order patterns, computes risk signals, may suggest kill actions</li>
<li><strong>Enforcement</strong> (Router): Makes the final decision on every order based on authoritative state</li>
<li><strong>Control Plane</strong> (Kill Switch): Maintains single source of truth for kill status</li>
<li><strong>Audit</strong> (Kafka + DynamoDB): Immutable record of all decisions</li>
</ul>
<p>This separation ensures that even if detection fails or is delayed, enforcement remains consistent. The kill switch state is authoritative and replayable.</p>
<h2 id="why-kafka-compaction-for-kill-state">Why Kafka Compaction for Kill State?</h2>
<p>One of the most interesting architectural choices in our demo is using a Kafka compacted topic for kill switch state. Here&rsquo;s why:</p>
<h3 id="the-problem">The Problem</h3>
<p>We need a &ldquo;configuration&rdquo; or &ldquo;state&rdquo; that:</p>
<ul>
<li>Is the single source of truth</li>
<li>Can be updated in real-time</li>
<li>Is immediately available to all consumers</li>
<li>Has a complete audit trail</li>
<li>Can be replayed to bootstrap new services</li>
</ul>
<h3 id="the-solution-log-compaction">The Solution: Log Compaction</h3>
<p>Kafka&rsquo;s log compaction retains the latest value for each key while preserving the full history of changes. For kill switch state:</p>
<pre tabindex="0"><code>Compaction Example:

Commands Topic (Full History):
ACCOUNT:12345 KILL   t=100
ACCOUNT:12345 UNKILL t=200
ACCOUNT:12345 KILL   t=300

State Topic (Compacted - Latest Only):
ACCOUNT:12345 KILL   t=300
</code></pre><p><strong>How it works:</strong></p>
<ul>
<li><strong>Key</strong>: Scope (e.g., &ldquo;ACCOUNT:12345&rdquo;, &ldquo;SYMBOL:AAPL&rdquo;, &ldquo;GLOBAL&rdquo;)</li>
<li><strong>Value</strong>: Current status (KILLED or ACTIVE) with metadata</li>
<li><strong>Compaction</strong>: Kafka automatically retains only the latest state per scope</li>
<li><strong>Replayability</strong>: New consumers can read the entire topic to bootstrap current state</li>
</ul>
<p>This gives us:</p>
<ol>
<li><strong>Single source of truth</strong>: The compacted topic is authoritative</li>
<li><strong>Fast bootstrap</strong>: New routers can quickly load all current kill states</li>
<li><strong>Audit trail</strong>: The commands topic retains full history</li>
<li><strong>Distributed config</strong>: No need for external config store</li>
</ol>
<h3 id="state-compaction-process">State Compaction Process</h3>
<p>The following sequence diagram illustrates how Kafka&rsquo;s log compaction maintains the latest state per scope:</p>
<pre class="mermaid">sequenceDiagram
    participant K1 as Kafka&lt;br/&gt;killswitch.commands.v1&lt;br/&gt;(Full History)
    participant KSA as Kill Switch&lt;br/&gt;Aggregator
    participant K2 as Kafka&lt;br/&gt;killswitch.state.v1&lt;br/&gt;(Compacted)
    participant KC as Kafka&lt;br/&gt;Compaction Process
    participant OR as Order Router&lt;br/&gt;(New Instance)

    Note over K1,OR: State Evolution Over Time
    
    Note over K1: t=100
    K1-&gt;&gt;KSA: KILL command&lt;br/&gt;ACCOUNT:12345
    KSA-&gt;&gt;K2: Publish state&lt;br/&gt;Key: ACCOUNT:12345&lt;br/&gt;Value: KILLED (t=100)
    
    Note over K1: t=200
    K1-&gt;&gt;KSA: UNKILL command&lt;br/&gt;ACCOUNT:12345
    KSA-&gt;&gt;K2: Publish state&lt;br/&gt;Key: ACCOUNT:12345&lt;br/&gt;Value: ACTIVE (t=200)
    
    Note over K1: t=300
    K1-&gt;&gt;KSA: KILL command&lt;br/&gt;ACCOUNT:12345
    KSA-&gt;&gt;K2: Publish state&lt;br/&gt;Key: ACCOUNT:12345&lt;br/&gt;Value: KILLED (t=300)
    
    Note over K2: Before Compaction:&lt;br/&gt;ACCOUNT:12345 KILLED (t=100)&lt;br/&gt;ACCOUNT:12345 ACTIVE (t=200)&lt;br/&gt;ACCOUNT:12345 KILLED (t=300)
    
    K2-&gt;&gt;KC: Compaction triggered&lt;br/&gt;(based on segment.ms&lt;br/&gt;and dirty ratio)
    
    KC-&gt;&gt;KC: Retain latest value&lt;br/&gt;per key
    
    Note over K2: After Compaction:&lt;br/&gt;ACCOUNT:12345 KILLED (t=300)&lt;br/&gt;(older values removed)
    
    Note over OR: New router starts up
    OR-&gt;&gt;K2: Read from beginning
    K2--&gt;&gt;OR: ACCOUNT:12345 = KILLED (t=300)
    
    Note over OR: Router bootstrapped&lt;br/&gt;with current state&lt;br/&gt;(fast, no history to read)
    
    Note over K1: Commands topic still has&lt;br/&gt;full history for audit
</pre>

<h3 id="compaction-configuration">Compaction Configuration</h3>
<pre tabindex="0"><code>cleanup.policy=compact
min.cleanable.dirty.ratio=0.01  # Compact frequently
segment.ms=60000                # Small segments for faster compaction
</code></pre><p>These settings ensure kill state updates propagate quickly while maintaining the full history in the commands topic.</p>
<h2 id="demo-architecture-walkthrough">Demo Architecture Walkthrough</h2>
<p>Our demo implements these patterns using serverless AWS services:</p>
<p><img src="/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/designing-pre-trade-risk-controls-on-aws_hu_9bc266b9d36092e9.webp" srcset="/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/designing-pre-trade-risk-controls-on-aws_hu_cfccd3f5cdfc5123.webp 750w, /posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/designing-pre-trade-risk-controls-on-aws_hu_9bc266b9d36092e9.webp 1320w" sizes="(max-width: 800px) 100vw, 750px"
       width="1320" height="1340" alt="Architecture diagram showing pre-trade risk controls on AWS" loading="lazy" decoding="async"></p>
<p><strong>Key architectural decisions mapped to regulatory requirements:</strong></p>
<table>
  <thead>
      <tr>
          <th>Requirement</th>
          <th>Implementation</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Direct and exclusive control</td>
          <td>Operator Console API with manual override capability</td>
      </tr>
      <tr>
          <td>Pre-trade risk controls</td>
          <td>Order Router checks kill state before routing every order</td>
      </tr>
      <tr>
          <td>Prevent erroneous orders</td>
          <td>Spark detects anomalous patterns in real-time</td>
      </tr>
      <tr>
          <td>Audit trail</td>
          <td>Immutable Kafka log + DynamoDB index for queries</td>
      </tr>
      <tr>
          <td>Supervisory procedures</td>
          <td>Documented thresholds, operator actions, correlation IDs</td>
      </tr>
  </tbody>
</table>
<h3 id="normal-order-flow">Normal Order Flow</h3>
<p>The following sequence diagram illustrates how orders flow through the system when no kill switches are active:</p>
<pre class="mermaid">sequenceDiagram
    participant OG as Order Generator&lt;br/&gt;(Lambda)
    participant K1 as Kafka&lt;br/&gt;orders.v1
    participant S as Spark&lt;br/&gt;Risk Detector
    participant K2 as Kafka&lt;br/&gt;risk_signals.v1
    participant OR as Order Router&lt;br/&gt;(Lambda)
    participant KS as Kafka&lt;br/&gt;killswitch.state.v1
    participant K3 as Kafka&lt;br/&gt;orders.gated.v1
    participant K4 as Kafka&lt;br/&gt;audit.v1
    participant DDB as DynamoDB&lt;br/&gt;Audit Index

    Note over OG,DDB: Normal Operation - No Kill Switches Active
    
    OG-&gt;&gt;K1: Publish order&lt;br/&gt;(5 orders/sec)
    Note right of K1: Key: account_id&lt;br/&gt;Partition by account
    
    K1-&gt;&gt;S: Consume orders
    S-&gt;&gt;S: Compute 60s window&lt;br/&gt;order_count = 50&lt;br/&gt;notional = $500K
    Note right of S: Below thresholds:&lt;br/&gt;order_rate &lt; 100&lt;br/&gt;notional &lt; $1M
    S-&gt;&gt;K2: Publish risk signal&lt;br/&gt;(metrics only, no alert)
    
    K1-&gt;&gt;OR: Consume order
    OR-&gt;&gt;KS: Check kill state&lt;br/&gt;for ACCOUNT:12345
    KS--&gt;&gt;OR: No kill state found&lt;br/&gt;(ACTIVE by default)
    
    Note over OR: Decision: ALLOW
    
    OR-&gt;&gt;K3: Forward order&lt;br/&gt;to gated topic
    OR-&gt;&gt;K4: Publish audit event&lt;br/&gt;decision=ALLOW
    OR-&gt;&gt;DDB: Write audit record&lt;br/&gt;(async, best effort)
    
    Note over OG,DDB: Order successfully routed
</pre>

<h3 id="topic-flow">Topic Flow</h3>
<ol>
<li><strong>orders.v1</strong>: Raw orders from generator</li>
<li><strong>risk_signals.v1</strong>: Windowed aggregations from Spark (order rate, notional, concentration)</li>
<li><strong>killswitch.commands.v1</strong>: Kill/unkill commands (from Spark or operator)</li>
<li><strong>killswitch.state.v1</strong>: Authoritative kill state (compacted)</li>
<li><strong>orders.gated.v1</strong>: Orders that passed kill switch check</li>
<li><strong>audit.v1</strong>: Immutable audit trail of all routing decisions</li>
</ol>
<h3 id="order-router-enforcement">Order Router Enforcement</h3>
<p>The following diagram shows the detailed logic of how the order router enforces kill switches:</p>
<pre class="mermaid">sequenceDiagram
    participant K1 as Kafka&lt;br/&gt;orders.v1
    participant OR as Order Router&lt;br/&gt;(Lambda)
    participant Cache as In-Memory&lt;br/&gt;Kill State Cache
    participant K2 as Kafka&lt;br/&gt;killswitch.state.v1
    participant K3 as Kafka&lt;br/&gt;orders.gated.v1
    participant K4 as Kafka&lt;br/&gt;audit.v1
    participant DDB as DynamoDB&lt;br/&gt;Audit Index

    Note over K1,DDB: Order Router Processing Logic
    
    K1-&gt;&gt;OR: Consume order&lt;br/&gt;account_id: 12345&lt;br/&gt;symbol: AAPL
    
    OR-&gt;&gt;Cache: Check kill state&lt;br/&gt;for scopes
    
    Note over Cache: Check hierarchy:&lt;br/&gt;1. GLOBAL&lt;br/&gt;2. ACCOUNT:12345&lt;br/&gt;3. SYMBOL:AAPL
    
    alt GLOBAL kill active
        Cache--&gt;&gt;OR: GLOBAL = KILLED
        Note over OR: Decision: DROP&lt;br/&gt;Reason: Global kill
    else ACCOUNT kill active
        Cache--&gt;&gt;OR: ACCOUNT:12345 = KILLED
        Note over OR: Decision: DROP&lt;br/&gt;Reason: Account kill
    else SYMBOL kill active
        Cache--&gt;&gt;OR: SYMBOL:AAPL = KILLED
        Note over OR: Decision: DROP&lt;br/&gt;Reason: Symbol kill
    else No kills active
        Cache--&gt;&gt;OR: All scopes ACTIVE
        Note over OR: Decision: ALLOW
        OR-&gt;&gt;K3: Forward order
    end
    
    OR-&gt;&gt;K4: Publish audit event&lt;br/&gt;decision: ALLOW/DROP&lt;br/&gt;scope_matches: [...]&lt;br/&gt;corr_id: uuid-789
    
    OR-&gt;&gt;DDB: Write audit record&lt;br/&gt;(async, best effort)
    
    Note over K2,OR: State updates arrive
    K2-&gt;&gt;OR: New state update
    OR-&gt;&gt;Cache: Update in-memory cache
    
    Note over Cache: Cache always reflects&lt;br/&gt;latest compacted state
</pre>

<h3 id="latency-considerations">Latency Considerations</h3>
<p>This design introduces additional latency relative to in-process risk checks. However, it provides centralized, authoritative enforcement and replayable state — properties essential for satisfying the &ldquo;direct and exclusive control&rdquo; requirement of SEC Rule 15c3-5.</p>
<p>In most retail and DMA (Direct Market Access) environments, the added milliseconds (approximately 50ms at most) are an acceptable tradeoff for deterministic control and auditability. This approach is not intended to reflect the architecture of any former employer but rather examines how brokerages can solve these regulatory challenges in a robust, scalable way.</p>
<h3 id="kill-switch-activation-sequence">Kill Switch Activation Sequence</h3>
<p>Here&rsquo;s what happens when Spark detects a threshold breach:</p>
<pre class="mermaid">sequenceDiagram
    participant OG as Order Generator&lt;br/&gt;(Lambda)
    participant K1 as Kafka&lt;br/&gt;orders.v1
    participant S as Spark&lt;br/&gt;Risk Detector
    participant K2 as Kafka&lt;br/&gt;risk_signals.v1
    participant K3 as Kafka&lt;br/&gt;killswitch.commands.v1
    participant KSA as Kill Switch&lt;br/&gt;Aggregator (Lambda)
    participant K4 as Kafka&lt;br/&gt;killswitch.state.v1
    participant DDB as DynamoDB&lt;br/&gt;State Cache
    participant OR as Order Router&lt;br/&gt;(Lambda)
    participant K5 as Kafka&lt;br/&gt;audit.v1

    Note over OG,K5: Panic Mode Triggered
    
    OG-&gt;&gt;K1: Publish orders&lt;br/&gt;(50 orders/sec)
    Note right of K1: High rate for&lt;br/&gt;ACCOUNT:12345
    
    K1-&gt;&gt;S: Consume orders
    S-&gt;&gt;S: Compute 60s window&lt;br/&gt;order_count = 150&lt;br/&gt;notional = $2.5M
    
    Note over S: BREACH DETECTED!&lt;br/&gt;order_count &gt; 100
    
    S-&gt;&gt;K2: Publish risk signal&lt;br/&gt;with breach flag
    S-&gt;&gt;K3: Publish KILL command&lt;br/&gt;scope: ACCOUNT:12345&lt;br/&gt;reason: &#34;Order rate breach&#34;&lt;br/&gt;corr_id: uuid-123
    
    Note over K3: Commands topic&lt;br/&gt;(full history retained)
    
    K3-&gt;&gt;KSA: Consume KILL command
    KSA-&gt;&gt;KSA: Process command&lt;br/&gt;Create state record
    
    KSA-&gt;&gt;K4: Publish state&lt;br/&gt;Key: ACCOUNT:12345&lt;br/&gt;Value: KILLED&lt;br/&gt;corr_id: uuid-123
    Note right of K4: Compacted topic&lt;br/&gt;(latest state per key)
    
    KSA-&gt;&gt;DDB: Update state cache&lt;br/&gt;(optional, for fast lookup)
    
    Note over K4,OR: State propagates to all routers
    
    K4-&gt;&gt;OR: Router reads state update
    OR-&gt;&gt;OR: Update in-memory cache&lt;br/&gt;ACCOUNT:12345 = KILLED
    
    K1-&gt;&gt;OR: New order from 12345
    OR-&gt;&gt;OR: Check kill state&lt;br/&gt;ACCOUNT:12345 = KILLED
    
    Note over OR: Decision: DROP
    
    OR-&gt;&gt;K5: Publish audit event&lt;br/&gt;decision=DROP&lt;br/&gt;reason: &#34;Kill switch active&#34;&lt;br/&gt;corr_id: uuid-123
    
    Note over OG,K5: Order blocked - Kill switch active
</pre>

<p><strong>Key observations:</strong></p>
<ul>
<li>Detection (Spark) is decoupled from enforcement (Router)</li>
<li>State updates flow through compacted topic (single source of truth)</li>
<li>Every decision is audited with correlation IDs</li>
<li>Manual override capability (operator can unkill)</li>
</ul>
<h3 id="spark-sql-for-risk-detection">Spark SQL for Risk Detection</h3>
<p>The Spark job uses Spark SQL for windowed aggregations:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-sql" data-lang="sql"><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span>
</span></span><span style="display:flex;"><span>    window(event_time, <span style="color:#e6db74">&#39;60 seconds&#39;</span>) <span style="color:#66d9ef">as</span> window,
</span></span><span style="display:flex;"><span>    account_id,
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">COUNT</span>(<span style="color:#f92672">*</span>) <span style="color:#66d9ef">as</span> order_count,
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">SUM</span>(qty <span style="color:#f92672">*</span> price) <span style="color:#66d9ef">as</span> total_notional,
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">COUNT</span>(<span style="color:#66d9ef">DISTINCT</span> symbol) <span style="color:#66d9ef">as</span> unique_symbols
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">FROM</span> orders
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">GROUP</span> <span style="color:#66d9ef">BY</span> window(event_time, <span style="color:#e6db74">&#39;60 seconds&#39;</span>), account_id
</span></span></code></pre></div><p>When thresholds are breached, Spark emits a kill command:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;cmd_id&#34;</span>: <span style="color:#e6db74">&#34;uuid&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;scope&#34;</span>: <span style="color:#e6db74">&#34;ACCOUNT:12345&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;action&#34;</span>: <span style="color:#e6db74">&#34;KILL&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;reason&#34;</span>: <span style="color:#e6db74">&#34;Order rate breach: 150 orders in 60s&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;triggered_by&#34;</span>: <span style="color:#e6db74">&#34;spark&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;metric&#34;</span>: <span style="color:#e6db74">&#34;order_rate_60s&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;value&#34;</span>: <span style="color:#ae81ff">150</span>
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><h3 id="enforcement-logic">Enforcement Logic</h3>
<p>The order router maintains an in-memory cache of kill state (bootstrapped from the compacted topic) and checks every order:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">check_kill_status</span>(order):
</span></span><span style="display:flex;"><span>    scopes <span style="color:#f92672">=</span> [
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;GLOBAL&#39;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">f</span><span style="color:#e6db74">&#39;ACCOUNT:</span><span style="color:#e6db74">{</span>order[<span style="color:#e6db74">&#34;account_id&#34;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&#39;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">f</span><span style="color:#e6db74">&#39;SYMBOL:</span><span style="color:#e6db74">{</span>order[<span style="color:#e6db74">&#34;symbol&#34;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&#39;</span>
</span></span><span style="display:flex;"><span>    ]
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">for</span> scope <span style="color:#f92672">in</span> scopes:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> scope <span style="color:#f92672">in</span> kill_state <span style="color:#f92672">and</span> kill_state[scope][<span style="color:#e6db74">&#39;status&#39;</span>] <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;KILLED&#39;</span>:
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span> <span style="color:#66d9ef">True</span>, scope, kill_state[scope][<span style="color:#e6db74">&#39;reason&#39;</span>]
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#66d9ef">False</span>, <span style="color:#66d9ef">None</span>, <span style="color:#66d9ef">None</span>
</span></span></code></pre></div><p>Every decision is audited with correlation IDs for traceability.</p>
<h2 id="alternative-approaches-for-high-frequency-trading">Alternative Approaches for High-Frequency Trading</h2>
<p>For high-frequency trading environments, the architecture described above would introduce unacceptable latency. In these ultra-low-latency scenarios, the pre-trade risk gate would be embedded directly in the order handling process—potentially implemented in hardware (e.g., FPGA)—to ensure deterministic, microsecond-level enforcement without introducing network or broker latency.</p>
<p>Key differences in HFT implementations:</p>
<ul>
<li><strong>Embedded Controls</strong>: Risk checks directly in the order path, not as external services</li>
<li><strong>Hardware Acceleration</strong>: FPGAs or dedicated ASICs for microsecond-level checks</li>
<li><strong>Local State</strong>: State maintained in local memory with minimal or no network calls</li>
<li><strong>Minimal Serialization</strong>: Custom binary protocols instead of JSON</li>
<li><strong>Deterministic Performance</strong>: Bounded, predictable latency for all operations</li>
</ul>
<p>In such environments, you wouldn&rsquo;t add Kafka to the hot path, as even the most optimized message broker would introduce unacceptable latency. Instead, while still maintaining the regulatory requirements for &ldquo;direct and exclusive control,&rdquo; risk configurations would be loaded at startup and updated via side channels, with enforcement happening directly within the order processing pipeline.</p>
<h2 id="why-this-matters-for-students">Why This Matters for Students</h2>
<p>This demo teaches several critical concepts:</p>
<ol>
<li><strong>Event-Driven Architecture</strong>: Using Kafka as the backbone for real-time systems</li>
<li><strong>Stream Processing</strong>: Spark Structured Streaming for windowed aggregations</li>
<li><strong>Separation of Concerns</strong>: Detection vs. enforcement vs. control</li>
<li><strong>Operational Patterns</strong>: Compaction, idempotency, correlation IDs</li>
<li><strong>Regulatory Thinking</strong>: How compliance requirements shape architecture</li>
<li><strong>Serverless at Scale</strong>: Building production-grade systems without managing servers</li>
</ol>
<p>Most importantly, it connects abstract concepts (regulations, risk management) to concrete implementations you can deploy and experiment with.</p>
<h2 id="try-it-yourself">Try It Yourself</h2>
<p>The complete demo is available in the repository. You can:</p>
<ol>
<li><strong>Deploy to AWS</strong>: Full serverless stack with Terraform</li>
<li><strong>Run locally</strong>: Docker Compose for quick iteration</li>
<li><strong>Experiment</strong>: Change thresholds, add new scopes, implement throttling</li>
<li><strong>Learn</strong>: Detailed workshop docs with exercises</li>
</ol>
<p>See the <a href="https://github.com/lukelittle/sec-15c3-5-market-access-controls-example">repository</a> for step-by-step instructions.</p>
<h3 id="suggested-exercises">Suggested Exercises</h3>
<ol>
<li>Add a SYMBOL-level kill switch that triggers on concentration</li>
<li>Implement throttling (rate limiting) instead of binary kill/allow</li>
<li>Add deduplication to prevent duplicate order IDs</li>
<li>Build a dashboard to visualize risk signals in real-time</li>
<li>Implement automatic unkill after a cooldown period</li>
</ol>
<h2 id="conclusion">Conclusion</h2>
<p>The Knight Capital incident taught the industry a painful lesson about the importance of centralized control and pre-trade risk management. SEC Rule 15c3-5 codified these lessons into regulatory requirements that all broker-dealers must follow.</p>
<p>Modern streaming architectures using Kafka and Spark provide elegant solutions to these requirements:</p>
<ul>
<li>Kafka&rsquo;s compacted topics give us authoritative, replayable state</li>
<li>Spark&rsquo;s streaming SQL enables real-time risk detection</li>
<li>Separation of detection and enforcement ensures consistent control</li>
<li>Immutable audit trails provide full traceability</li>
</ul>
<p>While this demo uses synthetic data and simplified logic, the architectural patterns are production-grade. Real broker-dealers use similar approaches to maintain the &ldquo;direct and exclusive control&rdquo; that regulations require and that Knight Capital lacked.</p>
<p>The next time you hear about a trading glitch or market disruption, ask: &ldquo;Where was the kill switch?&rdquo;</p>
<h2 id="sources-and-further-reading">Sources and Further Reading</h2>
<h3 id="primary-regulatory-sources">Primary Regulatory Sources</h3>
<ol>
<li>
<p><strong>SEC Rule 15c3-5 Final Adopting Release</strong><br>
Securities and Exchange Commission, Release No. 34-63241 (November 3, 2010)<br>
<a href="https://www.sec.gov/files/rules/final/2010/34-63241.pdf">https://www.sec.gov/files/rules/final/2010/34-63241.pdf</a></p>
</li>
<li>
<p><strong>SEC Small Entity Compliance Guide for Rule 15c3-5</strong><br>
<a href="https://www.sec.gov/files/rules/final/2010/34-63241-secg.htm">https://www.sec.gov/files/rules/final/2010/34-63241-secg.htm</a></p>
</li>
<li>
<p><strong>Code of Federal Regulations: 17 CFR § 240.15c3-5</strong><br>
<a href="https://www.law.cornell.edu/cfr/text/17/240.15c3-5">https://www.law.cornell.edu/cfr/text/17/240.15c3-5</a></p>
</li>
</ol>
<h3 id="knight-capital-incident">Knight Capital Incident</h3>
<ol start="4">
<li>
<p><strong>SEC Administrative Proceeding Against Knight Capital</strong><br>
File No. 3-15570 (October 16, 2013)<br>
Details the regulatory findings and penalties related to the incident.</p>
</li>
<li>
<p><strong>Nanex Research: Knight Capital&rsquo;s Trading Glitch</strong><br>
Technical analysis of the order flow during the incident (secondary source).</p>
</li>
</ol>
<h3 id="technical-resources">Technical Resources</h3>
<ol start="6">
<li>
<p><strong>Apache Kafka Documentation: Log Compaction</strong><br>
<a href="https://kafka.apache.org/documentation/#compaction">https://kafka.apache.org/documentation/#compaction</a></p>
</li>
<li>
<p><strong>Apache Spark Structured Streaming Guide</strong><br>
<a href="https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html">https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html</a></p>
</li>
</ol>
<hr>
<p><strong>Disclaimer</strong>: This blog post and associated demo are for educational purposes only. They do not constitute trading advice, legal advice, or compliance guidance. The architecture described does not represent any former employer&rsquo;s actual systems or implementations. The demo uses synthetic data and simplified logic to illustrate concepts rather than real production implementations. Actual production trading systems require extensive additional controls, testing, and regulatory review. Always consult with legal and compliance professionals when implementing market access systems.</p>
]]></content:encoded></item><item><title>Building a Cost Optimization Agent with AWS Bedrock and Cost Explorer</title><link>https://lukelittle.com/posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/</link><pubDate>Fri, 13 Feb 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/</guid><description>How to build a cost optimization agent that analyzes your AWS spend and provides actionable recommendations to reduce your bill</description><content:encoded><![CDATA[<p>Managing AWS costs becomes increasingly complex as infrastructure grows. Organizations often struggle with cloud cost management, spending valuable engineering time manually analyzing Cost Explorer data, identifying optimization opportunities, and implementing changes. Even with dedicated cost management tools, the analysis and remediation process remains largely manual, requiring specialized expertise to interpret cost data and translate it into actionable steps.</p>
<p>This post demonstrates how to build an automated agent that analyzes AWS costs and generates actionable recommendations to reduce cloud spend. By combining AWS Bedrock&rsquo;s analytical capabilities with Cost Explorer data, the system identifies cost outliers and provides specific optimization steps that go beyond basic visualizations to deliver meaningful insights.</p>
<h2 id="what-were-building">What we&rsquo;re building</h2>
<p>A cost optimization agent that:</p>
<ol>
<li>Runs on a schedule (weekly or monthly)</li>
<li>Fetches detailed cost data from AWS Cost Explorer</li>
<li>Analyzes spending patterns and identifies optimization opportunities</li>
<li>Generates a markdown report with specific recommendations</li>
<li>Sends a summary to Slack or email</li>
<li>Tracks recommendations and their potential savings</li>
</ol>
<p>This solution goes beyond basic cost visualization by providing specific, actionable steps to optimize AWS spending—essentially turning data into decisions.</p>
<h2 id="real-world-applications">Real-World Applications</h2>
<p>This solution addresses cost management challenges across different contexts:</p>
<p><strong>Enterprise FinOps Teams</strong>: In larger organizations, the agent provides consistent, ongoing cost analysis that augments the FinOps team&rsquo;s capabilities, ensuring no optimization opportunity goes unnoticed even as the infrastructure grows in complexity.</p>
<p><strong>Startups and Small Teams</strong>: For organizations without dedicated cloud financial analysts, the agent provides expert-level cost optimization recommendations that would otherwise require specialized knowledge or expensive consultants.</p>
<p><strong>Managed Service Providers</strong>: MSPs can deploy the agent across client environments, standardizing cost optimization practices while customizing thresholds and priorities for each client&rsquo;s specific needs.</p>
<p><strong>Development Environments</strong>: The agent can enforce stricter cost controls in non-production environments, identifying development and testing resources that can be safely downsized, scheduled, or terminated to reduce costs without affecting production workloads.</p>
<p><strong>Multi-Cloud Strategies</strong>: While this implementation focuses on AWS, the architecture pattern can be extended to analyze costs across multiple cloud providers, giving organizations a unified view of optimization opportunities.</p>
<h2 id="architecture-overview">Architecture overview</h2>
<p><img src="/posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/aws-bedrock-cost-optimization-architecture_hu_6cc321b9ea36da9.webp" srcset="/posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/aws-bedrock-cost-optimization-architecture_hu_a0c182f575f5e013.webp 750w, /posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/aws-bedrock-cost-optimization-architecture_hu_6cc321b9ea36da9.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="1113" alt="Architecture diagram showing a cost optimization agent built with AWS Bedrock and Cost Explorer API" loading="lazy" decoding="async"></p>
<p>Here&rsquo;s the high-level architecture:</p>
<pre tabindex="0"><code>EventBridge (scheduled) → Lambda → Bedrock Agent with Action Groups → S3 (report) → SNS (notifications)
</code></pre><p>The key components:</p>
<ul>
<li><strong>EventBridge</strong>: Triggers the agent on a schedule</li>
<li><strong>Lambda</strong>: Initializes and coordinates the analysis process</li>
<li><strong>Bedrock Agent</strong>: Orchestrates the data gathering and analysis</li>
<li><strong>Action Groups</strong>: Custom Lambda functions for specific tasks</li>
<li><strong>S3</strong>: Stores the generated reports</li>
<li><strong>SNS</strong>: Sends notifications with the summary</li>
<li><strong>DynamoDB</strong>: Tracks recommendations and their implementation status</li>
</ul>
<p>This event-driven architecture ensures the cost optimization process runs automatically on schedule, eliminating the need for manual intervention and ensuring consistent analysis.</p>
<h2 id="what-youll-need">What you&rsquo;ll need</h2>
<ul>
<li>AWS Account with Bedrock and Cost Explorer access</li>
<li>IAM role with appropriate permissions</li>
<li>S3 bucket for storing reports</li>
<li>SNS topic or email for notifications</li>
<li>Basic understanding of AWS services</li>
</ul>
<h2 id="step-1-create-the-action-group-lambda-functions">Step 1: Create the Action Group Lambda functions</h2>
<p>First, let&rsquo;s create the Lambda functions our Bedrock Agent will use as action groups:</p>
<h3 id="1-cost-data-retrieval-lambda">1. Cost Data Retrieval Lambda</h3>
<p>This function fetches comprehensive cost data from AWS Cost Explorer:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> datetime
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> dateutil.relativedelta <span style="color:#f92672">import</span> relativedelta
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        time_period <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;time_period&#39;</span>, <span style="color:#e6db74">&#39;MONTH&#39;</span>)
</span></span><span style="display:flex;"><span>        services <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;services&#39;</span>, [])
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Configure time period based on request</span>
</span></span><span style="display:flex;"><span>        end_date <span style="color:#f92672">=</span> datetime<span style="color:#f92672">.</span>datetime<span style="color:#f92672">.</span>now()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> time_period <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;MONTH&#39;</span>:
</span></span><span style="display:flex;"><span>            start_date <span style="color:#f92672">=</span> end_date <span style="color:#f92672">-</span> relativedelta(months<span style="color:#f92672">=</span><span style="color:#ae81ff">1</span>)
</span></span><span style="display:flex;"><span>            granularity <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;DAILY&#39;</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> time_period <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;QUARTER&#39;</span>:
</span></span><span style="display:flex;"><span>            start_date <span style="color:#f92672">=</span> end_date <span style="color:#f92672">-</span> relativedelta(months<span style="color:#f92672">=</span><span style="color:#ae81ff">3</span>)
</span></span><span style="display:flex;"><span>            granularity <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;MONTHLY&#39;</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> time_period <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;WEEK&#39;</span>:
</span></span><span style="display:flex;"><span>            start_date <span style="color:#f92672">=</span> end_date <span style="color:#f92672">-</span> relativedelta(weeks<span style="color:#f92672">=</span><span style="color:#ae81ff">1</span>)
</span></span><span style="display:flex;"><span>            granularity <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;DAILY&#39;</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">else</span>:
</span></span><span style="display:flex;"><span>            <span style="color:#75715e"># Default to monthly</span>
</span></span><span style="display:flex;"><span>            start_date <span style="color:#f92672">=</span> end_date <span style="color:#f92672">-</span> relativedelta(months<span style="color:#f92672">=</span><span style="color:#ae81ff">1</span>)
</span></span><span style="display:flex;"><span>            granularity <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;DAILY&#39;</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Format dates for Cost Explorer</span>
</span></span><span style="display:flex;"><span>        start_str <span style="color:#f92672">=</span> start_date<span style="color:#f92672">.</span>strftime(<span style="color:#e6db74">&#39;%Y-%m-</span><span style="color:#e6db74">%d</span><span style="color:#e6db74">&#39;</span>)
</span></span><span style="display:flex;"><span>        end_str <span style="color:#f92672">=</span> end_date<span style="color:#f92672">.</span>strftime(<span style="color:#e6db74">&#39;%Y-%m-</span><span style="color:#e6db74">%d</span><span style="color:#e6db74">&#39;</span>)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Initialize Cost Explorer client</span>
</span></span><span style="display:flex;"><span>        ce_client <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;ce&#39;</span>)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Get cost data, anomalies, and recommendations</span>
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> get_cost_data(ce_client, start_str, end_str, granularity)
</span></span><span style="display:flex;"><span>        anomalies <span style="color:#f92672">=</span> get_anomalies(ce_client, start_str, end_str)
</span></span><span style="display:flex;"><span>        reservation_recs <span style="color:#f92672">=</span> get_reservation_recommendations(ce_client)
</span></span><span style="display:flex;"><span>        savings_plans_recs <span style="color:#f92672">=</span> get_savings_plans_recommendations(ce_client)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Return the compiled data</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;cost_data&#39;</span>: response,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;anomalies&#39;</span>: anomalies,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;reservation_recommendations&#39;</span>: reservation_recs,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;savings_plans_recommendations&#39;</span>: savings_plans_recs
</span></span><span style="display:flex;"><span>            })
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Helper functions (implementation details omitted for brevity)</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_cost_data</span>(ce_client, start_str, end_str, granularity):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_anomalies</span>(ce_client, start_str, end_str):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_reservation_recommendations</span>(ce_client):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_savings_plans_recommendations</span>(ce_client):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span></code></pre></div><h3 id="2-resource-analysis-lambda">2. Resource Analysis Lambda</h3>
<p>This function analyzes AWS resources for optimization opportunities:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> datetime
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        resource_types <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;resource_types&#39;</span>, [<span style="color:#e6db74">&#39;ec2&#39;</span>, <span style="color:#e6db74">&#39;rds&#39;</span>, <span style="color:#e6db74">&#39;ebs&#39;</span>, <span style="color:#e6db74">&#39;lambda&#39;</span>])
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        results <span style="color:#f92672">=</span> {}
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Analyze each resource type</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;ec2&#39;</span> <span style="color:#f92672">in</span> resource_types:
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;ec2&#39;</span>] <span style="color:#f92672">=</span> analyze_ec2_instances()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;rds&#39;</span> <span style="color:#f92672">in</span> resource_types:
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;rds&#39;</span>] <span style="color:#f92672">=</span> analyze_rds_instances()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;ebs&#39;</span> <span style="color:#f92672">in</span> resource_types:
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;ebs&#39;</span>] <span style="color:#f92672">=</span> analyze_ebs_volumes()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;lambda&#39;</span> <span style="color:#f92672">in</span> resource_types:
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;lambda&#39;</span>] <span style="color:#f92672">=</span> analyze_lambda_functions()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps(results)
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">analyze_ec2_instances</span>():
</span></span><span style="display:flex;"><span>    ec2_client <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;ec2&#39;</span>)
</span></span><span style="display:flex;"><span>    cloudwatch <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;cloudwatch&#39;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Get instances and identify optimization opportunities</span>
</span></span><span style="display:flex;"><span>    instances <span style="color:#f92672">=</span> get_all_instances(ec2_client)
</span></span><span style="display:flex;"><span>    low_utilization <span style="color:#f92672">=</span> find_low_utilization_instances(instances, cloudwatch)
</span></span><span style="display:flex;"><span>    potential_downsizing <span style="color:#f92672">=</span> find_downsizing_opportunities(instances, cloudwatch)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;total_instances&#39;</span>: len(instances),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;running_instances&#39;</span>: count_running_instances(instances),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;stopped_instances&#39;</span>: count_stopped_instances(instances),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;low_utilization_instances&#39;</span>: low_utilization,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;potential_downsizing&#39;</span>: potential_downsizing
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Helper functions (implementation details omitted for brevity)</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_all_instances</span>(ec2_client):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">find_low_utilization_instances</span>(instances, cloudwatch):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">find_downsizing_opportunities</span>(instances, cloudwatch):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> []
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">count_running_instances</span>(instances):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#ae81ff">0</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">count_stopped_instances</span>(instances):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#ae81ff">0</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">analyze_rds_instances</span>():
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">analyze_ebs_volumes</span>():
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">analyze_lambda_functions</span>():
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {}
</span></span></code></pre></div><h3 id="3-report-generation-lambda">3. Report Generation Lambda</h3>
<p>This function generates cost optimization reports and sends notifications:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> time
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> datetime <span style="color:#f92672">import</span> datetime
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        cost_data <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;cost_data&#39;</span>, {})
</span></span><span style="display:flex;"><span>        resource_analysis <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;resource_analysis&#39;</span>, {})
</span></span><span style="display:flex;"><span>        destination <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;destination&#39;</span>, <span style="color:#e6db74">&#39;S3&#39;</span>)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Generate detailed markdown report</span>
</span></span><span style="display:flex;"><span>        report_content <span style="color:#f92672">=</span> generate_markdown_report(cost_data, resource_analysis)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Create a unique filename with timestamp</span>
</span></span><span style="display:flex;"><span>        timestamp <span style="color:#f92672">=</span> datetime<span style="color:#f92672">.</span>now()<span style="color:#f92672">.</span>strftime(<span style="color:#e6db74">&#39;%Y-%m-</span><span style="color:#e6db74">%d</span><span style="color:#e6db74">-%H-%M-%S&#39;</span>)
</span></span><span style="display:flex;"><span>        filename <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;cost-optimization-report-</span><span style="color:#e6db74">{</span>timestamp<span style="color:#e6db74">}</span><span style="color:#e6db74">.md&#34;</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        results <span style="color:#f92672">=</span> {}
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Save to S3 if requested</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> destination <span style="color:#f92672">in</span> [<span style="color:#e6db74">&#39;S3&#39;</span>, <span style="color:#e6db74">&#39;BOTH&#39;</span>]:
</span></span><span style="display:flex;"><span>            s3_url <span style="color:#f92672">=</span> save_to_s3(report_content, filename)
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;s3_url&#39;</span>] <span style="color:#f92672">=</span> s3_url
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Send notification if requested</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> destination <span style="color:#f92672">in</span> [<span style="color:#e6db74">&#39;SNS&#39;</span>, <span style="color:#e6db74">&#39;BOTH&#39;</span>]:
</span></span><span style="display:flex;"><span>            summary <span style="color:#f92672">=</span> generate_summary(cost_data, resource_analysis)
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;s3_url&#39;</span> <span style="color:#f92672">in</span> results:
</span></span><span style="display:flex;"><span>                summary <span style="color:#f92672">+=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#ae81ff">\n\n</span><span style="color:#e6db74">Detailed report: </span><span style="color:#e6db74">{</span>results[<span style="color:#e6db74">&#39;s3_url&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>            
</span></span><span style="display:flex;"><span>            send_notification(summary, timestamp)
</span></span><span style="display:flex;"><span>            results[<span style="color:#e6db74">&#39;sns_notification&#39;</span>] <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Sent&#39;</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Store recommendations in DynamoDB for tracking</span>
</span></span><span style="display:flex;"><span>        store_recommendations(cost_data, resource_analysis)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps(results)
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Helper functions (implementation details omitted for brevity)</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">generate_markdown_report</span>(cost_data, resource_analysis):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">&#34;&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">save_to_s3</span>(report_content, filename):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">&#34;&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">generate_summary</span>(cost_data, resource_analysis):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">&#34;&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">send_notification</span>(summary, timestamp):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">pass</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">store_recommendations</span>(cost_data, resource_analysis):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Implementation details omitted</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">pass</span>
</span></span></code></pre></div><h2 id="step-2-set-up-dynamodb-for-tracking-recommendations">Step 2: Set up DynamoDB for tracking recommendations</h2>
<p>You&rsquo;ll need DynamoDB tables to track reports and recommendations. In production, you&rsquo;d define these in your infrastructure-as-code using Terraform or CloudFormation. The tables need:</p>
<ol>
<li>
<p><strong>CostOptimizationRecommendations table</strong>:</p>
<ul>
<li>Partition key: <code>recommendation_id</code> (String)</li>
<li>PAY_PER_REQUEST billing mode for cost efficiency</li>
</ul>
</li>
<li>
<p><strong>CostOptimizationReports table</strong>:</p>
<ul>
<li>Partition key: <code>report_id</code> (String)</li>
<li>PAY_PER_REQUEST billing mode for cost efficiency</li>
</ul>
</li>
</ol>
<p>The first table stores individual recommendations with their implementation status, while the second table tracks metadata about generated reports.</p>
<h2 id="step-3-create-the-bedrock-agent">Step 3: Create the Bedrock Agent</h2>
<p>Now let&rsquo;s create the agent that will orchestrate the entire analysis process:</p>
<ol>
<li>
<p>In the Bedrock console, go to &ldquo;Agents&rdquo; → &ldquo;Create agent&rdquo;</p>
</li>
<li>
<p>Name it &ldquo;CostOptimizationAgent&rdquo;</p>
</li>
<li>
<p>Select Claude 3.5 Sonnet for the foundation model</p>
</li>
<li>
<p>Create three action groups:</p>
<p>a. <strong>GetCostData</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;GetCostData&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Retrieves cost and usage data from AWS Cost Explorer&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: Cost Explorer API\n  version: 1.0.0\npaths:\n  /getCostData:\n    post:\n      summary: Get cost and usage data from AWS Cost Explorer\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              properties:\n                time_period:\n                  type: string\n                  description: The time period to analyze (WEEK, MONTH, QUARTER)\n                services:\n                  type: array\n                  items:\n                    type: string\n                  description: Optional filter for specific AWS services\n      responses:\n        200:\n          description: Successful response with cost data&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-COST-DATA-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>b. <strong>AnalyzeResources</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;AnalyzeResources&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Analyzes AWS resources for optimization opportunities&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: Resource Analysis API\n  version: 1.0.0\npaths:\n  /analyzeResources:\n    post:\n      summary: Analyze AWS resources for optimization opportunities\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              properties:\n                resource_types:\n                  type: array\n                  items:\n                    type: string\n                  description: Resource types to analyze, e.g., &#39;ec2&#39;, &#39;rds&#39;, &#39;ebs&#39;, &#39;lambda&#39;\n      responses:\n        200:\n          description: Successful analysis of resources&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-RESOURCE-ANALYSIS-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>c. <strong>GenerateReport</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;GenerateReport&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Generates a cost optimization report and sends notifications&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: Report Generation API\n  version: 1.0.0\npaths:\n  /generateReport:\n    post:\n      summary: Generate a cost optimization report and send notifications\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              required:\n                - cost_data\n                - resource_analysis\n              properties:\n                cost_data:\n                  type: object\n                  description: Cost and usage data from Cost Explorer\n                resource_analysis:\n                  type: object\n                  description: Results of resource analysis\n                destination:\n                  type: string\n                  description: Where to send the report (S3, SNS, or BOTH)\n      responses:\n        200:\n          description: Successful report generation&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-REPORT-GENERATION-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div></li>
<li>
<p>Configure the agent&rsquo;s instructions:</p>
<pre tabindex="0"><code>You are a helpful cost optimization agent for AWS. Your purpose is to analyze AWS costs and resource usage to identify potential savings opportunities.

When asked to analyze costs:
1. Get cost and usage data using GetCostData action
2. Analyze resources for optimization opportunities using AnalyzeResources action
3. Generate a report of findings and recommendations using GenerateReport action

Your recommendations should be practical and actionable, focusing on:
- EC2 instance optimization (rightsizing, stopping idle instances)
- Unattached or underutilized EBS volumes
- Rarely used Lambda functions
- Reserved Instance or Savings Plans opportunities
- Multi-AZ configurations that might not be needed for non-production

Present findings clearly in order of potential savings, with the highest-impact opportunities first.
</code></pre></li>
</ol>
<p>These instructions are crucial as they define how the agent will behave when analyzing costs. The careful wording ensures it focuses on the most impactful optimization opportunities.</p>
<h2 id="step-4-create-the-main-lambda-function">Step 4: Create the main Lambda function</h2>
<p>Create the main Lambda function that will be triggered by the schedule:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> logging
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> time
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize Bedrock Runtime client</span>
</span></span><span style="display:flex;"><span>bedrock_agent_runtime <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;bedrock-agent-runtime&#39;</span>)
</span></span><span style="display:flex;"><span>logger <span style="color:#f92672">=</span> logging<span style="color:#f92672">.</span>getLogger()
</span></span><span style="display:flex;"><span>logger<span style="color:#f92672">.</span>setLevel(logging<span style="color:#f92672">.</span>INFO)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Get parameters from the event</span>
</span></span><span style="display:flex;"><span>        time_period <span style="color:#f92672">=</span> event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;time_period&#39;</span>, <span style="color:#e6db74">&#39;MONTH&#39;</span>)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Invoke the Bedrock agent</span>
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> bedrock_agent_runtime<span style="color:#f92672">.</span>invoke_agent(
</span></span><span style="display:flex;"><span>            agentId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ID&#39;</span>],
</span></span><span style="display:flex;"><span>            agentAliasId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ALIAS_ID&#39;</span>],
</span></span><span style="display:flex;"><span>            sessionId<span style="color:#f92672">=</span><span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;cost-analysis-</span><span style="color:#e6db74">{</span>int(time<span style="color:#f92672">.</span>time())<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>,
</span></span><span style="display:flex;"><span>            inputText<span style="color:#f92672">=</span><span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Analyze AWS costs for the past </span><span style="color:#e6db74">{</span>time_period<span style="color:#f92672">.</span>lower()<span style="color:#e6db74">}</span><span style="color:#e6db74">, look for optimization opportunities, and generate a comprehensive report with specific actionable recommendations.&#34;</span>
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Process agent response</span>
</span></span><span style="display:flex;"><span>        completion <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;&#39;</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">for</span> event <span style="color:#f92672">in</span> response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;completion&#39;</span>, []):
</span></span><span style="display:flex;"><span>            chunk <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;chunk&#39;</span>][<span style="color:#e6db74">&#39;bytes&#39;</span>]<span style="color:#f92672">.</span>decode())
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">if</span> chunk[<span style="color:#e6db74">&#39;type&#39;</span>] <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;message&#39;</span>:
</span></span><span style="display:flex;"><span>                completion <span style="color:#f92672">+=</span> chunk[<span style="color:#e6db74">&#39;message&#39;</span>][<span style="color:#e6db74">&#39;content&#39;</span>][<span style="color:#ae81ff">0</span>][<span style="color:#e6db74">&#39;text&#39;</span>]
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;Cost analysis completed&#39;</span>})
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        logger<span style="color:#f92672">.</span>error(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Error: </span><span style="color:#e6db74">{</span>str(e)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})
</span></span><span style="display:flex;"><span>        }
</span></span></code></pre></div><h2 id="step-5-set-up-the-eventbridge-rule">Step 5: Set up the EventBridge rule</h2>
<p>Create an EventBridge rule to run the analysis on a schedule. In a production environment, you&rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation, setting parameters like:</p>
<ul>
<li>Rule name: &ldquo;WeeklyCostOptimizationAnalysis&rdquo;</li>
<li>Schedule expression: &ldquo;cron(0 8 ? * MON *)&rdquo; (runs every Monday at 8 AM)</li>
<li>Target: Your main Lambda function</li>
</ul>
<p>EventBridge ensures the cost analysis runs automatically at your chosen interval without manual intervention.</p>
<h2 id="step-6-set-up-sns-for-notifications">Step 6: Set up SNS for notifications</h2>
<p>Set up an SNS topic for notifications by creating a topic and adding subscribers. In production, you&rsquo;d define this in your infrastructure-as-code, configuring:</p>
<ul>
<li>Topic name: &ldquo;CostOptimizationAlerts&rdquo;</li>
<li>Protocol: Email, SMS, or webhook (based on your preferred notification channel)</li>
<li>Subscribers: Finance team, cloud administrators, or a Slack webhook</li>
</ul>
<p>The notification system ensures key stakeholders are informed of optimization opportunities as they&rsquo;re identified.</p>
<h2 id="analysis-capabilities">Analysis Capabilities</h2>
<p>The agent can identify several types of cost optimization opportunities:</p>
<h3 id="1-ec2-instance-optimization">1. EC2 Instance Optimization</h3>
<ul>
<li><strong>Idle Instances</strong>: Identifies running instances with CPU utilization consistently below 5%</li>
<li><strong>Rightsizing Opportunities</strong>: Finds instances that could be downsized based on utilization patterns</li>
<li><strong>Stopped Instances</strong>: Locates instances that have been stopped for extended periods</li>
<li><strong>Instance Family Upgrades</strong>: Suggests moving to newer generation instance families</li>
<li><strong>Reserved Instance Coverage</strong>: Identifies on-demand instances that should be covered by RIs</li>
</ul>
<h3 id="2-storage-optimization">2. Storage Optimization</h3>
<ul>
<li><strong>Unattached EBS Volumes</strong>: Finds volumes not attached to instances</li>
<li><strong>Old Snapshots</strong>: Identifies EBS snapshots older than 6 months</li>
<li><strong>Underutilized Volumes</strong>: Locates volumes with consistently low I/O patterns</li>
<li><strong>Storage Class Transitions</strong>: Recommends moving infrequently accessed data to lower-cost storage tiers</li>
</ul>
<h3 id="3-database-optimization">3. Database Optimization</h3>
<ul>
<li><strong>Overprovisioned RDS Instances</strong>: Identifies oversized database instances</li>
<li><strong>Multi-AZ in Development</strong>: Flags multi-AZ deployments in non-production environments</li>
<li><strong>Idle Databases</strong>: Finds database instances with minimal connection counts</li>
<li><strong>Reserved Instance Opportunities</strong>: Suggests RIs for stable database workloads</li>
</ul>
<h3 id="4-serverless-optimization">4. Serverless Optimization</h3>
<ul>
<li><strong>Overallocated Memory</strong>: Identifies Lambda functions with excessive memory allocation</li>
<li><strong>Rarely Used Functions</strong>: Finds functions that are rarely invoked but consume resources</li>
<li><strong>Long-Running Functions</strong>: Suggests optimizations for functions that consistently run long</li>
</ul>
<h2 id="cost-considerations">Cost considerations</h2>
<p>This solution is cost-efficient:</p>
<ul>
<li>Lambda costs: Most usage will fall under the free tier</li>
<li>EventBridge: No additional cost for scheduled rules</li>
<li>Bedrock API: ~$0.015 per 1,000 tokens with Claude Sonnet</li>
<li>S3: Minimal storage costs for reports</li>
<li>DynamoDB: Pay-per-request pricing keeps costs very low</li>
<li>SNS: Practically free for email notifications</li>
</ul>
<p>For a weekly execution schedule, the infrastructure costs typically remain under $5 per month for most organizations due to the minimal compute resources required.</p>
<h2 id="extending-the-solution">Extending the solution</h2>
<p>Here are some ways to enhance your cost optimization agent:</p>
<p><strong>Multi-account analysis</strong>: Extend to analyze costs across an AWS Organization. This offers a comprehensive view of spending across your entire cloud estate.</p>
<p><strong>Implementation tracking</strong>: Track which recommendations were implemented and their actual savings. This helps quantify the ROI of the optimization agent.</p>
<p><strong>Automated remediation</strong>: Add capability to automatically implement low-risk optimizations like removing unattached EBS volumes. The agent could implement changes automatically during off-hours.</p>
<p><strong>Slack integration</strong>: Send reports directly to Slack channels, enabling team discussions around cost optimization opportunities and tagging responsible teams.</p>
<p><strong>Tagging compliance</strong>: Check for resources without proper cost allocation tags, ensuring your organization maintains visibility into spend by department, team, or project.</p>
<p><strong>Budget alerts integration</strong>: Combine cost optimization with proactive budget alerts, automatically triggering more aggressive analysis when a budget threshold is approaching.</p>
<p><strong>Custom thresholds</strong>: Allow different teams or environments to set custom thresholds for what constitutes underutilization based on their specific workload patterns.</p>
<h2 id="conclusion">Conclusion</h2>
<p>This solution leverages several AWS services to create an intelligent cost optimization system that analyzes cloud spending and provides specific recommendations for reducing costs.</p>
<p>Key advantages of this approach include:</p>
<ol>
<li><strong>Automation</strong>: Regular, scheduled analysis without manual intervention</li>
<li><strong>Actionable insights</strong>: Specific recommendations rather than just data visualization</li>
<li><strong>Comprehensive coverage</strong>: Analysis across multiple resource types (EC2, RDS, Lambda, etc.)</li>
<li><strong>Prioritization</strong>: Recommendations sorted by potential impact</li>
<li><strong>Integration</strong>: Works with existing AWS services and notification systems</li>
</ol>
<p>The architecture combines the data collection capabilities of AWS Cost Explorer with the analytical power of Amazon Bedrock to generate insights similar to those from a cloud financial analyst. By implementing this solution, organizations can transform cost management from a periodic, manual exercise into an ongoing, automated process.</p>
<p>The system is particularly effective at identifying unused resources, rightsizing opportunities, and reservation/Savings Plans recommendations - areas that often yield significant savings when properly optimized.</p>
]]></content:encoded></item><item><title>Building a GitHub PR Reviewer with Bedrock Agents and Action Groups</title><link>https://lukelittle.com/posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/</link><pubDate>Thu, 12 Feb 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/</guid><description>How to build an intelligent PR review agent that automatically analyzes pull requests and provides feedback using AWS Bedrock Agents</description><content:encoded><![CDATA[<p>Code reviews are essential for maintaining code quality, but they can be time-consuming and often repetitive. Developers find themselves commenting on the same issues across multiple pull requests: missing tests, inconsistent naming, inadequate error handling, and numerous other routine concerns. This creates a bottleneck in the development process, as team members wait for their code to be reviewed while reviewers struggle to balance thorough reviews with their own development work.</p>
<p>An AI assistant can address this challenge by analyzing pull requests before human reviewers, catching common issues and allowing the team to focus on more complex aspects of the code review. This approach doesn&rsquo;t replace human judgment but enhances it, ensuring that routine issues are caught consistently while freeing up developer time for deeper analysis.</p>
<p>This post demonstrates how to build an automated PR reviewer using AWS Bedrock Agents that analyzes code changes and provides feedback directly in GitHub.</p>
<h2 id="what-were-building">What we&rsquo;re building</h2>
<p>A PR review agent that:</p>
<ol>
<li>Gets triggered automatically when a new PR is opened or updated</li>
<li>Fetches the PR diff from GitHub</li>
<li>Analyzes the changes using AWS Bedrock</li>
<li>Posts a detailed review comment with suggestions</li>
<li>Tracks review history in DynamoDB</li>
</ol>
<p>The goal isn&rsquo;t to replace human reviewers, but to complement them by identifying common issues, style violations, and potential bugs before human review begins.</p>
<h2 id="real-world-applications">Real-World Applications</h2>
<p>This solution addresses common development challenges across different contexts:</p>
<p><strong>Enterprise Development Teams</strong>: In large organizations with strict coding standards, the agent ensures consistency across hundreds of developers, reducing the burden on senior engineers who often shoulder the bulk of review responsibilities.</p>
<p><strong>Open Source Projects</strong>: Maintainers can use the PR reviewer to handle the initial assessment of community contributions, ensuring they meet project guidelines before dedicating their limited time to review.</p>
<p><strong>Educational Settings</strong>: Computer science programs can deploy the agent to provide students with immediate feedback on their code submissions, helping them learn best practices without requiring instructor intervention for every issue.</p>
<p><strong>Continuous Integration Pipelines</strong>: The PR reviewer can become part of a broader CI/CD strategy, providing code quality assessments alongside traditional test runs and builds.</p>
<p><strong>Onboarding New Team Members</strong>: New developers on a project receive immediate feedback on their work that helps them understand team coding standards more quickly, accelerating their integration into the team.</p>
<h2 id="architecture-overview">Architecture overview</h2>
<p><img src="/posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/github-bedrock-pr-reviewer-architecture_hu_92007e8c5143642d.webp" srcset="/posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/github-bedrock-pr-reviewer-architecture_hu_2a9db8872adee776.webp 750w, /posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/github-bedrock-pr-reviewer-architecture_hu_92007e8c5143642d.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="712" alt="Architecture diagram showing a GitHub PR reviewer built with AWS Bedrock Agents" loading="lazy" decoding="async"></p>
<p>Here&rsquo;s how the system works:</p>
<pre tabindex="0"><code>GitHub Webhook → EventBridge → Lambda → Bedrock Agent with Action Groups → GitHub API → PR Comments
</code></pre><p>The key components:</p>
<ul>
<li><strong>GitHub Webhook</strong>: Triggers on PR events</li>
<li><strong>EventBridge</strong>: Routes events to Lambda</li>
<li><strong>Lambda</strong>: Processes GitHub events and invokes Bedrock Agent</li>
<li><strong>Bedrock Agent</strong>: Coordinates the review process</li>
<li><strong>Action Groups</strong>: Custom Lambda functions that the agent can call</li>
<li><strong>DynamoDB</strong>: Tracks review history and status</li>
</ul>
<p>This event-driven architecture ensures the review process begins automatically whenever a pull request is opened or updated, without requiring any manual intervention.</p>
<h2 id="what-youll-need">What you&rsquo;ll need</h2>
<ul>
<li>AWS Account with Bedrock access</li>
<li>GitHub repository where you want to enable reviews</li>
<li>GitHub Personal Access Token with repo permissions</li>
<li>Basic familiarity with AWS and GitHub</li>
</ul>
<h2 id="step-1-create-the-action-group-lambda-functions">Step 1: Create the Action Group Lambda functions</h2>
<p>First, let&rsquo;s create the Lambda functions our Bedrock Agent will use as action groups. We need three functions:</p>
<h3 id="1-get-pr-diff-lambda">1. Get PR Diff Lambda</h3>
<p>This function fetches the pull request details and diff from GitHub:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> requests
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> github <span style="color:#f92672">import</span> Github
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        repo_name <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;repo_name&#39;</span>]
</span></span><span style="display:flex;"><span>        pr_number <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;pr_number&#39;</span>]
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Initialize GitHub</span>
</span></span><span style="display:flex;"><span>        g <span style="color:#f92672">=</span> Github(os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;GITHUB_TOKEN&#39;</span>])
</span></span><span style="display:flex;"><span>        repo <span style="color:#f92672">=</span> g<span style="color:#f92672">.</span>get_repo(repo_name)
</span></span><span style="display:flex;"><span>        pr <span style="color:#f92672">=</span> repo<span style="color:#f92672">.</span>get_pull(int(pr_number))
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Get PR details and diff</span>
</span></span><span style="display:flex;"><span>        pr_title <span style="color:#f92672">=</span> pr<span style="color:#f92672">.</span>title
</span></span><span style="display:flex;"><span>        pr_description <span style="color:#f92672">=</span> pr<span style="color:#f92672">.</span>body <span style="color:#f92672">or</span> <span style="color:#e6db74">&#34;&#34;</span>
</span></span><span style="display:flex;"><span>        pr_author <span style="color:#f92672">=</span> pr<span style="color:#f92672">.</span>user<span style="color:#f92672">.</span>login
</span></span><span style="display:flex;"><span>        pr_files <span style="color:#f92672">=</span> [f<span style="color:#f92672">.</span>filename <span style="color:#66d9ef">for</span> f <span style="color:#f92672">in</span> pr<span style="color:#f92672">.</span>get_files()]
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        diff_url <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;https://github.com/</span><span style="color:#e6db74">{</span>repo_name<span style="color:#e6db74">}</span><span style="color:#e6db74">/pull/</span><span style="color:#e6db74">{</span>pr_number<span style="color:#e6db74">}</span><span style="color:#e6db74">.diff&#34;</span>
</span></span><span style="display:flex;"><span>        headers <span style="color:#f92672">=</span> {<span style="color:#e6db74">&#39;Authorization&#39;</span>: <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;token </span><span style="color:#e6db74">{</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;GITHUB_TOKEN&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>}
</span></span><span style="display:flex;"><span>        diff_response <span style="color:#f92672">=</span> requests<span style="color:#f92672">.</span>get(diff_url, headers<span style="color:#f92672">=</span>headers)
</span></span><span style="display:flex;"><span>        diff <span style="color:#f92672">=</span> diff_response<span style="color:#f92672">.</span>text
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Return everything the agent needs</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_title&#39;</span>: pr_title,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_description&#39;</span>: pr_description,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_author&#39;</span>: pr_author,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_files&#39;</span>: pr_files,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_diff&#39;</span>: diff
</span></span><span style="display:flex;"><span>            })
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span></code></pre></div><h3 id="2-analyze-code-lambda">2. Analyze Code Lambda</h3>
<p>This function sends the code diff to Bedrock for analysis:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        diff <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;diff&#39;</span>]
</span></span><span style="display:flex;"><span>        languages <span style="color:#f92672">=</span> payload<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;languages&#39;</span>, [])
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Create prompt for the LLM</span>
</span></span><span style="display:flex;"><span>        prompt <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        Please review the following code diff and provide actionable feedback:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        </span><span style="color:#e6db74">{</span>diff<span style="color:#e6db74">}</span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        In your analysis, please look for:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        1. Potential bugs or logic errors
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        2. Security vulnerabilities
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        3. Performance issues
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        4. Code style and best practices
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        5. Missing tests or documentation
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        Format your response as:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        ## Summary
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        (Brief overview of the changes and their purpose)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        ## Critical Issues
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        (List any serious problems that must be fixed)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        ## Suggestions
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        (List minor issues and improvements)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        ## Positive Notes
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        (Highlight good practices in the code)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Call Bedrock for analysis</span>
</span></span><span style="display:flex;"><span>        bedrock <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;bedrock-runtime&#39;</span>)
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> bedrock<span style="color:#f92672">.</span>invoke_model(
</span></span><span style="display:flex;"><span>            modelId<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;anthropic.claude-3-sonnet-20240229-v1:0&#39;</span>,
</span></span><span style="display:flex;"><span>            contentType<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;application/json&#39;</span>,
</span></span><span style="display:flex;"><span>            accept<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;application/json&#39;</span>,
</span></span><span style="display:flex;"><span>            body<span style="color:#f92672">=</span>json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#34;anthropic_version&#34;</span>: <span style="color:#e6db74">&#34;bedrock-2023-05-31&#34;</span>,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#34;max_tokens&#34;</span>: <span style="color:#ae81ff">2000</span>,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#34;messages&#34;</span>: [{<span style="color:#e6db74">&#34;role&#34;</span>: <span style="color:#e6db74">&#34;user&#34;</span>, <span style="color:#e6db74">&#34;content&#34;</span>: prompt}]
</span></span><span style="display:flex;"><span>            })
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Parse and return response</span>
</span></span><span style="display:flex;"><span>        response_body <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(response[<span style="color:#e6db74">&#39;body&#39;</span>]<span style="color:#f92672">.</span>read())
</span></span><span style="display:flex;"><span>        analysis <span style="color:#f92672">=</span> response_body[<span style="color:#e6db74">&#39;content&#39;</span>][<span style="color:#ae81ff">0</span>][<span style="color:#e6db74">&#39;text&#39;</span>]
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;analysis&#39;</span>: analysis})
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span></code></pre></div><h3 id="3-post-comment-lambda">3. Post Comment Lambda</h3>
<p>This function posts the review feedback to the GitHub PR and records the review in DynamoDB:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> time
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> github <span style="color:#f92672">import</span> Github
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract parameters from the event</span>
</span></span><span style="display:flex;"><span>        payload <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        repo_name <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;repo_name&#39;</span>]
</span></span><span style="display:flex;"><span>        pr_number <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;pr_number&#39;</span>]
</span></span><span style="display:flex;"><span>        comment <span style="color:#f92672">=</span> payload[<span style="color:#e6db74">&#39;comment&#39;</span>]
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Record review in DynamoDB</span>
</span></span><span style="display:flex;"><span>        dynamodb <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>resource(<span style="color:#e6db74">&#39;dynamodb&#39;</span>)
</span></span><span style="display:flex;"><span>        table <span style="color:#f92672">=</span> dynamodb<span style="color:#f92672">.</span>Table(os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;REVIEW_TABLE&#39;</span>])
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        review_id <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>repo_name<span style="color:#e6db74">}</span><span style="color:#e6db74">-</span><span style="color:#e6db74">{</span>pr_number<span style="color:#e6db74">}</span><span style="color:#e6db74">-</span><span style="color:#e6db74">{</span>int(time<span style="color:#f92672">.</span>time())<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>        table<span style="color:#f92672">.</span>put_item(
</span></span><span style="display:flex;"><span>            Item<span style="color:#f92672">=</span>{
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;review_id&#39;</span>: review_id,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;repo_name&#39;</span>: repo_name,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;pr_number&#39;</span>: pr_number,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;timestamp&#39;</span>: int(time<span style="color:#f92672">.</span>time()),
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;comment&#39;</span>: comment
</span></span><span style="display:flex;"><span>            }
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Post comment to GitHub</span>
</span></span><span style="display:flex;"><span>        g <span style="color:#f92672">=</span> Github(os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;GITHUB_TOKEN&#39;</span>])
</span></span><span style="display:flex;"><span>        repo <span style="color:#f92672">=</span> g<span style="color:#f92672">.</span>get_repo(repo_name)
</span></span><span style="display:flex;"><span>        pr <span style="color:#f92672">=</span> repo<span style="color:#f92672">.</span>get_pull(int(pr_number))
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Format the comment for GitHub</span>
</span></span><span style="display:flex;"><span>        formatted_comment <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">## AI Code Review 🤖
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74"></span><span style="color:#e6db74">{</span>comment<span style="color:#e6db74">}</span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">---
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">*This review was automatically generated by the PR Review Agent. [Learn more](https://example.com/pr-agent)*
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        pr<span style="color:#f92672">.</span>create_issue_comment(formatted_comment)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;review_id&#39;</span>: review_id,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;Comment posted successfully&#39;</span>
</span></span><span style="display:flex;"><span>            })
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span></code></pre></div><h2 id="step-2-set-up-dynamodb-for-tracking-reviews">Step 2: Set up DynamoDB for tracking reviews</h2>
<p>You&rsquo;ll need a DynamoDB table to track review history. In production, you&rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation. The table needs:</p>
<ul>
<li>Partition key: <code>review_id</code> (String)</li>
<li>PAY_PER_REQUEST billing mode for cost efficiency</li>
</ul>
<p>This table will store metadata about each review, including repository, PR number, timestamp, and the content of the review comment.</p>
<h2 id="step-3-create-the-bedrock-agent">Step 3: Create the Bedrock Agent</h2>
<p>Now let&rsquo;s create the agent that will orchestrate the entire review process:</p>
<ol>
<li>
<p>In the Bedrock console, go to &ldquo;Agents&rdquo; → &ldquo;Create agent&rdquo;</p>
</li>
<li>
<p>Name it &ldquo;PRReviewAgent&rdquo;</p>
</li>
<li>
<p>Select Claude 3.5 Sonnet for the foundation model</p>
</li>
<li>
<p>Create three action groups:</p>
<p>a. <strong>GetPRDiff</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;GetPRDiff&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Retrieves the diff for a GitHub pull request&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: GitHub PR Diff API\n  version: 1.0.0\npaths:\n  /getPRDiff:\n    post:\n      summary: Get the diff for a GitHub pull request\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              required:\n                - repo_name\n                - pr_number\n              properties:\n                repo_name:\n                  type: string\n                  description: The repository in format &#39;owner/repo&#39;\n                pr_number:\n                  type: integer\n                  description: The pull request number\n      responses:\n        200:\n          description: Successful response\n          content:\n            application/json:\n              schema:\n                type: object\n                properties:\n                  pr_title:\n                    type: string\n                  pr_description:\n                    type: string\n                  pr_author:\n                    type: string\n                  pr_files:\n                    type: array\n                    items:\n                      type: string\n                  pr_diff:\n                    type: string&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-GET-PR-DIFF-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>b. <strong>AnalyzeCode</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;AnalyzeCode&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Analyzes code for potential issues and improvements&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: Code Analysis API\n  version: 1.0.0\npaths:\n  /analyzeCode:\n    post:\n      summary: Analyze code diff for issues and suggestions\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              required:\n                - diff\n              properties:\n                diff:\n                  type: string\n                  description: The code diff to analyze\n                languages:\n                  type: array\n                  items:\n                    type: string\n                  description: Programming languages in the diff (optional)\n      responses:\n        200:\n          description: Successful analysis\n          content:\n            application/json:\n              schema:\n                type: object\n                properties:\n                  analysis:\n                    type: string&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-ANALYZE-CODE-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>c. <strong>PostComment</strong></p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupName&#34;</span>: <span style="color:#e6db74">&#34;PostComment&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;description&#34;</span>: <span style="color:#e6db74">&#34;Posts a review comment to a GitHub pull request&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;apiSchema&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;openapi&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;payload&#34;</span>: <span style="color:#e6db74">&#34;openapi: 3.0.0\ninfo:\n  title: GitHub Comment API\n  version: 1.0.0\npaths:\n  /postComment:\n    post:\n      summary: Post a comment to a GitHub pull request\n      requestBody:\n        required: true\n        content:\n          application/json:\n            schema:\n              type: object\n              required:\n                - repo_name\n                - pr_number\n                - comment\n              properties:\n                repo_name:\n                  type: string\n                  description: The repository in format &#39;owner/repo&#39;\n                pr_number:\n                  type: integer\n                  description: The pull request number\n                comment:\n                  type: string\n                  description: The comment text to post\n      responses:\n        200:\n          description: Successful response\n          content:\n            application/json:\n              schema:\n                type: object\n                properties:\n                  review_id:\n                    type: string\n                  status:\n                    type: string&#34;</span>
</span></span><span style="display:flex;"><span>  },
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;actionGroupExecutor&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;lambda&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;lambdaArn&#34;</span>: <span style="color:#e6db74">&#34;[YOUR-POST-COMMENT-LAMBDA-ARN]&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div></li>
<li>
<p>Configure the agent&rsquo;s instructions:</p>
<pre tabindex="0"><code>You are a helpful GitHub Pull Request reviewer designed to analyze code changes and provide constructive feedback. Your goal is to help developers improve their code by identifying issues and suggesting improvements.

When a PR is submitted:
1. Get the PR diff using the GetPRDiff action
2. Analyze the code for issues using the AnalyzeCode action
3. Post a helpful review comment using the PostComment action

Your reviews should be constructive and educational. Focus on:
- Potential bugs or logic errors
- Security vulnerabilities
- Performance issues
- Code style and best practices
- Missing tests or documentation

Use a professional tone and be concise but thorough in your feedback.
</code></pre></li>
</ol>
<p>These instructions are crucial as they define how the agent will behave when reviewing code. The careful wording encourages constructive feedback while maintaining a professional tone.</p>
<h2 id="step-4-create-the-main-lambda-function">Step 4: Create the main Lambda function</h2>
<p>Now create the main Lambda function that will be triggered by GitHub events and coordinate the review process:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> logging
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize Bedrock Runtime client</span>
</span></span><span style="display:flex;"><span>bedrock_agent_runtime <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;bedrock-agent-runtime&#39;</span>)
</span></span><span style="display:flex;"><span>logger <span style="color:#f92672">=</span> logging<span style="color:#f92672">.</span>getLogger()
</span></span><span style="display:flex;"><span>logger<span style="color:#f92672">.</span>setLevel(logging<span style="color:#f92672">.</span>INFO)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Parse the GitHub webhook event</span>
</span></span><span style="display:flex;"><span>        github_event <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>        headers <span style="color:#f92672">=</span> event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;headers&#39;</span>, {})
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Check if it&#39;s a pull request event we care about</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> headers<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;X-GitHub-Event&#39;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;pull_request&#39;</span>:
</span></span><span style="display:flex;"><span>            action <span style="color:#f92672">=</span> github_event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;action&#39;</span>)
</span></span><span style="display:flex;"><span>            
</span></span><span style="display:flex;"><span>            <span style="color:#75715e"># Process only opened or synchronized (updated) PRs</span>
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">if</span> action <span style="color:#f92672">in</span> (<span style="color:#e6db74">&#39;opened&#39;</span>, <span style="color:#e6db74">&#39;synchronize&#39;</span>):
</span></span><span style="display:flex;"><span>                pr <span style="color:#f92672">=</span> github_event[<span style="color:#e6db74">&#39;pull_request&#39;</span>]
</span></span><span style="display:flex;"><span>                repo_name <span style="color:#f92672">=</span> github_event[<span style="color:#e6db74">&#39;repository&#39;</span>][<span style="color:#e6db74">&#39;full_name&#39;</span>]
</span></span><span style="display:flex;"><span>                pr_number <span style="color:#f92672">=</span> pr[<span style="color:#e6db74">&#39;number&#39;</span>]
</span></span><span style="display:flex;"><span>                
</span></span><span style="display:flex;"><span>                <span style="color:#75715e"># Invoke the Bedrock agent to perform the review</span>
</span></span><span style="display:flex;"><span>                response <span style="color:#f92672">=</span> bedrock_agent_runtime<span style="color:#f92672">.</span>invoke_agent(
</span></span><span style="display:flex;"><span>                    agentId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ID&#39;</span>],
</span></span><span style="display:flex;"><span>                    agentAliasId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ALIAS_ID&#39;</span>],
</span></span><span style="display:flex;"><span>                    sessionId<span style="color:#f92672">=</span><span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>repo_name<span style="color:#e6db74">}</span><span style="color:#e6db74">-</span><span style="color:#e6db74">{</span>pr_number<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>,
</span></span><span style="display:flex;"><span>                    inputText<span style="color:#f92672">=</span><span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Review pull request #</span><span style="color:#e6db74">{</span>pr_number<span style="color:#e6db74">}</span><span style="color:#e6db74"> in repository </span><span style="color:#e6db74">{</span>repo_name<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>                )
</span></span><span style="display:flex;"><span>                
</span></span><span style="display:flex;"><span>                <span style="color:#75715e"># Process agent response</span>
</span></span><span style="display:flex;"><span>                completion <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;&#39;</span>
</span></span><span style="display:flex;"><span>                <span style="color:#66d9ef">for</span> event <span style="color:#f92672">in</span> response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;completion&#39;</span>, []):
</span></span><span style="display:flex;"><span>                    chunk <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;chunk&#39;</span>][<span style="color:#e6db74">&#39;bytes&#39;</span>]<span style="color:#f92672">.</span>decode())
</span></span><span style="display:flex;"><span>                    <span style="color:#66d9ef">if</span> chunk[<span style="color:#e6db74">&#39;type&#39;</span>] <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;message&#39;</span>:
</span></span><span style="display:flex;"><span>                        completion <span style="color:#f92672">+=</span> chunk[<span style="color:#e6db74">&#39;message&#39;</span>][<span style="color:#e6db74">&#39;content&#39;</span>][<span style="color:#ae81ff">0</span>][<span style="color:#e6db74">&#39;text&#39;</span>]
</span></span><span style="display:flex;"><span>                
</span></span><span style="display:flex;"><span>                <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;Agent invoked successfully&#39;</span>})
</span></span><span style="display:flex;"><span>                }
</span></span><span style="display:flex;"><span>            
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;Ignored PR action&#39;</span>})}
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;Ignored webhook event&#39;</span>})}
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        logger<span style="color:#f92672">.</span>error(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Error: </span><span style="color:#e6db74">{</span>str(e)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">500</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;error&#39;</span>: str(e)})}
</span></span></code></pre></div><h2 id="step-5-create-the-api-gateway">Step 5: Create the API Gateway</h2>
<ol>
<li>Create a new REST API in API Gateway</li>
<li>Add a POST method for the root resource</li>
<li>Set the integration type to Lambda Function and select your main function</li>
<li>Deploy the API to a stage and note the URL</li>
</ol>
<p>The API Gateway serves as the entry point for GitHub webhooks, receiving events when pull requests are opened or updated.</p>
<h2 id="step-6-set-up-the-github-webhook">Step 6: Set up the GitHub webhook</h2>
<ol>
<li>Go to your GitHub repository</li>
<li>Navigate to Settings → Webhooks → Add webhook</li>
<li>Set the Payload URL to your API Gateway URL</li>
<li>Set Content type to application/json</li>
<li>Select &ldquo;Let me select individual events&rdquo; and check &ldquo;Pull requests&rdquo;</li>
<li>Click &ldquo;Add webhook&rdquo;</li>
</ol>
<p>With the webhook in place, GitHub will now notify your system whenever pull requests are created or updated, triggering the automated review process.</p>
<h2 id="testing-the-system">Testing the system</h2>
<p>To test your PR review agent:</p>
<ol>
<li>Create a simple PR with some code changes</li>
<li>Check that the webhook is triggered (visible in GitHub webhook settings)</li>
<li>Verify your Lambda is invoked (check CloudWatch logs)</li>
<li>Check that a comment appears on the PR with the review</li>
</ol>
<p>You should see a comprehensive review comment that identifies potential issues, makes suggestions for improvement, and highlights positive aspects of the code changes.</p>
<h2 id="enhancing-the-agent">Enhancing the agent</h2>
<p>Here are some ways to improve your PR reviewer:</p>
<p><strong>Language-specific rules</strong>: Extend the AnalyzeCode lambda to apply language-specific linters or static analysis tools for more precise feedback based on the programming language.</p>
<p><strong>Contextual awareness</strong>: Include repository history, architecture documentation, or team standards in the analysis to provide more relevant and contextual suggestions.</p>
<p><strong>Review customization</strong>: Allow teams to set review focus areas through configuration. For example, some teams might prioritize security checks while others emphasize performance.</p>
<p><strong>Learning from feedback</strong>: Track which suggestions developers implement versus ignore to improve future recommendations and reduce false positives.</p>
<p><strong>PR metadata analysis</strong>: Consider PR size, files changed, and complexity when determining the review strategy. Large PRs might receive different handling than small, focused changes.</p>
<p><strong>Inline comments</strong>: Enhance the agent to post specific comments on individual lines in the diff rather than just a summary comment.</p>
<p><strong>Pre-commit integration</strong>: Offer developers the ability to run the same analysis locally before submitting their PR, using pre-commit hooks.</p>
<h2 id="cost-considerations">Cost considerations</h2>
<p>This architecture is cost-efficient for most development teams:</p>
<ul>
<li>Lambda costs: Typically covered by the free tier for small to medium teams</li>
<li>API Gateway: Approximately $1 per million requests</li>
<li>Bedrock API: Around $0.015 per 1,000 tokens with Claude Sonnet</li>
<li>DynamoDB: Pay-per-request pricing with minimal storage needs</li>
</ul>
<p>The total cost will vary based on your team&rsquo;s PR volume and the size of code changes being reviewed, but for most teams, it remains quite affordable compared to the developer time saved.</p>
<h2 id="security-considerations">Security considerations</h2>
<p>When implementing this:</p>
<ul>
<li>Store API tokens in AWS Secrets Manager</li>
<li>Set IAM permissions using least privilege</li>
<li>Consider the sensitivity of code being reviewed</li>
<li>Remember that agents aren&rsquo;t perfect at security analysis</li>
</ul>
<p>While the PR reviewer can identify many common security issues, it should not be your only security control. Critical security reviews should still be performed by security experts, especially for sensitive components.</p>
<h2 id="conclusion">Conclusion</h2>
<p>By combining AWS Bedrock Agents with Action Groups, this architecture creates an intelligent PR review system that helps maintain code quality while reducing the time developers spend on repetitive aspects of code reviews.</p>
<p>The solution provides several key benefits:</p>
<ol>
<li><strong>Consistency</strong>: Every PR receives the same baseline level of review</li>
<li><strong>Early detection</strong>: Catches common issues before human reviewers see the code</li>
<li><strong>Educational feedback</strong>: Provides context and explanations for suggested improvements</li>
<li><strong>Focus</strong>: Allows human reviewers to concentrate on architecture and business logic</li>
<li><strong>Historical tracking</strong>: Stores reviews for future reference and improvement</li>
</ol>
<p>While this implementation uses specific AWS services, the architectural patterns can be adapted to other cloud providers or integrated with different version control systems beyond GitHub.</p>
<p>This approach to automated code review represents a practical application of generative AI that delivers immediate value to development teams: better code quality, faster reviews, and more time for developers to focus on creative and complex challenges rather than routine feedback.</p>
]]></content:encoded></item><item><title>Building a Company Knowledge Bot: Slack + Bedrock Knowledge Bases</title><link>https://lukelittle.com/posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/</link><pubDate>Wed, 11 Feb 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/</guid><description>How to build a company documentation Q&amp;amp;A bot using AWS Bedrock Knowledge Bases and Slack - giving your team instant answers from your docs</description><content:encoded><![CDATA[<p>&ldquo;Where can I find our vacation policy?&rdquo; &ldquo;What&rsquo;s the process for requesting new hardware?&rdquo; &ldquo;Can you explain our security guidelines?&rdquo; These questions echo through company Slack channels daily, interrupting workflows and creating redundant work for team leads and HR staff. The same questions get asked repeatedly, and answers are buried in documentation that&rsquo;s difficult to navigate.</p>
<p>In this post, I&rsquo;ll show you how to build a simple yet powerful Q&amp;A bot for Slack that leverages your company&rsquo;s documentation to provide accurate, contextual answers. The best part? It runs entirely on AWS managed services, minimizing operational overhead while delivering immediate value to your organization.</p>
<h2 id="real-world-applications">Real-World Applications</h2>
<p>This solution addresses documentation challenges across different departments:</p>
<p><strong>HR and People Teams</strong>: Employees constantly ask about benefits, PTO policies, and workplace guidelines. An AI bot can instantly answer &ldquo;How many vacation days do I have?&rdquo; or &ldquo;What&rsquo;s our parental leave policy?&rdquo; by citing the exact paragraph from your handbook.</p>
<p><strong>Engineering Teams</strong>: Technical documentation grows exponentially with your codebase. When an engineer asks &ldquo;How do I set up the development environment?&rdquo; or &ldquo;What&rsquo;s our database migration process?&rdquo;, the bot can provide step-by-step instructions from your wiki.</p>
<p><strong>Product and Sales Teams</strong>: Sales representatives need quick access to product specifications, pricing details, and competitive positioning. A knowledge bot can answer &ldquo;What are the enterprise tier limits?&rdquo; during a client call without disrupting other team members.</p>
<p><strong>Customer Support</strong>: Support teams juggle hundreds of internal processes. When an agent needs to know &ldquo;What&rsquo;s our escalation policy?&rdquo; or &ldquo;How do I process a refund?&rdquo;, immediate answers improve customer response times.</p>
<p><strong>New Employee Onboarding</strong>: The first weeks at a new job involve absorbing massive amounts of information. A knowledge bot gives new hires an accessible way to ask questions without feeling like they&rsquo;re bothering colleagues.</p>
<h2 id="what-were-building">What We&rsquo;re Building</h2>
<p>A Slack bot that:</p>
<ol>
<li>Receives questions from users in a channel or DM</li>
<li>Uses AWS Bedrock Knowledge Base to search through your company documentation</li>
<li>Generates accurate answers with citations to source documents</li>
<li>Handles follow-up questions with conversation history</li>
</ol>
<p>The system uses Retrieval-Augmented Generation (RAG), combining the reasoning capabilities of large language models with retrieval from your own data sources—giving you the benefits of generative AI while keeping your data within your AWS account.</p>
<h2 id="understanding-bedrock-knowledge-bases">Understanding Bedrock Knowledge Bases</h2>
<p>AWS Bedrock Knowledge Bases represents a significant advancement in enterprise knowledge management. Let&rsquo;s explore how it works and why it&rsquo;s superior to traditional search or direct LLM prompting.</p>
<h3 id="the-rag-architecture">The RAG Architecture</h3>
<p>Retrieval-Augmented Generation (RAG) addresses a fundamental limitation of LLMs: they have no knowledge of your internal documents. RAG works by:</p>
<ol>
<li><strong>Document Processing</strong>: Your documents are divided into chunks of an appropriate size for retrieval</li>
<li><strong>Vector Embedding</strong>: Each chunk is converted into a numerical vector representation using an embedding model</li>
<li><strong>Vector Storage</strong>: These embeddings are stored in a vector database (OpenSearch Serverless in Bedrock&rsquo;s case)</li>
<li><strong>Semantic Search</strong>: When a question arrives, it&rsquo;s converted to the same vector space and semantically similar chunks are retrieved</li>
<li><strong>Context Augmentation</strong>: Retrieved chunks are injected as context into the prompt sent to the LLM</li>
<li><strong>Answer Generation</strong>: The LLM generates an answer based on this context, citing the relevant sources</li>
</ol>
<p>This approach dramatically improves accuracy by giving the model direct access to your internal knowledge, while maintaining the reasoning capabilities of foundation models.</p>
<h3 id="supported-document-types">Supported Document Types</h3>
<p>Bedrock Knowledge Bases supports a wide range of document formats:</p>
<ul>
<li>PDF files (text and scanned documents with OCR)</li>
<li>Microsoft Office (Word, PowerPoint, Excel)</li>
<li>Text and Markdown files</li>
<li>HTML and web pages</li>
<li>CSV and JSON data</li>
</ul>
<p>This versatility means you can ingest existing documentation without reformatting.</p>
<h3 id="synchronization-and-updates">Synchronization and Updates</h3>
<p>A key feature is automatic synchronization. When documents in your S3 bucket are updated, Bedrock Knowledge Bases can automatically detect these changes and update the vector store, ensuring your bot always has the latest information.</p>
<h3 id="semantic-vs-keyword-search">Semantic vs. Keyword Search</h3>
<p>Traditional search systems match keywords, but Bedrock Knowledge Bases understands concepts. If someone asks about &ldquo;time off,&rdquo; it can retrieve documents about &ldquo;vacation,&rdquo; &ldquo;PTO,&rdquo; and &ldquo;leave of absence&rdquo; because it understands these concepts are related—even if they don&rsquo;t share exact keywords.</p>
<h2 id="architecture-overview">Architecture Overview</h2>
<p><img src="/posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/slack-bedrock-knowledge-base-architecture_hu_dd9302381e1c01e.webp" srcset="/posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/slack-bedrock-knowledge-base-architecture_hu_e66429157f3fba3d.webp 750w, /posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/slack-bedrock-knowledge-base-architecture_hu_dd9302381e1c01e.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="783" alt="Architecture diagram showing a Slack bot connected to AWS Bedrock Knowledge Bases" loading="lazy" decoding="async"></p>
<p>Here&rsquo;s how the solution components work together:</p>
<p>When a user asks a question in Slack, the message triggers a webhook to API Gateway. This request is processed by a Lambda function that maintains conversation context in DynamoDB and communicates with Bedrock. The Bedrock Agent uses the Knowledge Base to search your documentation, retrieves relevant information, and formulates a response that&rsquo;s sent back to the user through Slack.</p>
<p>Each component serves a specific purpose in this flow:</p>
<ul>
<li><strong>Slack API</strong> handles the user interface, making the experience seamless within your existing communication platform.</li>
<li><strong>API Gateway</strong> provides a secure endpoint for Slack to send events to.</li>
<li><strong>Lambda</strong> orchestrates the process, maintaining conversation history and managing the interaction between Slack and Bedrock.</li>
<li><strong>Bedrock Agent</strong> uses Claude to interpret questions, retrieve information, and generate natural-sounding responses.</li>
<li><strong>Bedrock Knowledge Base</strong> indexes and searches your company documentation, finding the most relevant information for each question.</li>
<li><strong>DynamoDB</strong> stores conversation history so the bot can understand follow-up questions in context.</li>
</ul>
<p>This serverless architecture scales automatically with usage and requires minimal maintenance once deployed.</p>
<h2 id="prerequisites">Prerequisites</h2>
<p>Before starting, make sure you have:</p>
<ul>
<li>AWS Account with Bedrock access (you&rsquo;ll need quota for Claude models)</li>
<li>Slack workspace with permissions to create apps</li>
<li>Company documentation organized in a folder structure</li>
<li>Basic familiarity with AWS services and Terraform (or CloudFormation)</li>
</ul>
<h2 id="step-1-create-the-knowledge-base">Step 1: Create the Knowledge Base</h2>
<p>First, we&rsquo;ll create a Knowledge Base to store and index your company documentation:</p>
<ol>
<li>
<p>Upload documents to an S3 bucket:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span>aws s3 mb s3://your-company-docs
</span></span><span style="display:flex;"><span>aws s3 sync ./docs s3://your-company-docs/
</span></span></code></pre></div><p>Organize your documents logically—folders like HR, Engineering, and Sales help the system understand document context.</p>
</li>
<li>
<p>Create the Knowledge Base in Bedrock:</p>
<p>Navigate to AWS Bedrock in the console, select &ldquo;Knowledge bases&rdquo; → &ldquo;Create knowledge base,&rdquo; and follow the wizard:</p>
<ul>
<li>Name: &ldquo;CompanyDocs&rdquo;</li>
<li>Data source: Select your S3 bucket</li>
<li>Vector store: &ldquo;Create new Amazon OpenSearch Serverless vector store&rdquo;</li>
<li>Embedding model: &ldquo;Titan Embeddings G1&rdquo; (offers excellent performance for most use cases)</li>
<li>Enable automatic synchronization to keep your knowledge base updated</li>
</ul>
</li>
</ol>
<p>The initial data synchronization process will take several minutes depending on the volume of your documents. During this time, Bedrock is analyzing your documents, chunking them appropriately, and converting them into vector embeddings for semantic search.</p>
<h2 id="step-2-create-the-bedrock-agent">Step 2: Create the Bedrock Agent</h2>
<p>Now let&rsquo;s create an agent that will use our Knowledge Base:</p>
<ol>
<li>
<p>In the Bedrock console, go to &ldquo;Agents&rdquo; → &ldquo;Create agent&rdquo;</p>
</li>
<li>
<p>Name it &ldquo;CompanyDocsAssistant&rdquo; and select Claude 3.5 Sonnet for the foundation model</p>
</li>
<li>
<p>In the &ldquo;Action groups&rdquo; section, add a Knowledge Base action group and select the &ldquo;CompanyDocs&rdquo; knowledge base we created</p>
</li>
<li>
<p>Configure the agent&rsquo;s instructions with detailed guidance:</p>
<pre tabindex="0"><code>You are a helpful assistant that answers questions about company documentation, policies, and procedures.

When answering:
1. Be concise but thorough
2. Always cite sources by document name when you provide information
3. If you don&#39;t know or can&#39;t find relevant information, say so clearly
4. For follow-up questions, maintain context from previous exchanges
5. Format responses with appropriate Slack formatting (bullets, bold, etc.) where helpful
6. Present step-by-step procedures in numbered lists when applicable
</code></pre></li>
<li>
<p>For the IAM role, create a new service role with the necessary permissions to access your Knowledge Base</p>
</li>
</ol>
<p>The detailed instructions are crucial—they set the tone and behavior of your assistant, determining how it will respond to various types of questions.</p>
<h2 id="step-3-create-the-lambda-function">Step 3: Create the Lambda Function</h2>
<p>Next, create a Lambda function to handle Slack events and communicate with our Bedrock agent:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> logging
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> urllib.request
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> time
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> boto3.dynamodb.conditions <span style="color:#f92672">import</span> Key
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize clients</span>
</span></span><span style="display:flex;"><span>bedrock_agent_runtime <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;bedrock-agent-runtime&#39;</span>)
</span></span><span style="display:flex;"><span>dynamodb <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>resource(<span style="color:#e6db74">&#39;dynamodb&#39;</span>)
</span></span><span style="display:flex;"><span>conversation_table <span style="color:#f92672">=</span> dynamodb<span style="color:#f92672">.</span>Table(os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;CONVERSATION_TABLE&#39;</span>])
</span></span><span style="display:flex;"><span>logger <span style="color:#f92672">=</span> logging<span style="color:#f92672">.</span>getLogger()
</span></span><span style="display:flex;"><span>logger<span style="color:#f92672">.</span>setLevel(logging<span style="color:#f92672">.</span>INFO)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Parse the incoming event from Slack</span>
</span></span><span style="display:flex;"><span>    body <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event[<span style="color:#e6db74">&#39;body&#39;</span>])
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Handle URL verification challenge</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;type&#39;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;url_verification&#39;</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;challenge&#39;</span>: body[<span style="color:#e6db74">&#39;challenge&#39;</span>]})}
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Process message events (app_mention or direct message)</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;event&#39;</span>, {})<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;type&#39;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;app_mention&#39;</span> <span style="color:#f92672">or</span> \
</span></span><span style="display:flex;"><span>       (body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;event&#39;</span>, {})<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;type&#39;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;message&#39;</span> <span style="color:#f92672">and</span> 
</span></span><span style="display:flex;"><span>        body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;event&#39;</span>, {})<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;channel_type&#39;</span>) <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;im&#39;</span>):
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        event_data <span style="color:#f92672">=</span> body[<span style="color:#e6db74">&#39;event&#39;</span>]
</span></span><span style="display:flex;"><span>        user_id <span style="color:#f92672">=</span> event_data[<span style="color:#e6db74">&#39;user&#39;</span>]
</span></span><span style="display:flex;"><span>        channel_id <span style="color:#f92672">=</span> event_data[<span style="color:#e6db74">&#39;channel&#39;</span>]
</span></span><span style="display:flex;"><span>        text <span style="color:#f92672">=</span> event_data<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;text&#39;</span>, <span style="color:#e6db74">&#39;&#39;</span>)<span style="color:#f92672">.</span>replace(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;&lt;@</span><span style="color:#e6db74">{</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BOT_USER_ID&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&gt;&#34;</span>, <span style="color:#e6db74">&#39;&#39;</span>)<span style="color:#f92672">.</span>strip()
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Get conversation history and invoke Bedrock agent</span>
</span></span><span style="display:flex;"><span>        conversation_id <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>user_id<span style="color:#e6db74">}</span><span style="color:#e6db74">:</span><span style="color:#e6db74">{</span>channel_id<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>        history <span style="color:#f92672">=</span> get_conversation_history(conversation_id)
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> invoke_bedrock_agent(text, history, conversation_id)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Send response back to Slack</span>
</span></span><span style="display:flex;"><span>        send_slack_message(channel_id, response)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;ok&#39;</span>})}
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;status&#39;</span>: <span style="color:#e6db74">&#39;ignored&#39;</span>})}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">invoke_bedrock_agent</span>(question, history, conversation_id):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Format history for Bedrock and add the current question</span>
</span></span><span style="display:flex;"><span>        messages <span style="color:#f92672">=</span> format_conversation_history(history)
</span></span><span style="display:flex;"><span>        messages<span style="color:#f92672">.</span>append({
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;role&#39;</span>: <span style="color:#e6db74">&#39;user&#39;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;content&#39;</span>: [{<span style="color:#e6db74">&#39;text&#39;</span>: question}]
</span></span><span style="display:flex;"><span>        })
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Invoke the Bedrock agent</span>
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> bedrock_agent_runtime<span style="color:#f92672">.</span>invoke_agent(
</span></span><span style="display:flex;"><span>            agentId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ID&#39;</span>],
</span></span><span style="display:flex;"><span>            agentAliasId<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>environ[<span style="color:#e6db74">&#39;BEDROCK_AGENT_ALIAS_ID&#39;</span>],
</span></span><span style="display:flex;"><span>            sessionId<span style="color:#f92672">=</span>conversation_id,
</span></span><span style="display:flex;"><span>            inputText<span style="color:#f92672">=</span>question,
</span></span><span style="display:flex;"><span>            enableTrace<span style="color:#f92672">=</span><span style="color:#66d9ef">True</span>
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Extract and process the response</span>
</span></span><span style="display:flex;"><span>        completion <span style="color:#f92672">=</span> process_agent_response(response)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Store conversation in DynamoDB for history</span>
</span></span><span style="display:flex;"><span>        store_conversation_entry(conversation_id, <span style="color:#e6db74">&#39;user&#39;</span>, question)
</span></span><span style="display:flex;"><span>        store_conversation_entry(conversation_id, <span style="color:#e6db74">&#39;assistant&#39;</span>, completion)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> completion
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        logger<span style="color:#f92672">.</span>error(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Error invoking Bedrock agent: </span><span style="color:#e6db74">{</span>str(e)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;I&#39;m having trouble answering that right now. Technical details: </span><span style="color:#e6db74">{</span>str(e)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Additional helper functions for conversation history, messaging, etc.</span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># (implementation details omitted for brevity)</span>
</span></span></code></pre></div><p>This Lambda function handles:</p>
<ul>
<li>Receiving events from Slack</li>
<li>Maintaining conversation context</li>
<li>Communicating with the Bedrock agent</li>
<li>Sending responses back to the user</li>
</ul>
<p>The actual implementation includes additional helpers for conversation history management, message formatting, and error handling that we&rsquo;ve omitted here for brevity.</p>
<h2 id="step-4-set-up-the-dynamodb-table">Step 4: Set up the DynamoDB Table</h2>
<p>You&rsquo;ll need a DynamoDB table to track conversation history. In production, you&rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation. The table needs:</p>
<ul>
<li>Partition key: <code>conversation_id</code> (String)</li>
<li>Sort key: <code>timestamp</code> (Number)</li>
<li>PAY_PER_REQUEST billing mode for cost efficiency</li>
</ul>
<p>This table enables the bot to understand follow-up questions by maintaining context from previous exchanges.</p>
<h2 id="step-5-create-the-slack-app">Step 5: Create the Slack App</h2>
<ol>
<li>Go to <a href="https://api.slack.com/apps">api.slack.com/apps</a> and create a new app &ldquo;From scratch&rdquo;</li>
<li>Under &ldquo;OAuth &amp; Permissions,&rdquo; add these scopes:
<ul>
<li><code>app_mentions:read</code></li>
<li><code>chat:write</code></li>
<li><code>im:history</code></li>
<li><code>im:read</code></li>
</ul>
</li>
<li>Under &ldquo;Event Subscriptions&rdquo;:
<ul>
<li>Enable Events and set the Request URL to your API Gateway endpoint</li>
<li>Subscribe to bot events: <code>app_mention</code> and <code>message.im</code></li>
</ul>
</li>
<li>Install the app to your workspace and copy the Bot User OAuth Token</li>
</ol>
<p>The Slack app configuration establishes the permissions and event subscriptions needed for the bot to receive messages and respond to users.</p>
<h2 id="step-6-deploy-the-api-gateway">Step 6: Deploy the API Gateway</h2>
<p>Create an API Gateway to receive events from Slack:</p>
<ol>
<li>Create a new HTTP API with a POST route that integrates with your Lambda function</li>
<li>Deploy the API and note the URL</li>
<li>Update your Slack app&rsquo;s Event Subscriptions URL with this endpoint</li>
<li>Add environment variables to your Lambda function:
<ul>
<li><code>SLACK_BOT_TOKEN</code>: The OAuth token from your Slack app</li>
<li><code>BOT_USER_ID</code>: The user ID of your Slack bot</li>
<li><code>BEDROCK_AGENT_ID</code>: The ID of your Bedrock agent</li>
<li><code>BEDROCK_AGENT_ALIAS_ID</code>: The alias ID of your agent</li>
<li><code>CONVERSATION_TABLE</code>: Your DynamoDB table name</li>
</ul>
</li>
</ol>
<h2 id="cost-optimization">Cost Optimization</h2>
<p>This solution is cost-effective, but there are a few considerations:</p>
<ul>
<li><strong>Bedrock API calls</strong>: ~$0.015 per 1,000 tokens with Claude Sonnet</li>
<li><strong>Knowledge Base storage</strong>: ~$0.023/GB for S3 + OpenSearch Serverless vector storage</li>
<li><strong>Lambda</strong>: Free tier likely covers most usage patterns</li>
<li><strong>DynamoDB</strong>: Pay-per-request pricing keeps costs low</li>
<li><strong>API Gateway</strong>: ~$1 per million requests</li>
</ul>
<p>For a team of 20 people asking 10 questions per day, expect costs around $30-50 per month. You can implement usage tracking to monitor and control costs as adoption grows.</p>
<h2 id="extending-the-solution">Extending the Solution</h2>
<p>Here are some ways to enhance this basic implementation:</p>
<ul>
<li><strong>Multi-channel support</strong>: Monitor multiple Slack channels with channel-specific knowledge bases</li>
<li><strong>Document syncing</strong>: Set up automatic synchronization with your documentation systems</li>
<li><strong>Permissions</strong>: Implement access controls based on Slack user groups</li>
<li><strong>Analytics</strong>: Track common questions to identify gaps in your documentation</li>
<li><strong>Multi-model support</strong>: Use a simpler model for basic questions and Claude for complex ones</li>
<li><strong>Conversation summarization</strong>: Periodically summarize long conversations for better context management</li>
</ul>
<h2 id="conclusion">Conclusion</h2>
<p>With just a few AWS services, you&rsquo;ve built an intelligent assistant that makes your company&rsquo;s documentation accessible via Slack. No more hunting through SharePoint or Confluence—just ask the bot and get instant answers with citations to the source material.</p>
<p>The real power here is that your data remains within your AWS account, the system only has access to approved documents, and it continuously improves as you add more documentation. As AWS enhances Bedrock&rsquo;s capabilities, your bot will automatically benefit from these improvements without any changes to your architecture.</p>
<p>This solution demonstrates how easily companies can now deploy practical AI applications using managed services. What used to require a specialized ML team and months of development can now be built in days using serverless components.</p>
<p>What documentation would you connect to your knowledge bot first? Let me know <a href="https://www.linkedin.com/in/lucaslittle/">on LinkedIn</a>!</p>
]]></content:encoded></item><item><title>From Prompt to Production: Designing Safe Generative AI on AWS for Regulated Environments</title><link>https://lukelittle.com/posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/</link><pubDate>Sun, 01 Feb 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/</guid><description>How to architect production-grade generative AI systems that meet enterprise security, compliance, and governance requirements with AWS Bedrock</description><content:encoded><![CDATA[<h2 id="the-real-problem-production-not-prototypes">The Real Problem: Production, Not Prototypes</h2>
<p>Everyone can demo generative AI. Almost no one can run it safely in production.</p>
<p>Enterprises in finance, healthcare, and the public sector aren&rsquo;t blocked by technology capabilities—they&rsquo;re blocked by governance requirements that today&rsquo;s AI implementations rarely satisfy.</p>
<p>These organizations face three critical blockers:</p>
<ul>
<li><strong>Data leakage risk</strong>: Sensitive information, from PII to trade secrets, flowing through public model APIs</li>
<li><strong>Lack of auditability</strong>: No reliable record of prompts, responses, or who accessed what information</li>
<li><strong>Unclear ownership</strong>: Ambiguous rights over prompt engineering IP, training data, and generated outputs</li>
</ul>
<p>AWS customers don&rsquo;t want AI that behaves like a chatbot toy. They need AI that behaves like enterprise infrastructure: secured, monitored, audited, governed, and compliant with their existing security posture.</p>
<h2 id="design-goals-for-enterprise-ready-genai">Design Goals for Enterprise-Ready GenAI</h2>
<p>When designing generative AI systems for regulated environments, your architecture must satisfy these non-negotiable requirements:</p>
<ul>
<li>No public internet exposure for sensitive data</li>
<li>No training on customer data without explicit permission</li>
<li>Full audit trail of all prompts and responses</li>
<li>IAM-first access control integrated with enterprise identity</li>
<li>Serverless and scalable by default</li>
</ul>
<p>This checklist maps directly to AWS Well-Architected Framework principles, particularly in security and operational excellence.</p>
<h2 id="reference-architecture-overview">Reference Architecture Overview</h2>
<p>Here&rsquo;s a reference architecture that meets these requirements using AWS services:</p>
<p><img src="/posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/genai-reference-architecture_hu_4c74e40b746f2990.webp" srcset="/posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/genai-reference-architecture_hu_cee2e7aacadf3e55.webp 750w, /posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/genai-reference-architecture_hu_4c74e40b746f2990.webp 1301w" sizes="(max-width: 800px) 100vw, 750px"
       width="1301" height="960" alt="Reference Architecture for Safe Generative AI on AWS" loading="lazy" decoding="async"></p>
<h2 id="walking-the-architecture-building-for-security-and-scale">Walking the Architecture: Building for Security and Scale</h2>
<h3 id="1-edge--entry-cloudfront--api-gateway">1️⃣ Edge &amp; Entry: CloudFront + API Gateway</h3>
<p>The edge layer serves as your first line of defense:</p>
<ul>
<li>Global edge protection through CloudFront</li>
<li>Request validation and throttling via API Gateway</li>
<li>Clear API contract for AI access</li>
<li>WAF rules to block suspicious patterns</li>
</ul>
<p>This approach frames AI as just another AWS workload, not an exception to your security rules. Your existing infrastructure and compliance controls extend naturally to your AI services.</p>
<h3 id="2-prompt-handling-lambda">2️⃣ Prompt Handling: Lambda</h3>
<p>The Prompt Handler Lambda is where policy meets AI:</p>
<ul>
<li>Sanitizes inputs to prevent prompt injection</li>
<li>Injects system prompts to enforce guardrails</li>
<li>Enforces token limits (cost control)</li>
<li>Attaches request metadata (user ID, application, purpose)</li>
</ul>
<p>This layer ensures all model interactions are appropriately structured and traced. Every prompt includes context about who sent it, why, and what constraints apply.</p>
<h3 id="3-private-model-access-bedrock-via-vpc-endpoint">3️⃣ Private Model Access: Bedrock via VPC Endpoint</h3>
<p>The model interaction layer guarantees data privacy:</p>
<ul>
<li>No public internet egress</li>
<li>No customer-managed model hosting</li>
<li>No fine-tuning on customer prompts</li>
<li>VPC integration with existing security controls</li>
</ul>
<p>The model is consumed like a managed AWS service—not an external API. This distinction is critical for security teams evaluating AI adoption.</p>
<h3 id="4-response-filtering-post-processing-lambda">4️⃣ Response Filtering: Post-processing Lambda</h3>
<p>The response handler implements safety guardrails:</p>
<ul>
<li>Content moderation (PII, offensive content)</li>
<li>Output validation against schema</li>
<li>Optional redaction of sensitive information</li>
<li>Confidence scoring and hallucination detection</li>
</ul>
<p>This layer acknowledges and mitigates hallucination risk without fear-mongering, providing mechanisms to validate and filter model outputs.</p>
<h3 id="5-audit--evidence-dynamodb--s3">5️⃣ Audit &amp; Evidence: DynamoDB / S3</h3>
<p>The audit layer addresses compliance requirements:</p>
<ul>
<li>Persistent storage of prompt hashes</li>
<li>Model ID and version tracking</li>
<li>Timestamped responses</li>
<li>Immutable audit logs</li>
</ul>
<p>This creates a defensible evidentiary trail that satisfies governance requirements for regulated industries.</p>
<h2 id="why-this-works-for-regulated-industries">Why This Works for Regulated Industries</h2>
<p>This architecture succeeds where most AI implementations fail because it addresses the key requirements that matter to enterprise stakeholders:</p>
<p><strong>Security</strong>: IAM-based access control, VPC endpoints, and private networking eliminate public exposure risks. The system operates entirely within your security perimeter, following the principle of &ldquo;default deny&rdquo; with explicit allow policies.</p>
<p><strong>Compliance</strong>: Complete prompt/response traceability enables regulatory reporting and satisfies audit requirements. You can demonstrate who used the system, when, how, and what results they received—critical for SOC2, HIPAA, and FedRAMP.</p>
<p><strong>Cost control</strong>: Serverless scaling plus token limits provide predictable, manageable costs. Unlike self-hosted options, you&rsquo;re not paying for idle infrastructure, and unlike public APIs, you have fine-grained control over usage patterns.</p>
<p><strong>Operational clarity</strong>: The system is observable, debuggable, and auditable using the same tools you already use for the rest of your AWS infrastructure. There&rsquo;s no AI-specific monitoring to implement.</p>
<p>This approach works particularly well for financial services (handling sensitive financial data), healthcare (maintaining PHI compliance), and public sector (satisfying FedRAMP requirements)—precisely the industries with the most to gain from AI and the most stringent security requirements.</p>
<h2 id="implementation-considerations">Implementation Considerations</h2>
<p>When implementing this pattern in production, several practical considerations emerge:</p>
<p><strong>IAM roles and boundaries</strong>: Create specific IAM roles for each component with least-privilege access. The prompt handler needs Bedrock access but not S3 write access; the response filter needs DynamoDB write access but not Bedrock APIs. Use service control policies (SCPs) to enforce guardrails.</p>
<p><strong>VPC design</strong>: Depending on your existing network topology, you may need to adjust the VPC design. For large enterprises with transit gateways, consider routing AI traffic through dedicated VPCs with specific security monitoring.</p>
<p><strong>Cost management</strong>: Monitor token usage carefully. Implement token quotas at the API Gateway layer and consider using smaller context window models for initial responses, reserving larger context models for specific use cases.</p>
<p><strong>Scaling characteristics</strong>: Lambda&rsquo;s concurrency model handles traffic spikes well, but Bedrock has model-specific quotas and SLAs. Request quota increases proactively if you anticipate high volume. Consider implementing queue-based architectures for asynchronous workloads.</p>
<p><strong>Cross-account patterns</strong>: For large organizations, implement a hub-and-spoke model where a central AI governance account hosts the Bedrock endpoint, with workload accounts accessing it through cross-account roles. This centralizes auditing while enabling distributed usage.</p>
<h2 id="conclusion">Conclusion</h2>
<p>The gap between AI demos and AI in production isn&rsquo;t primarily a technical gap—it&rsquo;s a governance gap.</p>
<p>This reference architecture bridges that gap by treating generative AI as enterprise infrastructure rather than a standalone tool. It integrates with existing security controls, creates auditability, and provides the governance hooks necessary for regulated environments.</p>
<p>The result? AI that can safely navigate the journey from prompt to production, enabling organizations to capture AI&rsquo;s business value without compromising on security and compliance requirements.</p>
<p>In regulated environments, the future of AI isn&rsquo;t about building fancy demos—it&rsquo;s about building trust. By architecting generative AI systems that behave like proper enterprise infrastructure—secured, monitored, audited, and governed—we allow organizations to focus on business value rather than security firefighting.</p>
<p>Remember that this is a reference architecture, not a one-size-fits-all solution. Your specific implementation should be tailored to your compliance requirements, existing infrastructure, and risk profile. But the principles outlined here—isolation, auditability, IAM-first access, and metadata enrichment—remain universal best practices for any enterprise AI deployment.</p>
]]></content:encoded></item><item><title>FastMCP and the Vinyl Collection Chatbot: Serverless Agentic AI in Action</title><link>https://lukelittle.com/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/</link><pubDate>Sat, 24 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/</guid><description>Building AI agent integrations with FastMCP and the Model Context Protocol—the universal standard that makes agents actually useful in production</description><content:encoded><![CDATA[<h2 id="what-is-the-model-context-protocol">What is the Model Context Protocol?</h2>
<p>The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Think of it as a universal adapter that lets any AI agent talk to any tool or data source without custom integration code.</p>
<p>Anthropic announced MCP in November 2024 and donated it to the Linux Foundation&rsquo;s Agentic AI Foundation about a year later, in December 2025. The adoption has been swift: OpenAI integrated it into ChatGPT, Google DeepMind uses it for Gemini agents, AWS built AgentCore around it, and development tools like Zed, Sourcegraph, Replit, and Codeium all support it. In just a few months, the community has built thousands of MCP servers. The protocol has become the de-facto standard for agent-to-tool communication.</p>
<p>If you&rsquo;ve read my previous articles on AWS DevOps Agent, Security Agent, and Kiro, you&rsquo;ve seen what these frontier agents do. This article explains how they actually work—the protocol layer that makes integration possible. More importantly, it shows how you can extend these agents to work with your proprietary systems.</p>
<h2 id="the-nm-integration-problem">The N×M integration problem</h2>
<p>Here&rsquo;s why MCP matters: You&rsquo;ve built an amazing AI agent. It reasons brilliantly, writes elegant code, debugs complex issues. But it can&rsquo;t access your company&rsquo;s data. Your customer records are in Salesforce. Your code is in GitHub. Your metrics are in Datadog. Your tickets are in Jira.</p>
<p>The traditional approach is building custom integrations. One for each pairing. Want your agent to read Salesforce? Build a Salesforce connector. Want it to access GitHub? Build a GitHub connector. Want it to work with both? Build both connectors. Want to switch LLM providers? Rebuild everything.</p>
<p>This is the N×M integration problem: N agents times M data sources equals N×M custom integrations. As your ecosystem grows, the complexity becomes unmanageable. Three agents talking to four services means twelve custom connectors. Ten agents and fifty systems means five hundred integrations to build and maintain.</p>
<p>Each connector requires understanding the target system&rsquo;s API, building authentication flows, handling rate limits and retries, writing serialization and deserialization logic, maintaining the connector as APIs change, and rebuilding for each new agent or LLM provider. This doesn&rsquo;t scale.</p>
<p>With MCP, you build the Salesforce MCP server once. After that, DevOps Agent, Security Agent, Kiro, Claude Desktop—any MCP client—can access Salesforce. No custom integration needed. This is why every major AI company adopted MCP within months of its announcement. It solves an existential scaling problem.</p>
<h2 id="how-mcp-works-the-architecture">How MCP works: The architecture</h2>
<p>MCP follows a client-server architecture using JSON-RPC 2.0 for communication. The protocol exposes three core primitives that servers can implement.</p>
<p><strong>Resources</strong> are like RESTful GET endpoints. They load information into LLM context—files, database records, API responses. When you ask an agent to retrieve a customer&rsquo;s support ticket history, it&rsquo;s accessing a resource.</p>
<p><strong>Tools</strong> are like RESTful POST endpoints. They execute code and produce side effects—creating Jira tickets, deploying code, sending Slack messages. When you tell an agent to create a high-priority ticket for a customer, it&rsquo;s using a tool.</p>
<p><strong>Prompts</strong> are reusable templates that guide interactions with structured patterns. Think &ldquo;Analyze this code for security vulnerabilities&rdquo; or &ldquo;Generate release notes from commits.&rdquo; They standardize how agents approach common tasks.</p>
<h2 id="the-communication-flow">The communication flow</h2>
<p>Here&rsquo;s how an agent actually communicates with an MCP server:</p>
<pre class="mermaid">sequenceDiagram
    participant Agent as MCP Client (Agent)
    participant Server as MCP Server (Tool)
    
    Agent-&gt;&gt;Server: 1. Initialize Connection
    Agent-&gt;&gt;Server: 2. List Available Tools
    Server-&gt;&gt;Agent: 3. Tool Definitions (JSON Schema)
    Note over Agent: Agent decides&lt;br/&gt;which tool to call
    Agent-&gt;&gt;Server: 4. Call Tool (with parameters)
    Note over Server: Validates params&lt;br/&gt;Executes logic&lt;br/&gt;Queries external API
    Server-&gt;&gt;Agent: 5. Tool Result
    Note over Agent: Agent uses result&lt;br/&gt;in response
</pre>

<p>The agent initializes a connection, requests the list of available tools, receives their JSON Schema definitions, decides which tool to call based on the user&rsquo;s request, sends the tool invocation with parameters, waits for the server to validate inputs and execute the logic, receives the result, and uses it to generate the final response. It&rsquo;s a clean request-response pattern.</p>
<h2 id="transport-mechanisms">Transport mechanisms</h2>
<p>MCP supports two primary transport methods. Standard Input/Output (stdio) is designed for local integration—the server runs as a subprocess and communicates via stdin/stdout. This is what Claude Desktop uses to run local MCP servers. It&rsquo;s simple, fast, and synchronous, but only works locally.</p>
<p>Server-Sent Events (SSE) is designed for remote integration. It uses HTTP-based streaming where the server pushes updates to the client. This is what you&rsquo;d use for cloud-hosted MCP servers and enterprise integrations. It&rsquo;s network-capable and supports real-time updates, but requires HTTP server infrastructure.</p>
<h2 id="the-protocol-layer">The protocol layer</h2>
<p>MCP messages are structured as JSON-RPC 2.0 calls. Here&rsquo;s what a tool call looks like:
Request from client to server:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;jsonrpc&#34;</span>: <span style="color:#e6db74">&#34;2.0&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;id&#34;</span>: <span style="color:#ae81ff">1</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;method&#34;</span>: <span style="color:#e6db74">&#34;tools/call&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;params&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;name&#34;</span>: <span style="color:#e6db74">&#34;get_customer_info&#34;</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;arguments&#34;</span>: {
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;customer_id&#34;</span>: <span style="color:#e6db74">&#34;cust_12345&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>Response from server to client:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;jsonrpc&#34;</span>: <span style="color:#e6db74">&#34;2.0&#34;</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;id&#34;</span>: <span style="color:#ae81ff">1</span>,
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;result&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;content&#34;</span>: [
</span></span><span style="display:flex;"><span>      {
</span></span><span style="display:flex;"><span>        <span style="color:#f92672">&#34;type&#34;</span>: <span style="color:#e6db74">&#34;text&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#f92672">&#34;text&#34;</span>: <span style="color:#e6db74">&#34;{\&#34;name\&#34;: \&#34;Acme Corp\&#34;, \&#34;tier\&#34;: \&#34;Enterprise\&#34;, \&#34;health\&#34;: \&#34;green\&#34;}&#34;</span>
</span></span><span style="display:flex;"><span>      }
</span></span><span style="display:flex;"><span>    ]
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>The beauty of this: you rarely write this JSON by hand. MCP SDKs for Python, TypeScript, Java, C#, and Kotlin abstract it away completely. Which brings us to FastMCP.</p>
<h2 id="building-with-fastmcp">Building with FastMCP</h2>
<p>Building MCP servers from scratch means handling JSON-RPC protocol details, managing connections, writing serialization boilerplate. It&rsquo;s tedious work that distracts from the actual business logic you want to implement.</p>
<p>FastMCP is the solution—a decorator-based Python framework that makes building MCP servers feel like writing FastAPI applications. Think of it as FastAPI for AI agents with the same elegant decorator patterns, Flask for tool integration with minimal boilerplate, or Express.js for MCP with simple and intuitive design.</p>
<p>Here&rsquo;s a complete, working MCP server:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize server</span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;My First Server&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Define a tool</span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">add_numbers</span>(a: int, b: int) <span style="color:#f92672">-&gt;</span> int:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Add two numbers together&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> a <span style="color:#f92672">+</span> b
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Run it</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">if</span> __name__ <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;__main__&#34;</span>:
</span></span><span style="display:flex;"><span>    mcp<span style="color:#f92672">.</span>run()
</span></span></code></pre></div><p>That&rsquo;s it. You now have an MCP server that exposes a tool, automatically generates JSON Schema from type hints, handles serialization and deserialization, provides error handling, supports both stdio and SSE transports, and works with any MCP client. No boilerplate. No protocol details. Just your logic.</p>
<h2 id="a-real-example-github-integration">A real example: GitHub integration</h2>
<p>Let&rsquo;s build something practical—an MCP server that lets agents interact with GitHub. This demonstrates how FastMCP handles real-world integrations with external APIs, authentication, and error handling.</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> github <span style="color:#f92672">import</span> Github
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize server</span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(
</span></span><span style="display:flex;"><span>    name<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;GitHub Integration&#34;</span>,
</span></span><span style="display:flex;"><span>    instructions<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;Use this server to interact with GitHub repositories&#34;</span>
</span></span><span style="display:flex;"><span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize GitHub client</span>
</span></span><span style="display:flex;"><span>github_token <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;GITHUB_TOKEN&#34;</span>)
</span></span><span style="display:flex;"><span>gh <span style="color:#f92672">=</span> Github(github_token)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_repository_info</span>(owner: str, repo: str) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Get information about a GitHub repository.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Args:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        owner: Repository owner (username or organization)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        repo: Repository name
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Returns:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        Repository details including stars, forks, issues
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    repository <span style="color:#f92672">=</span> gh<span style="color:#f92672">.</span>get_repo(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>owner<span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>repo<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;name&#34;</span>: repository<span style="color:#f92672">.</span>name,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;description&#34;</span>: repository<span style="color:#f92672">.</span>description,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;stars&#34;</span>: repository<span style="color:#f92672">.</span>stargazers_count,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;forks&#34;</span>: repository<span style="color:#f92672">.</span>forks_count,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;open_issues&#34;</span>: repository<span style="color:#f92672">.</span>open_issues_count,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;language&#34;</span>: repository<span style="color:#f92672">.</span>language,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;created_at&#34;</span>: repository<span style="color:#f92672">.</span>created_at<span style="color:#f92672">.</span>isoformat(),
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;updated_at&#34;</span>: repository<span style="color:#f92672">.</span>updated_at<span style="color:#f92672">.</span>isoformat()
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">create_issue</span>(
</span></span><span style="display:flex;"><span>    owner: str,
</span></span><span style="display:flex;"><span>    repo: str,
</span></span><span style="display:flex;"><span>    title: str,
</span></span><span style="display:flex;"><span>    body: str,
</span></span><span style="display:flex;"><span>    labels: list[str] <span style="color:#f92672">=</span> <span style="color:#66d9ef">None</span>
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Create a new issue in a GitHub repository.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Args:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        owner: Repository owner
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        repo: Repository name
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        title: Issue title
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        body: Issue description
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        labels: Optional list of labels to apply
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Returns:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        Created issue details
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    repository <span style="color:#f92672">=</span> gh<span style="color:#f92672">.</span>get_repo(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>owner<span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>repo<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    issue <span style="color:#f92672">=</span> repository<span style="color:#f92672">.</span>create_issue(
</span></span><span style="display:flex;"><span>        title<span style="color:#f92672">=</span>title,
</span></span><span style="display:flex;"><span>        body<span style="color:#f92672">=</span>body,
</span></span><span style="display:flex;"><span>        labels<span style="color:#f92672">=</span>labels <span style="color:#f92672">or</span> []
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;number&#34;</span>: issue<span style="color:#f92672">.</span>number,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;url&#34;</span>: issue<span style="color:#f92672">.</span>html_url,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;state&#34;</span>: issue<span style="color:#f92672">.</span>state,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;created_at&#34;</span>: issue<span style="color:#f92672">.</span>created_at<span style="color:#f92672">.</span>isoformat()
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">list_pull_requests</span>(
</span></span><span style="display:flex;"><span>    owner: str,
</span></span><span style="display:flex;"><span>    repo: str,
</span></span><span style="display:flex;"><span>    state: str <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;open&#34;</span>
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> list[dict]:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    List pull requests in a repository.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Args:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        owner: Repository owner
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        repo: Repository name
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        state: PR state (&#39;open&#39;, &#39;closed&#39;, &#39;all&#39;)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Returns:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        List of pull requests
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    repository <span style="color:#f92672">=</span> gh<span style="color:#f92672">.</span>get_repo(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>owner<span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>repo<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    pulls <span style="color:#f92672">=</span> repository<span style="color:#f92672">.</span>get_pulls(state<span style="color:#f92672">=</span>state)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> [
</span></span><span style="display:flex;"><span>        {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;number&#34;</span>: pr<span style="color:#f92672">.</span>number,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;title&#34;</span>: pr<span style="color:#f92672">.</span>title,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;author&#34;</span>: pr<span style="color:#f92672">.</span>user<span style="color:#f92672">.</span>login,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;state&#34;</span>: pr<span style="color:#f92672">.</span>state,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;created_at&#34;</span>: pr<span style="color:#f92672">.</span>created_at<span style="color:#f92672">.</span>isoformat(),
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;url&#34;</span>: pr<span style="color:#f92672">.</span>html_url
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">for</span> pr <span style="color:#f92672">in</span> pulls[:<span style="color:#ae81ff">10</span>]  <span style="color:#75715e"># Limit to 10 for brevity</span>
</span></span><span style="display:flex;"><span>    ]
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.resource</span>(<span style="color:#e6db74">&#34;repo://</span><span style="color:#e6db74">{owner}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{repo}</span><span style="color:#e6db74">/README&#34;</span>)
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_readme</span>(owner: str, repo: str) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Get the README content for a repository.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    This is exposed as a resource (not a tool) because it&#39;s
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    primarily for loading context into the LLM.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    repository <span style="color:#f92672">=</span> gh<span style="color:#f92672">.</span>get_repo(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>owner<span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>repo<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    readme <span style="color:#f92672">=</span> repository<span style="color:#f92672">.</span>get_readme()
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> readme<span style="color:#f92672">.</span>decoded_content<span style="color:#f92672">.</span>decode(<span style="color:#e6db74">&#39;utf-8&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">if</span> __name__ <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;__main__&#34;</span>:
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Run with stdio (for local Claude Desktop)</span>
</span></span><span style="display:flex;"><span>    mcp<span style="color:#f92672">.</span>run(transport<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;stdio&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Or run with SSE (for remote access)</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># mcp.run(transport=&#34;sse&#34;, port=8000)</span>
</span></span></code></pre></div><p>FastMCP handled type validation—ensuring owner and repo are strings—error handling so GitHub API failures return proper MCP errors, schema generation from docstrings and type hints, authentication flow using the GitHub token from environment variables, and transport abstraction so the same code works with stdio or SSE.</p>
<p>You focused on your business logic—what the tool actually does—and clear documentation where docstrings become tool descriptions. The framework handles everything else.</p>
<h2 id="advanced-fastmcp-features">Advanced FastMCP features</h2>
<p>FastMCP isn&rsquo;t just decorators—it&rsquo;s a full-featured framework with capabilities that become important as your integrations grow more sophisticated.</p>
<p><strong>Context injection</strong> lets you access MCP context and capabilities within your tools. This is useful for long-running operations where you want to send progress updates to the client:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP, Context
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> mcp.server.session <span style="color:#f92672">import</span> ServerSession
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Progress Example&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">def</span> <span style="color:#a6e22e">long_running_task</span>(
</span></span><span style="display:flex;"><span>    task_name: str,
</span></span><span style="display:flex;"><span>    ctx: Context[ServerSession, <span style="color:#66d9ef">None</span>],
</span></span><span style="display:flex;"><span>    steps: int <span style="color:#f92672">=</span> <span style="color:#ae81ff">5</span>
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Execute a task with progress updates.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Send progress notifications to client</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">await</span> ctx<span style="color:#f92672">.</span>info(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Starting: </span><span style="color:#e6db74">{</span>task_name<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">for</span> i <span style="color:#f92672">in</span> range(steps):
</span></span><span style="display:flex;"><span>        progress <span style="color:#f92672">=</span> (i <span style="color:#f92672">+</span> <span style="color:#ae81ff">1</span>) <span style="color:#f92672">/</span> steps <span style="color:#f92672">*</span> <span style="color:#ae81ff">100</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">await</span> ctx<span style="color:#f92672">.</span>progress(progress, <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Step </span><span style="color:#e6db74">{</span>i<span style="color:#f92672">+</span><span style="color:#ae81ff">1</span><span style="color:#e6db74">}</span><span style="color:#e6db74">/</span><span style="color:#e6db74">{</span>steps<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Do actual work here</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">await</span> asyncio<span style="color:#f92672">.</span>sleep(<span style="color:#ae81ff">1</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">await</span> ctx<span style="color:#f92672">.</span>info(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Completed: </span><span style="color:#e6db74">{</span>task_name<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Task </span><span style="color:#e6db74">{</span>task_name<span style="color:#e6db74">}</span><span style="color:#e6db74"> completed successfully&#34;</span>
</span></span></code></pre></div><p><strong>Sampling</strong> lets your server request LLM completions. This is powerful for agentic workflows where the server orchestrates LLM calls without needing API keys:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">def</span> <span style="color:#a6e22e">generate_commit_message</span>(
</span></span><span style="display:flex;"><span>    diff: str,
</span></span><span style="display:flex;"><span>    ctx: Context[ServerSession, <span style="color:#66d9ef">None</span>]
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Generate a commit message from a git diff.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Ask the client&#39;s LLM to generate the message</span>
</span></span><span style="display:flex;"><span>    result <span style="color:#f92672">=</span> <span style="color:#66d9ef">await</span> ctx<span style="color:#f92672">.</span>session<span style="color:#f92672">.</span>create_message(
</span></span><span style="display:flex;"><span>        messages<span style="color:#f92672">=</span>[{
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;role&#34;</span>: <span style="color:#e6db74">&#34;user&#34;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;content&#34;</span>: <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Generate a concise commit message for this diff:</span><span style="color:#ae81ff">\n\n</span><span style="color:#e6db74">{</span>diff<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>        }],
</span></span><span style="display:flex;"><span>        max_tokens<span style="color:#f92672">=</span><span style="color:#ae81ff">100</span>
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> result<span style="color:#f92672">.</span>content[<span style="color:#ae81ff">0</span>]<span style="color:#f92672">.</span>text
</span></span></code></pre></div><p><strong>Elicitation</strong> lets servers request additional information mid-operation. This is useful when the tool needs clarification from the user:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> mcp.types <span style="color:#f92672">import</span> ElicitRequestTextParams
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">def</span> <span style="color:#a6e22e">deploy_to_environment</span>(
</span></span><span style="display:flex;"><span>    service: str,
</span></span><span style="display:flex;"><span>    ctx: Context[ServerSession, <span style="color:#66d9ef">None</span>]
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Deploy a service to an environment.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Ask user which environment</span>
</span></span><span style="display:flex;"><span>    result <span style="color:#f92672">=</span> <span style="color:#66d9ef">await</span> ctx<span style="color:#f92672">.</span>elicit(
</span></span><span style="display:flex;"><span>        ElicitRequestTextParams(
</span></span><span style="display:flex;"><span>            mode<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;text&#34;</span>,
</span></span><span style="display:flex;"><span>            message<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;Which environment? (staging/production)&#34;</span>,
</span></span><span style="display:flex;"><span>            placeholder<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;staging&#34;</span>
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    environment <span style="color:#f92672">=</span> result<span style="color:#f92672">.</span>value
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Validate input</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> environment <span style="color:#f92672">not</span> <span style="color:#f92672">in</span> [<span style="color:#e6db74">&#34;staging&#34;</span>, <span style="color:#e6db74">&#34;production&#34;</span>]:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">ValueError</span>(<span style="color:#e6db74">&#34;Environment must be &#39;staging&#39; or &#39;production&#39;&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Proceed with deployment</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Deploying </span><span style="color:#e6db74">{</span>service<span style="color:#e6db74">}</span><span style="color:#e6db74"> to </span><span style="color:#e6db74">{</span>environment<span style="color:#e6db74">}</span><span style="color:#e6db74">...&#34;</span>
</span></span></code></pre></div><p><strong>Filesystem roots</strong> define security boundaries for which directories agents can access. FastMCP enforces that paths must be within configured roots that the client specifies when connecting:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">read_file</span>(path: str, ctx: Context[ServerSession, <span style="color:#66d9ef">None</span>]) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Read a file from the allowed workspace.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># FastMCP enforces that path must be within configured roots</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Client specifies roots when connecting:</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e">#   roots=[&#34;file:///safe/workspace&#34;]</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">with</span> open(path, <span style="color:#e6db74">&#39;r&#39;</span>) <span style="color:#66d9ef">as</span> f:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> f<span style="color:#f92672">.</span>read()
</span></span></code></pre></div><h2 id="case-study-vinyl-collection-chatbot">Case study: Vinyl collection chatbot</h2>
<p>To see how easy FastMCP makes building real-world integrations, I built a serverless chatbot that answers questions about my vinyl collection. The data comes from a Discogs export—a CSV file containing my complete record collection that I store in S3.</p>
<p>Ask it &ldquo;What Grimes records do I have?&rdquo; and it queries the CSV, parses the Discogs data, and returns results. Ask it &ldquo;What is vinyl?&rdquo; and it just answers from general knowledge. The bot intelligently decides when to use tools and when to rely on its training data.</p>
<p>Here&rsquo;s the complete implementation of the MCP server that queries the Discogs collection data:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> csv
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> io <span style="color:#f92672">import</span> StringIO
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize FastMCP server</span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;vinyl-collection-server&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># S3 client for accessing Discogs export</span>
</span></span><span style="display:flex;"><span>s3_client <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;s3&#39;</span>)
</span></span><span style="display:flex;"><span>DATA_BUCKET <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;my-vinyl-collection&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">query_vinyl_collection</span>(query_type: str, search_term: str, limit: int <span style="color:#f92672">=</span> <span style="color:#ae81ff">10</span>) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Query Luke&#39;s vinyl record collection from Discogs export data.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Args:
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        query_type: One of: artist, label, year, title, all
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        search_term: What to search for
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">        limit: Max results (default 10)
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Download Discogs CSV export from S3</span>
</span></span><span style="display:flex;"><span>    response <span style="color:#f92672">=</span> s3_client<span style="color:#f92672">.</span>get_object(Bucket<span style="color:#f92672">=</span>DATA_BUCKET, Key<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;discogs.csv&#39;</span>)
</span></span><span style="display:flex;"><span>    csv_content <span style="color:#f92672">=</span> response[<span style="color:#e6db74">&#39;Body&#39;</span>]<span style="color:#f92672">.</span>read()<span style="color:#f92672">.</span>decode(<span style="color:#e6db74">&#39;utf-8&#39;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Parse CSV</span>
</span></span><span style="display:flex;"><span>    records <span style="color:#f92672">=</span> []
</span></span><span style="display:flex;"><span>    csv_reader <span style="color:#f92672">=</span> csv<span style="color:#f92672">.</span>DictReader(StringIO(csv_content))
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">for</span> row <span style="color:#f92672">in</span> csv_reader:
</span></span><span style="display:flex;"><span>        records<span style="color:#f92672">.</span>append(row)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Filter based on query type</span>
</span></span><span style="display:flex;"><span>    matches <span style="color:#f92672">=</span> []
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">for</span> record <span style="color:#f92672">in</span> records:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> query_type <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;artist&#34;</span> <span style="color:#f92672">and</span> search_term<span style="color:#f92672">.</span>lower() <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Artist&#39;</span>]<span style="color:#f92672">.</span>lower():
</span></span><span style="display:flex;"><span>            matches<span style="color:#f92672">.</span>append(record)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> query_type <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;label&#34;</span> <span style="color:#f92672">and</span> search_term<span style="color:#f92672">.</span>lower() <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Label&#39;</span>]<span style="color:#f92672">.</span>lower():
</span></span><span style="display:flex;"><span>            matches<span style="color:#f92672">.</span>append(record)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> query_type <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;year&#34;</span> <span style="color:#f92672">and</span> search_term <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Released&#39;</span>]:
</span></span><span style="display:flex;"><span>            matches<span style="color:#f92672">.</span>append(record)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> query_type <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;title&#34;</span> <span style="color:#f92672">and</span> search_term<span style="color:#f92672">.</span>lower() <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Title&#39;</span>]<span style="color:#f92672">.</span>lower():
</span></span><span style="display:flex;"><span>            matches<span style="color:#f92672">.</span>append(record)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">elif</span> query_type <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;all&#34;</span>:
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">if</span> (search_term<span style="color:#f92672">.</span>lower() <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Artist&#39;</span>]<span style="color:#f92672">.</span>lower() <span style="color:#f92672">or</span> 
</span></span><span style="display:flex;"><span>                search_term<span style="color:#f92672">.</span>lower() <span style="color:#f92672">in</span> record[<span style="color:#e6db74">&#39;Title&#39;</span>]<span style="color:#f92672">.</span>lower()):
</span></span><span style="display:flex;"><span>                matches<span style="color:#f92672">.</span>append(record)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Format results</span>
</span></span><span style="display:flex;"><span>    results <span style="color:#f92672">=</span> []
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">for</span> record <span style="color:#f92672">in</span> matches[:limit]:
</span></span><span style="display:flex;"><span>        results<span style="color:#f92672">.</span>append(
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>record[<span style="color:#e6db74">&#39;Artist&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74"> - </span><span style="color:#e6db74">{</span>record[<span style="color:#e6db74">&#39;Title&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74"> &#34;</span>
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;(</span><span style="color:#e6db74">{</span>record[<span style="color:#e6db74">&#39;Label&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">, </span><span style="color:#e6db74">{</span>record[<span style="color:#e6db74">&#39;Released&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">)&#34;</span>
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#e6db74">&#34;</span><span style="color:#ae81ff">\n</span><span style="color:#e6db74">&#34;</span><span style="color:#f92672">.</span>join(results) <span style="color:#66d9ef">if</span> results <span style="color:#66d9ef">else</span> <span style="color:#e6db74">&#34;No matches found&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">if</span> __name__ <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;__main__&#34;</span>:
</span></span><span style="display:flex;"><span>    mcp<span style="color:#f92672">.</span>run(transport<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;sse&#34;</span>, port<span style="color:#f92672">=</span><span style="color:#ae81ff">8080</span>)
</span></span></code></pre></div><p>That&rsquo;s it. No JSON schema writing. No protocol handling. No tool registration boilerplate. FastMCP generates the schema from type hints and docstrings, handles the Model Context Protocol communication, converts everything to Bedrock&rsquo;s format, and manages tool execution. You write the business logic for parsing Discogs data. FastMCP handles everything else.</p>
<p>The whole system—frontend, Lambda, Bedrock integration, S3 data access—deployed in under two minutes. This isn&rsquo;t a toy example. It&rsquo;s running in production, costs about fifteen dollars a month, and demonstrates why FastMCP is becoming the standard way to build tools for AI agents.</p>
<p><img src="/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/vinyl-chatbot-architecture_hu_7d494f0977af11b7.webp" srcset="/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/vinyl-chatbot-architecture_hu_ab8250641e881360.webp 750w, /posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/vinyl-chatbot-architecture_hu_7d494f0977af11b7.webp 1101w" sizes="(max-width: 800px) 100vw, 750px"
       width="1101" height="760" alt="Vinyl Chatbot Architecture Diagram" loading="lazy" decoding="async"></p>
<p>This is the full loop: user asks a question, Bedrock decides whether to use a tool, FastMCP executes it against the Discogs export data, and the result is returned—all serverless.</p>
<h2 id="how-aws-frontier-agents-use-mcp">How AWS frontier agents use MCP</h2>
<p>Now let&rsquo;s connect this to the agents I&rsquo;ve covered in previous articles. These patterns show how MCP enables coordination between specialized agents.</p>
<p><strong>DevOps Agent scenario:</strong> A Lambda function starts timing out in production. DevOps Agent detects the anomaly through CloudWatch MCP Server where it reads metrics, logs, and traces. It queries the GitHub MCP Server to analyze recent code changes that might have introduced the issue. It creates an incident through the PagerDuty MCP Server and notifies the on-call engineer. Finally, it posts a summary to the incidents channel via the Slack MCP Server.</p>
<p>Without MCP, AWS would need custom connectors for every observability tool, ticketing system, and chat platform—an integration nightmare that doesn&rsquo;t scale. With MCP, DevOps Agent uses a standard MCP client and any MCP-compatible tool works instantly.</p>
<p><strong>Security Agent scenario:</strong> During a penetration test, Security Agent finds a SQL injection vulnerability. It reads the vulnerable code through the GitHub MCP Server, creates a security ticket via the Jira MCP Server, requests a code fix from Kiro through MCP, and once Kiro generates the fix, it creates a pull request back through GitHub&rsquo;s MCP Server and notifies the security team via Slack&rsquo;s MCP Server.</p>
<p>The agents coordinate through MCP without custom integration code. Security Agent discovers, Kiro fixes, DevOps Agent validates deployment—all using the same protocol.</p>
<p><strong>Kiro scenario:</strong> A developer asks Kiro to &ldquo;Add authentication to the user API endpoint.&rdquo; Kiro reads the existing code through the Filesystem MCP Server, checks authentication patterns in other services via the GitHub MCP Server, queries your company&rsquo;s authentication library documentation through an Internal Auth MCP Server, writes the updated code back through the Filesystem MCP Server, and creates a pull request via GitHub&rsquo;s MCP Server.</p>
<p>Kiro doesn&rsquo;t just generate code—it actively researches your codebase and standards through MCP servers, understanding context before making changes.</p>
<h2 id="building-enterprise-mcp-servers-real-patterns">Building enterprise MCP servers: Real patterns</h2>
<p>Here are three production-grade patterns that demonstrate how FastMCP handles enterprise integrations.</p>
<p><strong>Salesforce MCP Server</strong> shows integration with a major CRM system:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> simple_salesforce <span style="color:#f92672">import</span> Salesforce
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Salesforce CRM&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>sf <span style="color:#f92672">=</span> Salesforce(
</span></span><span style="display:flex;"><span>    username<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;SF_USERNAME&#34;</span>),
</span></span><span style="display:flex;"><span>    password<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;SF_PASSWORD&#34;</span>),
</span></span><span style="display:flex;"><span>    security_token<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;SF_SECURITY_TOKEN&#34;</span>)
</span></span><span style="display:flex;"><span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_account</span>(account_id: str) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Retrieve Salesforce account details.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> sf<span style="color:#f92672">.</span>Account<span style="color:#f92672">.</span>get(account_id)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">create_case</span>(
</span></span><span style="display:flex;"><span>    account_id: str,
</span></span><span style="display:flex;"><span>    subject: str,
</span></span><span style="display:flex;"><span>    description: str,
</span></span><span style="display:flex;"><span>    priority: str <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;Medium&#34;</span>
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Create a support case in Salesforce.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> sf<span style="color:#f92672">.</span>Case<span style="color:#f92672">.</span>create({
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;AccountId&#39;</span>: account_id,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;Subject&#39;</span>: subject,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;Description&#39;</span>: description,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;Priority&#39;</span>: priority
</span></span><span style="display:flex;"><span>    })
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">search_accounts</span>(query: str, limit: int <span style="color:#f92672">=</span> <span style="color:#ae81ff">10</span>) <span style="color:#f92672">-&gt;</span> list[dict]:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Search Salesforce accounts by name.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    soql <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;SELECT Id, Name, Industry, AnnualRevenue FROM Account WHERE Name LIKE &#39;%</span><span style="color:#e6db74">{</span>query<span style="color:#e6db74">}</span><span style="color:#e6db74">%&#39; LIMIT </span><span style="color:#e6db74">{</span>limit<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> sf<span style="color:#f92672">.</span>query(soql)[<span style="color:#e6db74">&#39;records&#39;</span>]
</span></span></code></pre></div><p>With this server deployed, any agent can now access Salesforce—DevOps Agent, Security Agent, Kiro, or any future agent you build. No custom integration needed per agent. Build once, use everywhere.</p>
<p><strong>Internal Wiki MCP Server</strong> demonstrates knowledge base integration:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> elasticsearch <span style="color:#f92672">import</span> Elasticsearch
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> os
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Company Wiki&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>es <span style="color:#f92672">=</span> Elasticsearch(
</span></span><span style="display:flex;"><span>    os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;ELASTICSEARCH_URL&#34;</span>),
</span></span><span style="display:flex;"><span>    api_key<span style="color:#f92672">=</span>os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;ELASTICSEARCH_API_KEY&#34;</span>)
</span></span><span style="display:flex;"><span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.resource</span>(<span style="color:#e6db74">&#34;wiki://</span><span style="color:#e6db74">{article_id}</span><span style="color:#e6db74">&#34;</span>)
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_wiki_article</span>(article_id: str) <span style="color:#f92672">-&gt;</span> str:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Get wiki article by ID.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    result <span style="color:#f92672">=</span> es<span style="color:#f92672">.</span>get(index<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;wiki&#34;</span>, id<span style="color:#f92672">=</span>article_id)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> result[<span style="color:#e6db74">&#39;_source&#39;</span>][<span style="color:#e6db74">&#39;content&#39;</span>]
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">search_wiki</span>(
</span></span><span style="display:flex;"><span>    query: str,
</span></span><span style="display:flex;"><span>    limit: int <span style="color:#f92672">=</span> <span style="color:#ae81ff">5</span>
</span></span><span style="display:flex;"><span>) <span style="color:#f92672">-&gt;</span> list[dict]:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Search company wiki for relevant articles.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Returns articles ranked by relevance.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    response <span style="color:#f92672">=</span> es<span style="color:#f92672">.</span>search(
</span></span><span style="display:flex;"><span>        index<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;wiki&#34;</span>,
</span></span><span style="display:flex;"><span>        body<span style="color:#f92672">=</span>{
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;query&#34;</span>: {
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#34;multi_match&#34;</span>: {
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#34;query&#34;</span>: query,
</span></span><span style="display:flex;"><span>                    <span style="color:#e6db74">&#34;fields&#34;</span>: [<span style="color:#e6db74">&#34;title^2&#34;</span>, <span style="color:#e6db74">&#34;content&#34;</span>, <span style="color:#e6db74">&#34;tags&#34;</span>]
</span></span><span style="display:flex;"><span>                }
</span></span><span style="display:flex;"><span>            },
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;size&#34;</span>: limit
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> [
</span></span><span style="display:flex;"><span>        {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;id&#34;</span>: hit[<span style="color:#e6db74">&#39;_id&#39;</span>],
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;title&#34;</span>: hit[<span style="color:#e6db74">&#39;_source&#39;</span>][<span style="color:#e6db74">&#39;title&#39;</span>],
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;excerpt&#34;</span>: hit[<span style="color:#e6db74">&#39;_source&#39;</span>][<span style="color:#e6db74">&#39;content&#39;</span>][:<span style="color:#ae81ff">200</span>] <span style="color:#f92672">+</span> <span style="color:#e6db74">&#34;...&#34;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;score&#34;</span>: hit[<span style="color:#e6db74">&#39;_score&#39;</span>],
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#34;url&#34;</span>: <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;https://wiki.company.com/articles/</span><span style="color:#e6db74">{</span>hit[<span style="color:#e6db74">&#39;_id&#39;</span>]<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">for</span> hit <span style="color:#f92672">in</span> response[<span style="color:#e6db74">&#39;hits&#39;</span>][<span style="color:#e6db74">&#39;hits&#39;</span>]
</span></span><span style="display:flex;"><span>    ]
</span></span></code></pre></div><p>This lets agents access institutional knowledge—Security Agent can read security policies, DevOps Agent can consult runbooks, and Kiro can reference coding standards, all from your internal wiki.</p>
<p><strong>Multi-Tool Orchestration MCP Server</strong> aggregates data from multiple sources:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> datadog_api_client <span style="color:#f92672">import</span> ApiClient, Configuration
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> datadog_api_client.v1.api.metrics_api <span style="color:#f92672">import</span> MetricsApi
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Operations Dashboard&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Initialize multiple clients</span>
</span></span><span style="display:flex;"><span>datadog_config <span style="color:#f92672">=</span> Configuration()
</span></span><span style="display:flex;"><span>datadog_config<span style="color:#f92672">.</span>api_key[<span style="color:#e6db74">&#39;apiKeyAuth&#39;</span>] <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;DD_API_KEY&#34;</span>)
</span></span><span style="display:flex;"><span>datadog_config<span style="color:#f92672">.</span>api_key[<span style="color:#e6db74">&#39;appKeyAuth&#39;</span>] <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;DD_APP_KEY&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_service_health</span>(service_name: str) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Get comprehensive health status for a service.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    Aggregates data from multiple monitoring systems.
</span></span></span><span style="display:flex;"><span><span style="color:#e6db74">    &#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Get Datadog metrics</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">with</span> ApiClient(datadog_config) <span style="color:#66d9ef">as</span> api_client:
</span></span><span style="display:flex;"><span>        api_instance <span style="color:#f92672">=</span> MetricsApi(api_client)
</span></span><span style="display:flex;"><span>        metrics <span style="color:#f92672">=</span> api_instance<span style="color:#f92672">.</span>query_metrics(
</span></span><span style="display:flex;"><span>            _from<span style="color:#f92672">=</span>int(time<span style="color:#f92672">.</span>time()) <span style="color:#f92672">-</span> <span style="color:#ae81ff">3600</span>,
</span></span><span style="display:flex;"><span>            to<span style="color:#f92672">=</span>int(time<span style="color:#f92672">.</span>time()),
</span></span><span style="display:flex;"><span>            query<span style="color:#f92672">=</span><span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;avg:system.cpu.user</span><span style="color:#ae81ff">{{</span><span style="color:#e6db74">service:</span><span style="color:#e6db74">{</span>service_name<span style="color:#e6db74">}</span><span style="color:#ae81ff">}}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Get PagerDuty incidents (if integrated)</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Get deployment status from CI/CD</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Aggregate everything</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;service&#34;</span>: service_name,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;status&#34;</span>: <span style="color:#e6db74">&#34;healthy&#34;</span>,  <span style="color:#75715e"># or &#34;degraded&#34;, &#34;down&#34;</span>
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;cpu_usage&#34;</span>: metrics[<span style="color:#e6db74">&#39;series&#39;</span>][<span style="color:#ae81ff">0</span>][<span style="color:#e6db74">&#39;pointlist&#39;</span>][<span style="color:#f92672">-</span><span style="color:#ae81ff">1</span>][<span style="color:#ae81ff">1</span>],
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;open_incidents&#34;</span>: <span style="color:#ae81ff">0</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;last_deployment&#34;</span>: <span style="color:#e6db74">&#34;2025-01-25T10:30:00Z&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;error_rate&#34;</span>: <span style="color:#ae81ff">0.02</span>  <span style="color:#75715e"># 2%</span>
</span></span><span style="display:flex;"><span>    }
</span></span></code></pre></div><p>This gives agents a unified view across disparate monitoring tools—one call returns comprehensive health status aggregated from Datadog, PagerDuty, and your CI/CD system.</p>
<h2 id="deploying-mcp-servers-production-patterns">Deploying MCP servers: Production patterns</h2>
<p>Claude Desktop is the easiest way to test MCP servers during development. Create your MCP server using the patterns shown above, then configure Claude Desktop to use it. Edit the configuration file at <code>~/Library/Application Support/Claude/claude_desktop_config.json</code> on Mac or <code>%APPDATA%\Claude\claude_desktop_config.json</code> on Windows:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-json" data-lang="json"><span style="display:flex;"><span>{
</span></span><span style="display:flex;"><span>  <span style="color:#f92672">&#34;mcpServers&#34;</span>: {
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;github&#34;</span>: {
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;command&#34;</span>: <span style="color:#e6db74">&#34;python&#34;</span>,
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;args&#34;</span>: [<span style="color:#e6db74">&#34;/path/to/github_server.py&#34;</span>],
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;env&#34;</span>: {
</span></span><span style="display:flex;"><span>        <span style="color:#f92672">&#34;GITHUB_TOKEN&#34;</span>: <span style="color:#e6db74">&#34;your_token_here&#34;</span>
</span></span><span style="display:flex;"><span>      }
</span></span><span style="display:flex;"><span>    },
</span></span><span style="display:flex;"><span>    <span style="color:#f92672">&#34;salesforce&#34;</span>: {
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;command&#34;</span>: <span style="color:#e6db74">&#34;python&#34;</span>,
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;args&#34;</span>: [<span style="color:#e6db74">&#34;/path/to/salesforce_server.py&#34;</span>],
</span></span><span style="display:flex;"><span>      <span style="color:#f92672">&#34;env&#34;</span>: {
</span></span><span style="display:flex;"><span>        <span style="color:#f92672">&#34;SF_USERNAME&#34;</span>: <span style="color:#e6db74">&#34;your_username&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#f92672">&#34;SF_PASSWORD&#34;</span>: <span style="color:#e6db74">&#34;your_password&#34;</span>
</span></span><span style="display:flex;"><span>      }
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>Restart Claude Desktop and your agents will have access to these tools.</p>
<p>For production deployments, use remote MCP servers with SSE transport:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># server.py</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Production GitHub Server&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># ... (your tools here)</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">if</span> __name__ <span style="color:#f92672">==</span> <span style="color:#e6db74">&#34;__main__&#34;</span>:
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Run with SSE on custom port</span>
</span></span><span style="display:flex;"><span>    mcp<span style="color:#f92672">.</span>run(transport<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;sse&#34;</span>, port<span style="color:#f92672">=</span><span style="color:#ae81ff">8080</span>)
</span></span></code></pre></div><p>You can deploy this on AWS Lambda with a Lambda function URL, ECS or Fargate for containerized auto-scaling workloads, traditional EC2 instances, or AWS App Runner for fully managed hosting.</p>
<p>For AWS Bedrock AgentCore integration, configure the client like this:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> mcp <span style="color:#f92672">import</span> ClientSession, SseServerParameters
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> mcp.client.sse <span style="color:#f92672">import</span> sse_client
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>server_params <span style="color:#f92672">=</span> SseServerParameters(
</span></span><span style="display:flex;"><span>    url<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;https://mcp.company.com/github&#34;</span>,
</span></span><span style="display:flex;"><span>    headers<span style="color:#f92672">=</span>{<span style="color:#e6db74">&#34;Authorization&#34;</span>: <span style="color:#e6db74">&#34;Bearer YOUR_TOKEN&#34;</span>}
</span></span><span style="display:flex;"><span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">with</span> sse_client(server_params) <span style="color:#66d9ef">as</span> (read, write):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">async</span> <span style="color:#66d9ef">with</span> ClientSession(read, write) <span style="color:#66d9ef">as</span> session:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">await</span> session<span style="color:#f92672">.</span>initialize()
</span></span><span style="display:flex;"><span>        tools <span style="color:#f92672">=</span> <span style="color:#66d9ef">await</span> session<span style="color:#f92672">.</span>list_tools()
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Use tools...</span>
</span></span></code></pre></div><h2 id="aws-bedrock-agentcore-and-mcp-integration">AWS Bedrock AgentCore and MCP integration</h2>
<p>This is where everything comes together. AgentCore natively supports MCP through its Gateway service.</p>
<pre class="mermaid">graph TB
    subgraph AgentCore[&#34;AWS Bedrock AgentCore&#34;]
        Runtime[AgentCore Runtime&lt;br/&gt;LangChain/LangGraph/Custom]
        Gateway[AgentCore Gateway&lt;br/&gt;Tool Discovery&lt;br/&gt;MCP Client&lt;br/&gt;Policy Enforcement&lt;br/&gt;Authentication]
        Runtime --&gt; Gateway
    end
    
    Gateway --&gt; GH[GitHub&lt;br/&gt;MCP Server]
    Gateway --&gt; SF[Salesforce&lt;br/&gt;MCP Server]
    Gateway --&gt; Custom[Your Custom&lt;br/&gt;MCP Server]
</pre>

<p>Integrating your FastMCP server with AgentCore follows three steps:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># Step 1: Build your MCP server with FastMCP</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> fastmcp <span style="color:#f92672">import</span> FastMCP
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>mcp <span style="color:#f92672">=</span> FastMCP(<span style="color:#e6db74">&#34;Custom Enterprise Tool&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">query_internal_database</span>(sql: str) <span style="color:#f92672">-&gt;</span> list[dict]:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Query internal PostgreSQL database.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Your implementation</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">pass</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Deploy to AWS (Lambda, ECS, etc.)</span>
</span></span><span style="display:flex;"><span>mcp<span style="color:#f92672">.</span>run(transport<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;sse&#34;</span>, port<span style="color:#f92672">=</span><span style="color:#ae81ff">8080</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Step 2: Register with AgentCore Gateway</span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># In AgentCore Console or via SDK:</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">import</span> boto3
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>agentcore <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>client(<span style="color:#e6db74">&#39;bedrock-agentcore&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>agentcore<span style="color:#f92672">.</span>register_mcp_server(
</span></span><span style="display:flex;"><span>    name<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;internal-database&#34;</span>,
</span></span><span style="display:flex;"><span>    url<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;https://internal-mcp.company.com&#34;</span>,
</span></span><span style="display:flex;"><span>    authentication<span style="color:#f92672">=</span>{
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;type&#39;</span>: <span style="color:#e6db74">&#39;bearer&#39;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;tokenSecret&#39;</span>: <span style="color:#e6db74">&#39;arn:aws:secretsmanager:...&#39;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># Step 3: Use in your agent</span>
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> langchain.agents <span style="color:#f92672">import</span> create_react_agent
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> bedrock_agentcore <span style="color:#f92672">import</span> BedrockAgentCoreApp
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>app <span style="color:#f92672">=</span> BedrockAgentCoreApp()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@app.entrypoint</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">my_agent</span>(request):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Your agent automatically has access to all</span>
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># MCP servers registered in AgentCore Gateway</span>
</span></span><span style="display:flex;"><span>    agent <span style="color:#f92672">=</span> create_react_agent(model, tools)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> agent<span style="color:#f92672">.</span>invoke(request)
</span></span></code></pre></div><p>AgentCore Gateway gives you automatic tool discovery where it lists all registered MCP servers, policy enforcement so you can define which agents can use which tools, authentication handling through AgentCore Identity for OAuth and API keys, observability with every MCP call logged to CloudWatch, and scaling managed entirely by AgentCore infrastructure.</p>
<h2 id="security-and-best-practices">Security and best practices</h2>
<p>Input validation is critical. Always validate and sanitize inputs before execution:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">query_database</span>(sql: str) <span style="color:#f92672">-&gt;</span> list[dict]:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Execute a SQL query (SELECT only).&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># ✅ GOOD: Validate before execution</span>
</span></span><span style="display:flex;"><span>    dangerous_keywords <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#39;DROP&#39;</span>, <span style="color:#e6db74">&#39;DELETE&#39;</span>, <span style="color:#e6db74">&#39;UPDATE&#39;</span>, <span style="color:#e6db74">&#39;INSERT&#39;</span>, <span style="color:#e6db74">&#39;ALTER&#39;</span>]
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> any(keyword <span style="color:#f92672">in</span> sql<span style="color:#f92672">.</span>upper() <span style="color:#66d9ef">for</span> keyword <span style="color:#f92672">in</span> dangerous_keywords):
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">ValueError</span>(<span style="color:#e6db74">&#34;Only SELECT queries allowed&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Additional validation</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> <span style="color:#e6db74">&#39;;&#39;</span> <span style="color:#f92672">in</span> sql:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">ValueError</span>(<span style="color:#e6db74">&#34;Multiple statements not allowed&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> execute_query(sql)
</span></span></code></pre></div><p>Apply the principle of least privilege by granting minimal permissions. Instead of using full admin access tokens, use read-only or scoped tokens:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># BAD: Full admin access</span>
</span></span><span style="display:flex;"><span>github_token <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;GITHUB_ADMIN_TOKEN&#34;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># GOOD: Read-only or scoped token</span>
</span></span><span style="display:flex;"><span>github_token <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>getenv(<span style="color:#e6db74">&#34;GITHUB_READONLY_TOKEN&#34;</span>)
</span></span></code></pre></div><p>Rate limiting protects your backend systems from abuse:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> functools <span style="color:#f92672">import</span> lru_cache
</span></span><span style="display:flex;"><span><span style="color:#f92672">from</span> time <span style="color:#f92672">import</span> time
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@lru_cache</span>(maxsize<span style="color:#f92672">=</span><span style="color:#ae81ff">128</span>)
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">rate_limit_key</span>(user_id: str) <span style="color:#f92672">-&gt;</span> int:
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> int(time() <span style="color:#f92672">/</span> <span style="color:#ae81ff">60</span>)  <span style="color:#75715e"># 1-minute windows</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">expensive_operation</span>(user_id: str, query: str) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Rate-limited operation.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Simple rate limiting (use Redis in production)</span>
</span></span><span style="display:flex;"><span>    key <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;</span><span style="color:#e6db74">{</span>user_id<span style="color:#e6db74">}</span><span style="color:#e6db74">:</span><span style="color:#e6db74">{</span>rate_limit_key(user_id)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> request_count(key) <span style="color:#f92672">&gt;</span> <span style="color:#ae81ff">10</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">Exception</span>(<span style="color:#e6db74">&#34;Rate limit exceeded: 10 requests per minute&#34;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    increment_count(key)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> perform_operation(query)
</span></span></code></pre></div><p>Error handling should provide helpful errors without leaking internal details:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">from</span> mcp.types <span style="color:#f92672">import</span> ToolError
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">@mcp.tool</span>()
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_sensitive_data</span>(resource_id: str) <span style="color:#f92672">-&gt;</span> dict:
</span></span><span style="display:flex;"><span>    <span style="color:#e6db74">&#34;&#34;&#34;Access protected resource.&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> fetch_resource(resource_id)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">PermissionError</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># GOOD: Generic error message</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> ToolError(<span style="color:#e6db74">&#34;Access denied: insufficient permissions&#34;</span>)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># BAD: Would leak internal details</span>
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># raise ToolError(f&#34;Database error: {str(e)}&#34;)</span>
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># GOOD: Log internally, return generic error</span>
</span></span><span style="display:flex;"><span>        logger<span style="color:#f92672">.</span>error(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Internal error: </span><span style="color:#e6db74">{</span>e<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>, exc_info<span style="color:#f92672">=</span><span style="color:#66d9ef">True</span>)
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> ToolError(<span style="color:#e6db74">&#34;An internal error occurred&#34;</span>)
</span></span></code></pre></div><h2 id="the-future-of-mcp">The future of MCP</h2>
<p>Based on current trajectory and industry patterns, MCP is evolving rapidly. In the near term—the next six to twelve months—expect broader ecosystem adoption with every major AI platform supporting MCP (OpenAI, Google, Microsoft, AWS), IDE integration in VS Code, JetBrains, and Zed, and browser extensions for Arc and Chrome. Enhanced security will standardize OAuth 2.0 flows, introduce fine-grained permission models, and establish audit logging specifications. Performance optimizations will add caching strategies, batch operations, and streaming support for large results.</p>
<p>The medium term—one to two years out—will bring enterprise features like an MCP server marketplace similar to AWS Marketplace, a certified and verified server registry, and SLA guarantees with monitoring. Advanced capabilities will enable multi-step workflows that chain tools across servers, state management for sessions and transactions, and push notifications where servers can alert clients. Developer experience improvements will include visual server builders for low-code MCP development, testing frameworks like pytest for MCP, and observability SDKs with OpenTelemetry integration.</p>
<p>Long term—two to five years—MCP becomes the HTTP of agentic AI. Every enterprise system will provide an MCP interface, legacy systems will get wrapped in MCP adapters, and universal agent communication becomes standard. Economic models will emerge with pay-per-call MCP servers, SaaS subscriptions for premium MCP servers, and even agent-to-agent commerce. The platform shift means developers stop building custom integrations entirely, &ldquo;MCP-first&rdquo; becomes the default architecture, and agents compose tools like developers compose libraries today.</p>
<h2 id="next-steps">Next steps</h2>
<p>If you&rsquo;re building agents, start by learning FastMCP basics in about thirty minutes—install it with pip, build the hello world server, and test with Claude Desktop. Then build an MCP server for your most-used tool in a couple hours—GitHub, Jira, Salesforce, or your internal API—following the patterns in this article and deploying locally to test thoroughly. Finally, integrate with your agents in about an hour by adding it to your Claude Desktop config or integrating with AgentCore Gateway, then verify agents can discover and use the tools.</p>
<p>If you&rsquo;re at an enterprise, inventory your tool landscape over a day—list all systems agents need to access, identify which have existing MCP servers, and plan which you need to build custom. Build two to three pilot MCP servers in a week, starting with high-value, low-risk systems, using FastMCP for rapid development, and deploying to staging to test with real agents. Establish governance over two weeks by defining which agents can use which tools, implementing policy enforcement through AgentCore Policy, and setting up observability and monitoring. Then scale gradually over time, adding MCP servers incrementally, monitoring usage and costs, and iterating based on agent behavior.</p>
<h2 id="why-this-matters">Why this matters</h2>
<p>Here&rsquo;s the reality: The value of AI agents isn&rsquo;t in the agents themselves—it&rsquo;s in what they can access. An agent that can reason brilliantly but can&rsquo;t touch your systems is a toy. An agent that can access your systems through MCP is a tool. A fleet of specialized agents, coordinating through MCP, is a transformation.</p>
<p>DevOps Agent is powerful. Security Agent is powerful. Kiro is powerful. But they&rsquo;re only as powerful as the tools you give them through MCP.</p>
<p>That&rsquo;s why understanding MCP isn&rsquo;t optional—it&rsquo;s foundational. It&rsquo;s the difference between agents that demo well versus agents that actually work, prototypes in staging versus production deployments, vendor lock-in versus true interoperability.</p>
<p>FastMCP makes building these integrations trivial. The protocol is stable. The ecosystem is growing rapidly. The question isn&rsquo;t &ldquo;Should we adopt MCP?&rdquo; It&rsquo;s &ldquo;Which systems do we connect first?&rdquo;</p>
<h2 id="resources">Resources</h2>
<p>For further exploration, start with the MCP Official Specification at modelcontextprotocol.io, the FastMCP GitHub repository at github.com/jlowin/fastmcp, and FastMCP Documentation at gofastmcp.com. The MCP Python SDK is available at github.com/modelcontextprotocol/python-sdk. AWS AgentCore Gateway documentation can be found in the AWS Bedrock AgentCore Documentation, and the MCP Server Registry is at github.com/modelcontextprotocol/servers.</p>
<p>If you&rsquo;re building agents or integrating AI into enterprise systems, MCP is infrastructure you need to understand. Not as a nice-to-have, but as a prerequisite.
Start with FastMCP. Build a server for your most-used tool. See how quickly you can give agents superpowers.
The agents are here. The protocol is stable. The ecosystem is exploding.
Time to plug in.</p>
]]></content:encoded></item><item><title>Your AI Security Engineer: Inside AWS Security Agent</title><link>https://lukelittle.com/posts/2026/01/your-ai-security-engineer-inside-aws-security-agent/</link><pubDate>Fri, 23 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/your-ai-security-engineer-inside-aws-security-agent/</guid><description>How AWS&amp;#39;s frontier agent is changing application security - an autonomous security engineer that works while you ship</description><content:encoded><![CDATA[<p>Here&rsquo;s what should make every security leader uncomfortable: organizations routinely deploy vulnerable code to production to meet delivery deadlines.</p>
<p>Not because they don&rsquo;t care about security. Because security can&rsquo;t keep up.</p>
<p>Over 60% of organizations update their web applications weekly or more frequently. Nearly 75% test those applications for security monthly or less. The math doesn&rsquo;t work. The gap between development velocity and security validation grows wider every sprint.</p>
<p>At re:Invent 2025, AWS CEO Matt Garman announced <strong>AWS Security Agent</strong>—not as another security scanning tool to add to the pile, but as a fundamentally different approach to the problem.</p>
<p>Security Agent is a frontier agent that operates autonomously throughout the development lifecycle, conducting design reviews, analyzing code, and executing penetration tests on-demand, matching the pace of modern development instead of being its bottleneck.</p>
<p>You can watch the AWS Security Agent announcement here: <a href="https://www.youtube.com/watch?v=oMY0tUDEhtY">https://www.youtube.com/watch?v=oMY0tUDEhtY</a></p>
<p>This is AWS&rsquo;s bet that security doesn&rsquo;t scale through more manual reviews—it scales through intelligent automation that understands your applications, your standards, and your threats.</p>
<h2 id="what-makes-an-agent-frontier-class">What makes an agent &ldquo;frontier-class&rdquo;</h2>
<p>AWS uses the term &ldquo;frontier agent&rdquo; to mean something specific. It&rsquo;s not just GPT-4 with security tools.</p>
<p><strong>1. Autonomous goal-directed behavior</strong></p>
<p>Traditional security: &ldquo;Run this SAST scan and give me findings.&rdquo;<br>
Frontier agent: &ldquo;Validate this application meets our security requirements&rdquo; → agent figures out how</p>
<p>You give it an objective, it decomposes the problem, forms hypotheses, collects evidence, and executes—without asking you for step-by-step guidance.</p>
<p><strong>2. Multi-agent coordination</strong></p>
<p>Security Agent doesn&rsquo;t work alone. When conducting a penetration test, it spawns specialized sub-agents—one for authentication bypass, another for authorization flaws, a third for injection attacks. These agents work concurrently, investigating multiple attack vectors simultaneously and coordinating across findings.</p>
<p><strong>3. Long-running and context-aware</strong></p>
<p>Here&rsquo;s the paradigm shift: Security Agent maintains persistent understanding of your applications.</p>
<p>Traditional security tools forget everything between scans. Security Agent learns:</p>
<ul>
<li>Your organizational security requirements</li>
<li>Your application architecture and data flows</li>
<li>Your code patterns and common vulnerabilities</li>
<li>Your historical findings and remediation approaches</li>
</ul>
<p>When testing your API, it doesn&rsquo;t just throw generic payloads. It understands your authentication mechanism, maps your business logic, and targets application-specific vulnerabilities.</p>
<h2 id="how-it-actually-works">How it actually works</h2>
<p>Security Agent operates across three phases of the development lifecycle, each with different capabilities:</p>
<p><strong>Phase 1: Design Security Review</strong></p>
<p>Before code exists, upload design documents, architecture diagrams, and threat models. Security Agent analyzes against:</p>
<ul>
<li>AWS security best practices</li>
<li>Your organization&rsquo;s security requirements</li>
<li>Common architectural vulnerabilities</li>
<li>Threat modeling patterns</li>
</ul>
<p>Output: Security risk analysis with specific remediation guidance, in minutes instead of days.</p>
<p><strong>Phase 2: Code Security Review</strong></p>
<p>During development, GitHub integration provides automated security feedback on every pull request. Security Agent validates:</p>
<ul>
<li>Organizational security requirements (approved libraries, logging standards, data policies)</li>
<li>OWASP Top 10 vulnerabilities</li>
<li>Code-level security anti-patterns</li>
<li>Compliance with your defined security standards</li>
</ul>
<p>Developers get immediate feedback in their workflow—no context switching required.</p>
<p><strong>Phase 3: On-Demand Penetration Testing</strong></p>
<p>Whenever needed—pre-deployment, post-change, or on a schedule—Security Agent conducts comprehensive penetration testing.</p>
<p>Unlike traditional scanners, it:</p>
<ul>
<li>Builds understanding from your source code and architecture</li>
<li>Creates customized attack plans based on your specific stack</li>
<li>Executes multi-step attack chains (not just single-payload scans)</li>
<li>Tests business logic vulnerabilities</li>
<li>Validates findings to eliminate false positives</li>
<li>Generates pull requests with remediation code</li>
</ul>
<h2 id="the-penetration-testing-loop">The penetration testing loop</h2>
<p>When you trigger a pentest, here&rsquo;s what happens:</p>
<pre class="mermaid">graph TB
    Input[Target URLs + Code + Docs] --&gt; Context[Build Application Understanding]
    
    Context --&gt; Discovery[Discover Attack Surface&lt;br/&gt;Map endpoints, APIs, flows]
    
    Discovery --&gt; Planning[Planning Agent&lt;br/&gt;Create customized attack plan]
    
    Planning --&gt; Testing[Specialized Testing Agents]
    
    subgraph Testing[&#34; &#34;]
        Auth[Auth Bypass]
        Authz[Authorization]
        Inject[Injection]
        Logic[Business Logic]
    end
    
    Testing --&gt; Validate[Validator Agent&lt;br/&gt;Eliminate false positives]
    
    Validate --&gt; Remediate[Remediation Agent&lt;br/&gt;Generate code fixes]
    
    Remediate --&gt; PR[Pull Request with Fix]
</pre>

<p>The clever part is the <strong>context-aware testing</strong>. Security Agent analyzes your source code to understand:</p>
<ul>
<li>Which endpoints are public vs. internal</li>
<li>What authentication patterns you use</li>
<li>How data flows through your system</li>
<li>What your actual threat model looks like</li>
</ul>
<p>When it sees JWT tokens, it doesn&rsquo;t just test for SQL injection—it focuses on JWT-specific attacks like algorithm confusion, token tampering, and replay attacks.</p>
<h2 id="agent-spaces-and-organizational-requirements">Agent Spaces and organizational requirements</h2>
<p>Everything starts with an <strong>Agent Space</strong>—the workspace where Security Agent operates and the security boundary for what it can access.</p>
<p>You might structure Agent Spaces as:</p>
<ul>
<li><strong>Per-application:</strong> One space for your customer portal, another for your admin dashboard</li>
<li><strong>Per-team:</strong> One space per development team managing their portfolio</li>
<li><strong>Per-environment:</strong> Separate spaces for staging vs. production testing</li>
</ul>
<p>The powerful part: you define your organization&rsquo;s security standards once, centrally:</p>
<pre tabindex="0"><code>Authentication Requirements:
- OAuth 2.0 for all API endpoints
- JWT tokens with 15-minute expiration
- MFA required for admin functions

Data Protection Requirements:
- PII encrypted at rest (KMS)
- TLS 1.3 for data in transit
- No credit card data in logs

Logging Requirements:
- Correlation IDs on all requests
- Auth failures logged with IP
- No PII in application logs
</code></pre><p>These requirements automatically apply to all Agent Spaces, enforced during both design reviews and code reviews. Consistent enforcement across the organization—no more &ldquo;this team follows standards, that team doesn&rsquo;t.&rdquo;</p>
<h2 id="deploying-it-practical-walkthrough">Deploying it (practical walkthrough)</h2>
<p>AWS provides console-based setup. Here&rsquo;s the flow:</p>
<p><strong>Step 1: Create Agent Space</strong></p>
<pre tabindex="0"><code>AWS Console → Security Agent → Create Agent Space
- Name: &#34;production-security&#34;
- Agent role: Auto-created
</code></pre><p><strong>Step 2: Define Security Requirements</strong></p>
<pre tabindex="0"><code>Security Requirements → Add requirements:
- Authentication standards
- Authorization patterns  
- Data protection policies
- Logging requirements
- Compliance frameworks
</code></pre><p><strong>Step 3: Configure Penetration Testing</strong></p>
<pre tabindex="0"><code>Agent Space → Enable penetration testing
- Add target domains (verify ownership)
- Configure CloudWatch logging
- Set up VPC access (for private apps)
- Add credentials to Secrets Manager
</code></pre><p><strong>Step 4: Integrate with GitHub</strong></p>
<pre tabindex="0"><code>Install AWS Security Agent GitHub App
- Authorize repository access
- Enable code review for Agent Space
- Auto-review triggered on PRs
</code></pre><p><strong>Step 5: Execute Penetration Test</strong></p>
<pre tabindex="0"><code>Security Agent Web App → Create pentest
- Target: https://staging.app.example.com
- Authentication: From Secrets Manager
- Attach: Source code + design docs
- Enable automatic remediation PRs
</code></pre><p>Watch as Security Agent:</p>
<ul>
<li>Discovers attack surface</li>
<li>Executes targeted scenarios</li>
<li>Validates findings</li>
<li>Creates PRs with fixes</li>
</ul>
<h2 id="testing-it">Testing it</h2>
<p>AWS provides real test scenarios. Run these before connecting production:</p>
<p><strong>Test 1: API authentication bypass</strong></p>
<p>Deploy an API with intentionally weak JWT validation:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># Weak JWT verification</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">verify_token</span>(token):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Missing signature validation</span>
</span></span><span style="display:flex;"><span>    payload <span style="color:#f92672">=</span> jwt<span style="color:#f92672">.</span>decode(token, verify<span style="color:#f92672">=</span><span style="color:#66d9ef">False</span>)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> payload[<span style="color:#e6db74">&#39;user_id&#39;</span>]
</span></span></code></pre></div><p>Trigger Security Agent pentest, watch it:</p>
<ul>
<li>Identify JWT usage</li>
<li>Test algorithm confusion</li>
<li>Attempt signature bypass</li>
<li>Generate exploit proof</li>
<li>Create PR with proper validation</li>
</ul>
<p><strong>Test 2: SQL injection in query</strong></p>
<p>Deploy an endpoint with SQL injection:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># Vulnerable query</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_user</span>(user_id):
</span></span><span style="display:flex;"><span>    query <span style="color:#f92672">=</span> <span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;SELECT * FROM users WHERE id = </span><span style="color:#e6db74">{</span>user_id<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> db<span style="color:#f92672">.</span>execute(query)
</span></span></code></pre></div><p>Security Agent should:</p>
<ul>
<li>Detect SQL construction</li>
<li>Test injection vectors</li>
<li>Confirm exploitability</li>
<li>Recommend parameterized queries</li>
</ul>
<h2 id="whats-not-ready-yet-the-honest-limitations">What&rsquo;s not ready yet (the honest limitations)</h2>
<h3 id="1-us-east-1-only">1. us-east-1 only</h3>
<p>Security Agent is currently only available in us-east-1.</p>
<p>If you have data residency requirements (GDPR, finance, healthcare), this is a blocker. Applications in other regions must be accessible from us-east-1.</p>
<p>Mitigation: Test staging/dev environments. AWS will expand regions post-GA.</p>
<h3 id="2-github-only-code-review">2. GitHub-only code review</h3>
<p>Currently only GitHub is supported for automated code review.</p>
<p>What&rsquo;s missing:</p>
<ul>
<li>GitLab</li>
<li>Bitbucket</li>
<li>AWS CodeCommit</li>
<li>Azure DevOps</li>
</ul>
<p>Workaround: You can still use penetration testing by providing code via S3.</p>
<h3 id="3-not-a-replacement-for-professional-pentesting">3. Not a replacement for professional pentesting</h3>
<p>Security Agent is powerful but not guaranteed to discover all vulnerabilities. It&rsquo;s best used for continuous testing at development velocity.</p>
<p>AWS&rsquo;s position: &ldquo;AWS Security Agent is not a professional penetration testing service, and we encourage users to integrate AWS Security Agent into their security review workflow.&rdquo;</p>
<p>The right mental model: Security Agent is your continuous validation layer. Professional pentesters are your comprehensive pre-launch audit.</p>
<h3 id="4-false-positive-management">4. False positive management</h3>
<p>While Validator Agents significantly reduce false positives, AI-powered testing will never be 100% perfect.</p>
<p>What AWS does:</p>
<ul>
<li>Only reports high/medium confidence findings</li>
<li>Hides unverified findings by default</li>
<li>Provides reproducible exploit paths</li>
</ul>
<p>What you should do:</p>
<ul>
<li>Review findings with security expertise</li>
<li>Validate critical findings independently</li>
<li>Use CloudWatch logs to understand agent reasoning</li>
</ul>
<h3 id="5-learning-curve">5. Learning curve</h3>
<p>Security Agent builds topology understanding over time. Early investigations might be less accurate.</p>
<p>Mitigation:</p>
<ul>
<li>Run test assessments to let it learn</li>
<li>Tag resources consistently</li>
<li>Document dependencies explicitly</li>
</ul>
<h2 id="should-you-actually-use-this">Should you actually use this?</h2>
<p><strong>Use it if:</strong></p>
<ul>
<li>Your development velocity is outpacing security capacity</li>
<li>You deploy weekly but test security monthly</li>
<li>You knowingly ship vulnerable code to meet deadlines</li>
<li>You want to scale security across your entire portfolio</li>
<li>You&rsquo;re AWS-native (tight integration benefits)</li>
<li>You&rsquo;re comfortable with preview-phase tech</li>
</ul>
<p><strong>Wait if:</strong></p>
<ul>
<li>You need multi-region support now</li>
<li>You use source control other than GitHub (for code review)</li>
<li>Your security process already keeps pace with development</li>
<li>You need production SLAs (preview = no SLAs)</li>
<li>You require deterministic pricing</li>
</ul>
<p><strong>Key insight:</strong> Security Agent amplifies good practices and exposes bad ones. If your infrastructure is poorly tagged, deployments aren&rsquo;t tracked, and requirements are scattered, it&rsquo;ll struggle. But if you have solid foundations, it can be transformative.</p>
<h2 id="my-honest-take">My honest take</h2>
<p>This is the future of application security. Not because AI replaces security engineers, but because it handles undifferentiated heavy lifting.</p>
<p>The traditional model: Security is a gate. Development builds features, security reviews them, findings go back to development, repeat until deadlines force compromise.</p>
<p>The agentic model: Security is embedded. AI agents continuously validate security throughout development, provide real-time guidance, and scale security expertise to match development velocity.</p>
<p>Security Agent doesn&rsquo;t eliminate the need for security expertise—it amplifies what your security team can accomplish. One senior security engineer using Security Agent can cover more applications than a team of five without it.</p>
<p>The question isn&rsquo;t whether agentic security is coming—it&rsquo;s here. The question is whether you&rsquo;re ready when GA drops.</p>
<p>If you&rsquo;re experimenting with this or have questions, <a href="https://www.linkedin.com/in/lucaslittle/">reach out on LinkedIn</a>. The technology is moving fast, and we&rsquo;re all figuring it out together.</p>
<p><strong>Resources:</strong></p>
<ul>
<li><a href="https://docs.aws.amazon.com/securityagent/latest/userguide/">AWS Security Agent User Guide</a></li>
<li><a href="https://aws.amazon.com/security-agent">AWS Security Agent Product Page</a></li>
<li><a href="https://aws.amazon.com/ai/frontier-agents">AWS Frontier Agents Overview</a></li>
</ul>
]]></content:encoded></item><item><title>Enhancing Security: Adding AWS Cognito Authentication to Your Serverless App</title><link>https://lukelittle.com/posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/</link><pubDate>Thu, 22 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/</guid><description>How we added secure user authentication to our serverless survey app using AWS Cognito</description><content:encoded><![CDATA[<p>Our serverless survey application is a great example of a modern cloud native application. It&rsquo;s fast, scalable, and cost-effective. But it&rsquo;s missing one critical feature: user authentication. In this post, we&rsquo;ll walk through how to add robust, secure authentication using AWS Cognito.</p>
<h2 id="why-add-authentication">Why Add Authentication?</h2>
<p>Right now, anyone can vote, and anyone can reset the entire survey. In a real-world application, we need to control access. Authentication allows us to:</p>
<ul>
<li><strong>Prevent abuse:</strong> Stop users from voting multiple times.</li>
<li><strong>Secure administrative functions:</strong> Only allow authorized users to reset the survey.</li>
<li><strong>Personalize the user experience:</strong> (Future enhancement) Show users their past votes.</li>
</ul>
<h2 id="exploring-the-options">Exploring the Options</h2>
<p>When adding authentication to a serverless app on AWS, there are a few common patterns:</p>
<ol>
<li>
<p><strong>API Keys:</strong> The simplest approach. We could generate an API key and require it for certain API endpoints.</p>
<ul>
<li><strong>Pros:</strong> Easy to implement.</li>
<li><strong>Cons:</strong> Not true user authentication. Keys can be shared or leaked. Doesn&rsquo;t scale for managing individual users.</li>
</ul>
</li>
<li>
<p><strong>Lambda Authorizers (Custom Authorizers):</strong> We can write a custom Lambda function that is triggered by API Gateway before the main handler. This function would be responsible for validating a token (e.g., a session token you manage yourself).</p>
<ul>
<li><strong>Pros:</strong> Full control over the authentication logic.</li>
<li><strong>Cons:</strong> You have to build and manage the entire user lifecycle yourself (sign-up, sign-in, password reset, etc.). This is a lot of work and easy to get wrong.</li>
</ul>
</li>
<li>
<p><strong>AWS Cognito:</strong> A fully managed user identity and authentication service. Cognito handles all the heavy lifting of user management, including sign-up, sign-in, password recovery, and multi-factor authentication (MFA). It integrates seamlessly with API Gateway.</p>
<ul>
<li><strong>Pros:</strong> Secure, scalable, and feature-rich. Offloads the undifferentiated heavy lifting of authentication.</li>
<li><strong>Cons:</strong> Can seem complex at first due to the number of features.</li>
</ul>
</li>
</ol>
<p>For our application, <strong>AWS Cognito is the clear winner.</strong> It provides the best balance of security, features, and ease of integration.</p>
<h2 id="understanding-the-current-architecture">Understanding the Current Architecture</h2>
<p>Before we add authentication, let&rsquo;s visualize how our serverless application currently works:</p>
<p>Right now, anyone can call our API endpoints. There&rsquo;s no way to verify who&rsquo;s making the request or prevent abuse.</p>
<h2 id="the-plan-integrating-cognito">The Plan: Integrating Cognito</h2>
<p><img src="/posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/adding-cognito-to-our-survey-app_hu_d7c3c1099a40640b.webp" srcset="/posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/adding-cognito-to-our-survey-app_hu_4b43f59e39e0c5ea.webp 750w, /posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/adding-cognito-to-our-survey-app_hu_d7c3c1099a40640b.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="925" alt="Enhancing Security: Adding AWS Cognito Authentication to Your Serverless App" loading="lazy" decoding="async"></p>
<p>We&rsquo;ll use a <strong>Cognito User Pool</strong> to manage our users. Think of it as a user database with built-in authentication logic. Here&rsquo;s the high-level plan:</p>
<ol>
<li>
<p><strong>Infrastructure (Terraform):</strong></p>
<ul>
<li>Create a Cognito User Pool to store and manage users.</li>
<li>Create a Cognito User Pool Client, which allows our frontend application to interact with the User Pool.</li>
<li>Add an authorizer to our API Gateway to protect our endpoints.</li>
</ul>
</li>
<li>
<p><strong>Frontend (HTML/JavaScript):</strong></p>
<ul>
<li>Create a login page (<code>login.html</code>).</li>
<li>Add sign-up and sign-in forms.</li>
<li>Use the <a href="https://github.com/aws-amplify/amplify-js/tree/main/packages/auth">Amazon Cognito Identity SDK for JavaScript</a> to communicate with Cognito.</li>
<li>On successful sign-in, store the JWT (JSON Web Token) from Cognito in local storage.</li>
<li>Include the JWT in the <code>Authorization</code> header of all subsequent API requests.</li>
<li>Add a &ldquo;Logout&rdquo; button.</li>
</ul>
</li>
<li>
<p><strong>Backend (Lambda):</strong></p>
<ul>
<li>Update our Lambda functions to expect and validate the JWT from Cognito. API Gateway will do most of the validation for us.</li>
<li>We can use the claims inside the JWT to identify the user. For instance, we can use the user&rsquo;s <code>sub</code> (subject) claim instead of the <code>sessionId</code> to track votes.</li>
</ul>
</li>
</ol>
<p>Here&rsquo;s what the authenticated architecture will look like:</p>
<h2 id="lets-get-building">Let&rsquo;s Get Building!</h2>
<h3 id="step-1-beefing-up-our-terraform">Step 1: Beefing up our Terraform</h3>
<p>First, we need to add the Cognito resources to our <code>terraform/main.tf</code> file.</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-terraform" data-lang="terraform"><span style="display:flex;"><span><span style="color:#75715e"># --- Cognito User Pool ---
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_cognito_user_pool&#34;</span> <span style="color:#e6db74">&#34;survey_user_pool&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;SurveyUserPool&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">auto_verified_attributes</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;email&#34;</span>]
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_cognito_user_pool_client&#34;</span> <span style="color:#e6db74">&#34;survey_user_pool_client&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;SurveyUserPoolClient&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">user_pool_id</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_cognito_user_pool</span>.<span style="color:#a6e22e">survey_user_pool</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">generate_secret</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">false</span><span style="color:#75715e"> # This is a public client
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">explicit_auth_flows</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;ALLOW_USER_PASSWORD_AUTH&#34;</span>, <span style="color:#e6db74">&#34;ALLOW_REFRESH_TOKEN_AUTH&#34;</span>]
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>We also need to update our API Gateway to use a Cognito authorizer.</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-terraform" data-lang="terraform"><span style="display:flex;"><span><span style="color:#75715e"># --- API Gateway ---
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_apigatewayv2_api&#34;</span> <span style="color:#e6db74">&#34;survey_api&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span>          <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;CrackingTheCloudSurveyAPI&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">protocol_type</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;HTTP&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">cors_configuration</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">allow_origins</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;*&#34;</span>]
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">allow_methods</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;POST&#34;</span>, <span style="color:#e6db74">&#34;GET&#34;</span>, <span style="color:#e6db74">&#34;OPTIONS&#34;</span>]
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">allow_headers</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;*&#34;</span>]
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_apigatewayv2_authorizer&#34;</span> <span style="color:#e6db74">&#34;cognito_authorizer&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">api_id</span>           <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_apigatewayv2_api</span>.<span style="color:#a6e22e">survey_api</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">authorizer_type</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;JWT&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">identity_sources</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;</span><span style="color:#960050;background-color:#1e0010">$</span><span style="color:#e6db74">request.header.Authorization&#34;</span>]
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span>             <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;CognitoAuthorizer&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">jwt_configuration</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">audience</span> <span style="color:#f92672">=</span> [<span style="color:#a6e22e">aws_cognito_user_pool_client</span>.<span style="color:#a6e22e">survey_user_pool_client</span>.<span style="color:#a6e22e">id</span>]
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">issuer</span>   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;https://</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">aws_cognito_user_pool</span>.<span style="color:#a6e22e">survey_user_pool</span>.<span style="color:#a6e22e">endpoint</span><span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"># Update the routes to use the authorizer
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_apigatewayv2_route&#34;</span> <span style="color:#e6db74">&#34;vote_route&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">api_id</span>    <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_apigatewayv2_api</span>.<span style="color:#a6e22e">survey_api</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">route_key</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;POST /vote&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">target</span>    <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;integrations/</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">aws_apigatewayv2_integration</span>.<span style="color:#a6e22e">vote_integration</span>.<span style="color:#a6e22e">id</span><span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">authorization_type</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;JWT&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">authorizer_id</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_apigatewayv2_authorizer</span>.<span style="color:#a6e22e">cognito_authorizer</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_apigatewayv2_route&#34;</span> <span style="color:#e6db74">&#34;reset_route&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">api_id</span>    <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_apigatewayv2_api</span>.<span style="color:#a6e22e">survey_api</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">route_key</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;POST /reset&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">target</span>    <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;integrations/</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">aws_apigatewayv2_integration</span>.<span style="color:#a6e22e">reset_integration</span>.<span style="color:#a6e22e">id</span><span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # Note: we might want more fine-grained control here in a real app
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#a6e22e">authorization_type</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;JWT&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">authorizer_id</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_apigatewayv2_authorizer</span>.<span style="color:#a6e22e">cognito_authorizer</span>.<span style="color:#a6e22e">id</span>
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><h3 id="step-2-the-frontend---where-the-magic-happens">Step 2: The Frontend - Where the Magic Happens</h3>
<p>This is where we&rsquo;ll see the biggest changes. We need a way for users to sign up and sign in.</p>
<p>First, let&rsquo;s create a new <code>login.html</code> page. This will be a simple page with forms for sign-up and sign-in.</p>
<p>We&rsquo;ll also need to include the AWS Cognito Identity SDK in our project. We can either download it and host it ourselves, or use a CDN. For simplicity, we&rsquo;ll use a CDN in our HTML files.</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-html" data-lang="html"><span style="display:flex;"><span><span style="color:#75715e">&lt;!-- In login.html, index.html, etc. --&gt;</span>
</span></span><span style="display:flex;"><span>&lt;<span style="color:#f92672">script</span> <span style="color:#a6e22e">src</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;https://cdn.jsdelivr.net/npm/amazon-cognito-identity-js@5.2.10/dist/amazon-cognito-identity.min.js&#34;</span>&gt;&lt;/<span style="color:#f92672">script</span>&gt;
</span></span></code></pre></div><p>Our <code>main.js</code> will need a significant update. Here are the key parts:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-javascript" data-lang="javascript"><span style="display:flex;"><span><span style="color:#75715e">// Add these variables at the top of main.js
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">API_URL</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;${API_URL}&#39;</span>;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">COGNITO_USER_POOL_ID</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;${COGNITO_USER_POOL_ID}&#39;</span>;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">COGNITO_CLIENT_ID</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;${COGNITO_CLIENT_ID}&#39;</span>;
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">poolData</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">UserPoolId</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">COGNITO_USER_POOL_ID</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">ClientId</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">COGNITO_CLIENT_ID</span>
</span></span><span style="display:flex;"><span>};
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">userPool</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">CognitoUserPool</span>(<span style="color:#a6e22e">poolData</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e">// On page load, check if the user is logged in
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>window.<span style="color:#a6e22e">onload</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">function</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">idToken</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">localStorage</span>.<span style="color:#a6e22e">getItem</span>(<span style="color:#e6db74">&#39;idToken&#39;</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> (<span style="color:#f92672">!</span><span style="color:#a6e22e">idToken</span> <span style="color:#f92672">&amp;&amp;</span> <span style="color:#f92672">!</span>window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span>.<span style="color:#a6e22e">endsWith</span>(<span style="color:#e6db74">&#39;login.html&#39;</span>)) {
</span></span><span style="display:flex;"><span>        window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;login.html&#39;</span>;
</span></span><span style="display:flex;"><span>    } <span style="color:#66d9ef">else</span> <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">idToken</span> <span style="color:#f92672">&amp;&amp;</span> window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span>.<span style="color:#a6e22e">endsWith</span>(<span style="color:#e6db74">&#39;index.html&#39;</span>)) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">addSignOutButton</span>();
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>};
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">addSignOutButton</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">container</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signout-container&#39;</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">container</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">signOutButton</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">createElement</span>(<span style="color:#e6db74">&#39;button&#39;</span>);
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">signOutButton</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Sign Out&#39;</span>;
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">signOutButton</span>.<span style="color:#a6e22e">onclick</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">signOutUser</span>;
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">container</span>.<span style="color:#a6e22e">appendChild</span>(<span style="color:#a6e22e">signOutButton</span>);
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">signUpUser</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">email</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signUpEmail&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">password</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signUpPassword&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">messageDiv</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;message&#39;</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">attributeList</span> <span style="color:#f92672">=</span> [];
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">dataEmail</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Name</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;email&#39;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Value</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">email</span>,
</span></span><span style="display:flex;"><span>    };
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">attributeEmail</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">CognitoUserAttribute</span>(<span style="color:#a6e22e">dataEmail</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">attributeList</span>.<span style="color:#a6e22e">push</span>(<span style="color:#a6e22e">attributeEmail</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">userPool</span>.<span style="color:#a6e22e">signUp</span>(<span style="color:#a6e22e">email</span>, <span style="color:#a6e22e">password</span>, <span style="color:#a6e22e">attributeList</span>, <span style="color:#66d9ef">null</span>, <span style="color:#66d9ef">function</span>(<span style="color:#a6e22e">err</span>, <span style="color:#a6e22e">result</span>){
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">err</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">message</span> <span style="color:#f92672">||</span> <span style="color:#a6e22e">JSON</span>.<span style="color:#a6e22e">stringify</span>(<span style="color:#a6e22e">err</span>);
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span>;
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Sign up successful! Please check your email for a verification code.&#39;</span>;
</span></span><span style="display:flex;"><span>        document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;confirm-container&#39;</span>).<span style="color:#a6e22e">style</span>.<span style="color:#a6e22e">display</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;block&#39;</span>;
</span></span><span style="display:flex;"><span>    });
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">confirmSignUpUser</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">email</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signUpEmail&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">code</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;confirmationCode&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">messageDiv</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;message&#39;</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">userData</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Username</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">email</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Pool</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">userPool</span>
</span></span><span style="display:flex;"><span>    };
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">cognitoUser</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">CognitoUser</span>(<span style="color:#a6e22e">userData</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">cognitoUser</span>.<span style="color:#a6e22e">confirmRegistration</span>(<span style="color:#a6e22e">code</span>, <span style="color:#66d9ef">true</span>, <span style="color:#66d9ef">function</span>(<span style="color:#a6e22e">err</span>, <span style="color:#a6e22e">result</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">err</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">message</span> <span style="color:#f92672">||</span> <span style="color:#a6e22e">JSON</span>.<span style="color:#a6e22e">stringify</span>(<span style="color:#a6e22e">err</span>);
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span>;
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Confirmation successful! You can now sign in.&#39;</span>;
</span></span><span style="display:flex;"><span>        document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;confirm-container&#39;</span>).<span style="color:#a6e22e">style</span>.<span style="color:#a6e22e">display</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;none&#39;</span>;
</span></span><span style="display:flex;"><span>    });
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">signInUser</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">email</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signInEmail&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">password</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;signInPassword&#39;</span>).<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">messageDiv</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;message&#39;</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">authenticationData</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Username</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">email</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Password</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">password</span>,
</span></span><span style="display:flex;"><span>    };
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">authenticationDetails</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">AuthenticationDetails</span>(<span style="color:#a6e22e">authenticationData</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">userData</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Username</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">email</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">Pool</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">userPool</span>
</span></span><span style="display:flex;"><span>    };
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">cognitoUser</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">CognitoUser</span>(<span style="color:#a6e22e">userData</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">cognitoUser</span>.<span style="color:#a6e22e">authenticateUser</span>(<span style="color:#a6e22e">authenticationDetails</span>, {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">onSuccess</span><span style="color:#f92672">:</span> <span style="color:#66d9ef">function</span> (<span style="color:#a6e22e">result</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">idToken</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">result</span>.<span style="color:#a6e22e">getIdToken</span>().<span style="color:#a6e22e">getJwtToken</span>();
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">localStorage</span>.<span style="color:#a6e22e">setItem</span>(<span style="color:#e6db74">&#39;idToken&#39;</span>, <span style="color:#a6e22e">idToken</span>);
</span></span><span style="display:flex;"><span>            window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;index.html&#39;</span>;
</span></span><span style="display:flex;"><span>        },
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">onFailure</span><span style="color:#f92672">:</span> <span style="color:#66d9ef">function</span>(<span style="color:#a6e22e">err</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">message</span> <span style="color:#f92672">||</span> <span style="color:#a6e22e">JSON</span>.<span style="color:#a6e22e">stringify</span>(<span style="color:#a6e22e">err</span>);
</span></span><span style="display:flex;"><span>        },
</span></span><span style="display:flex;"><span>    });
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">signOutUser</span>() {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">localStorage</span>.<span style="color:#a6e22e">removeItem</span>(<span style="color:#e6db74">&#39;idToken&#39;</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">cognitoUser</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">userPool</span>.<span style="color:#a6e22e">getCurrentUser</span>();
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">cognitoUser</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">cognitoUser</span>.<span style="color:#a6e22e">signOut</span>();
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;login.html&#39;</span>;
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">function</span> <span style="color:#a6e22e">vote</span>(<span style="color:#a6e22e">option</span>) {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">idToken</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">localStorage</span>.<span style="color:#a6e22e">getItem</span>(<span style="color:#e6db74">&#39;idToken&#39;</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> (<span style="color:#f92672">!</span><span style="color:#a6e22e">idToken</span>) {
</span></span><span style="display:flex;"><span>        window.<span style="color:#a6e22e">location</span>.<span style="color:#a6e22e">href</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;login.html&#39;</span>;
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span>;
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">messageDiv</span> <span style="color:#f92672">=</span> document.<span style="color:#a6e22e">getElementById</span>(<span style="color:#e6db74">&#39;message&#39;</span>);
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Submitting your vote...&#39;</span>;
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">response</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">await</span> <span style="color:#a6e22e">fetch</span>(<span style="color:#e6db74">`</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">API_URL</span><span style="color:#e6db74">}</span><span style="color:#e6db74">vote`</span>, {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">method</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;POST&#39;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">headers</span><span style="color:#f92672">:</span> {
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;Content-Type&#39;</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;application/json&#39;</span>,
</span></span><span style="display:flex;"><span>                <span style="color:#e6db74">&#39;Authorization&#39;</span><span style="color:#f92672">:</span> <span style="color:#e6db74">`Bearer </span><span style="color:#e6db74">${</span><span style="color:#a6e22e">idToken</span><span style="color:#e6db74">}</span><span style="color:#e6db74">`</span>
</span></span><span style="display:flex;"><span>            },
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">body</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">JSON</span>.<span style="color:#a6e22e">stringify</span>({ <span style="color:#a6e22e">vote</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">option</span> }),
</span></span><span style="display:flex;"><span>        });
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> (<span style="color:#f92672">!</span><span style="color:#a6e22e">response</span>.<span style="color:#a6e22e">ok</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">throw</span> <span style="color:#66d9ef">new</span> Error(<span style="color:#e6db74">`HTTP error! status: $</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">response</span>.<span style="color:#a6e22e">status</span><span style="color:#e6db74">}</span><span style="color:#e6db74">`</span>);
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Thank you for voting!&#39;</span>;
</span></span><span style="display:flex;"><span>        document.<span style="color:#a6e22e">querySelectorAll</span>(<span style="color:#e6db74">&#39;.vote-btn&#39;</span>).<span style="color:#a6e22e">forEach</span>(<span style="color:#a6e22e">button</span> =&gt; {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">button</span>.<span style="color:#a6e22e">disabled</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>;
</span></span><span style="display:flex;"><span>        });
</span></span><span style="display:flex;"><span>        
</span></span><span style="display:flex;"><span>    } <span style="color:#66d9ef">catch</span> (<span style="color:#a6e22e">error</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">console</span>.<span style="color:#a6e22e">error</span>(<span style="color:#e6db74">&#39;Error submitting vote:&#39;</span>, <span style="color:#a6e22e">error</span>);
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">messageDiv</span>.<span style="color:#a6e22e">textContent</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#39;Sorry, there was an error submitting your vote.&#39;</span>;
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><h3 id="step-3-backend-adjustments">Step 3: Backend Adjustments</h3>
<p>Our backend Lambda functions need a small change. Since we are now identifying users by their JWT, we can remove the <code>sessionId</code> logic. The <code>vote.py</code> function will now get the user&rsquo;s unique ID from the JWT claims that API Gateway passes along.</p>
<p>Here&rsquo;s how the Lambda function processes an authenticated request:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#75715e"># In backend/vote.py</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span>:
</span></span><span style="display:flex;"><span>        body <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;body&#39;</span>, <span style="color:#e6db74">&#39;</span><span style="color:#e6db74">{}</span><span style="color:#e6db74">&#39;</span>))
</span></span><span style="display:flex;"><span>        vote_option <span style="color:#f92672">=</span> body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;vote&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Get user id from the authorizer context</span>
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># API Gateway extracts this from the JWT and passes it to Lambda</span>
</span></span><span style="display:flex;"><span>        user_id <span style="color:#f92672">=</span> event[<span style="color:#e6db74">&#39;requestContext&#39;</span>][<span style="color:#e6db74">&#39;authorizer&#39;</span>][<span style="color:#e6db74">&#39;jwt&#39;</span>][<span style="color:#e6db74">&#39;claims&#39;</span>][<span style="color:#e6db74">&#39;sub&#39;</span>]
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> <span style="color:#f92672">not</span> user_id:
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">400</span>, <span style="color:#f92672">...</span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> vote_option <span style="color:#f92672">not</span> <span style="color:#f92672">in</span> [<span style="color:#e6db74">&#39;no&#39;</span>, <span style="color:#e6db74">&#39;aws&#39;</span>, <span style="color:#e6db74">&#39;other&#39;</span>]:
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">400</span>, <span style="color:#f92672">...</span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        item <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;id&#39;</span>: user_id,      <span style="color:#75715e"># Use the cognito user id as the primary key</span>
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;vote&#39;</span>: vote_option
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>        table<span style="color:#f92672">.</span>put_item(Item<span style="color:#f92672">=</span>item)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {<span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>, <span style="color:#f92672">...</span>}
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">except</span> <span style="color:#a6e22e">Exception</span> <span style="color:#66d9ef">as</span> e:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># ...</span>
</span></span></code></pre></div><p><strong>What&rsquo;s happening here?</strong></p>
<ul>
<li>The <code>sub</code> (subject) claim in the JWT is a unique identifier for each Cognito user</li>
<li>API Gateway validates the JWT before the Lambda even runs</li>
<li>If the JWT is invalid or missing, the request never reaches our Lambda</li>
<li>We use the user&rsquo;s <code>sub</code> as the primary key in DynamoDB, ensuring one vote per user</li>
</ul>
<h2 id="security-is-not-a-feature-its-a-foundation">Security is Not a Feature, It&rsquo;s a Foundation</h2>
<p>Let&rsquo;s talk about the security improvements we&rsquo;ve made:</p>
<ul>
<li><strong>Managed User Directory:</strong> We are not storing passwords. Cognito handles all password policies, hashing, and storage, following best practices.</li>
<li><strong>JWT Authentication:</strong> We&rsquo;re using the industry standard for API authentication. The JWTs are signed by Cognito, and our API Gateway verifies this signature on every request. This prevents token tampering.</li>
<li><strong>Secure Token Storage:</strong> We are storing the JWT in <code>localStorage</code>. This is a common practice, but it&rsquo;s important to be aware of the risks (like XSS attacks).  For higher security applications, we could store tokens in memory and use refresh tokens to get new access tokens.</li>
<li><strong>HTTPS Everywhere:</strong> Our entire application, from the frontend on CloudFront to the API Gateway, enforces HTTPS. This prevents eavesdropping.</li>
<li><strong>Least Privilege:</strong> Our Lambda functions have fine-grained IAM roles. They can only access the DynamoDB table they need.</li>
</ul>
<h2 id="understanding-our-lambda-functions">Understanding Our Lambda Functions</h2>
<p>Let&rsquo;s break down what each Lambda function does in our application:</p>
<h3 id="vote-lambda-function">Vote Lambda Function</h3>
<p>This function processes vote submissions from authenticated users.</p>
<p><strong>Key responsibilities:</strong></p>
<ul>
<li>Extract the vote option from the request body</li>
<li>Get the user&rsquo;s unique ID from the JWT claims (provided by API Gateway)</li>
<li>Validate the vote is one of the allowed options</li>
<li>Store the vote in DynamoDB using the user ID as the key (prevents duplicate votes)</li>
</ul>
<h3 id="results-lambda-function">Results Lambda Function</h3>
<p>This function retrieves and counts all votes. No authentication required - results are public!</p>
<p><strong>Key responsibilities:</strong></p>
<ul>
<li>Scan the entire DynamoDB table to get all votes</li>
<li>Handle pagination (DynamoDB returns max 1MB per request)</li>
<li>Count votes for each option using Python&rsquo;s <code>Counter</code></li>
<li>Return totals as JSON: <code>{&quot;no&quot;: 5, &quot;aws&quot;: 12, &quot;other&quot;: 3}</code></li>
</ul>
<h3 id="reset-lambda-function">Reset Lambda Function</h3>
<p>This function deletes all votes. Requires authentication to prevent abuse.</p>
<p><strong>Key responsibilities:</strong></p>
<ul>
<li>Scan the table to get all item IDs</li>
<li>Use batch writer to efficiently delete items (groups of 25)</li>
<li>Handle pagination for large datasets</li>
<li>Return success confirmation</li>
</ul>
<h3 id="email-filter-lambda-function">Email Filter Lambda Function</h3>
<p>This is a special type of Lambda called a <strong>Cognito Trigger</strong>. It runs automatically before a user signs up.</p>
<p><strong>Key responsibilities:</strong></p>
<ul>
<li>Extract the email from the sign-up request</li>
<li>Check if it ends with <code>@charlotte.edu</code></li>
<li>Allow or reject the sign-up based on the domain</li>
<li>This enforces organization-level access control</li>
</ul>
<h2 id="the-final-result">The Final Result</h2>
<p>After implementing these changes, our complete authentication flow looks like this:</p>
<p>We&rsquo;ve now added a robust and secure authentication layer to our serverless application, moving it from a simple demo to a more production-ready state. This is the power of leveraging managed services like AWS Cognito!</p>
<h2 id="bonus-restricting-sign-ups-to-a-specific-email-domain">Bonus: Restricting Sign-ups to a Specific Email Domain</h2>
<p>A common requirement is to restrict application access to users from a specific organization. We can easily extend our Cognito setup to only allow sign-ups from users with a <code>@charlotte.edu</code> email address.</p>
<p>This is accomplished using a Cognito <strong>Pre Sign-up Lambda Trigger</strong>. This trigger fires just before Cognito creates a new user, giving us a chance to run custom validation logic.</p>
<p><strong>What&rsquo;s a Lambda Trigger?</strong>
Think of it like a hook or event listener. Cognito has several points in the user lifecycle where it can automatically invoke a Lambda function:</p>
<ul>
<li><strong>Pre Sign-up</strong>: Before creating a new user (we use this one!)</li>
<li><strong>Post Confirmation</strong>: After a user confirms their email</li>
<li><strong>Pre Authentication</strong>: Before signing in</li>
<li><strong>Post Authentication</strong>: After signing in</li>
</ul>
<p>These triggers let you customize Cognito&rsquo;s behavior without modifying AWS&rsquo;s code.</p>
<h3 id="step-1-create-the-email-filter-lambda">Step 1: Create the Email Filter Lambda</h3>
<p>We&rsquo;ll create a new Lambda function in <code>backend/email-filter.py</code>:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> json
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># This trigger is invoked before a user is signed up</span>
</span></span><span style="display:flex;"><span>    email <span style="color:#f92672">=</span> event[<span style="color:#e6db74">&#39;request&#39;</span>][<span style="color:#e6db74">&#39;userAttributes&#39;</span>]<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;email&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> email <span style="color:#f92672">and</span> email<span style="color:#f92672">.</span>endswith(<span style="color:#e6db74">&#39;@charlotte.edu&#39;</span>):
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Allow sign-up</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> event
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">else</span>:
</span></span><span style="display:flex;"><span>        <span style="color:#75715e"># Block sign-up</span>
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">Exception</span>(<span style="color:#e6db74">&#34;Only users with a @charlotte.edu email address are allowed to sign up.&#34;</span>)
</span></span></code></pre></div><h3 id="step-2-update-terraform">Step 2: Update Terraform</h3>
<p>Now, we&rsquo;ll update our <code>terraform/main.tf</code> to create this new Lambda and associate it with our User Pool.</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-terraform" data-lang="terraform"><span style="display:flex;"><span><span style="color:#66d9ef">data</span> <span style="color:#e6db74">&#34;archive_file&#34;</span> <span style="color:#e6db74">&#34;email_filter_zip&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">type</span>        <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;zip&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">source_file</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;../backend/email-filter.py&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">output_path</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;email-filter.zip&#34;</span>
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_lambda_function&#34;</span> <span style="color:#e6db74">&#34;email_filter_function&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">function_name</span> <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;CrackingTheCloudEmailFilter&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">role</span>          <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_iam_role</span>.<span style="color:#a6e22e">lambda_exec_role</span>.<span style="color:#a6e22e">arn</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">handler</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;email-filter.handler&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">runtime</span>       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;python3.9&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">filename</span>      <span style="color:#f92672">=</span> data.<span style="color:#a6e22e">archive_file</span>.<span style="color:#a6e22e">email_filter_zip</span>.<span style="color:#a6e22e">output_path</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">source_code_hash</span> <span style="color:#f92672">=</span> data.<span style="color:#a6e22e">archive_file</span>.<span style="color:#a6e22e">email_filter_zip</span>.<span style="color:#a6e22e">output_base64sha256</span>
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_cognito_user_pool&#34;</span> <span style="color:#e6db74">&#34;survey_user_pool&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">name</span>                     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;SurveyUserPool&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">auto_verified_attributes</span> <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#34;email&#34;</span>]
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">lambda_config</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">pre_sign_up</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_lambda_function</span>.<span style="color:#a6e22e">email_filter_function</span>.<span style="color:#a6e22e">arn</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_lambda_permission&#34;</span> <span style="color:#e6db74">&#34;cognito_permission_email_filter&#34;</span> {
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">statement_id</span>  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;AllowCognitoToInvokeEmailFilter&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">action</span>        <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;lambda:InvokeFunction&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">function_name</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_lambda_function</span>.<span style="color:#a6e22e">email_filter_function</span>.<span style="color:#a6e22e">function_name</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">principal</span>     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;cognito-idp.amazonaws.com&#34;</span>
</span></span><span style="display:flex;"><span>  <span style="color:#a6e22e">source_arn</span>    <span style="color:#f92672">=</span> <span style="color:#a6e22e">aws_cognito_user_pool</span>.<span style="color:#a6e22e">survey_user_pool</span>.<span style="color:#a6e22e">arn</span>
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>With these changes, any attempt to sign up with an email address that does not end in <code>@charlotte.edu</code> will be rejected by Cognito. This is a powerful way to enforce organization-specific access control.</p>
<h2 id="key-takeaways-for-students">Key Takeaways for Students</h2>
<p><strong>What you&rsquo;ve learned:</strong></p>
<ol>
<li><strong>Serverless Authentication</strong>: No need to build your own user management system from scratch</li>
<li><strong>JWT Tokens</strong>: Industry-standard way to authenticate API requests</li>
<li><strong>Lambda Triggers</strong>: Extend AWS services with custom logic at specific lifecycle events</li>
<li><strong>API Gateway Authorizers</strong>: Validate tokens before requests reach your backend code</li>
<li><strong>Infrastructure as Code</strong>: All of this is defined in Terraform and can be deployed in minutes</li>
</ol>
<p><strong>Real-world applications:</strong></p>
<ul>
<li>Student portals restricted to university email domains</li>
<li>Internal company tools that only employees can access</li>
<li>Multi-tenant SaaS applications where each organization has isolated access</li>
<li>Mobile apps that need secure backend APIs</li>
</ul>
<p><strong>Cost considerations:</strong></p>
<ul>
<li>Cognito: First 50,000 monthly active users are free</li>
<li>Lambda: First 1 million requests per month are free</li>
<li>DynamoDB: 25 GB of storage free</li>
<li>API Gateway: First 1 million API calls per month are free</li>
</ul>
<p>For a student project or small application, this entire stack runs essentially for free!</p>
<h2 id="next-steps">Next Steps</h2>
<p>Want to take this further? Here are some ideas:</p>
<ol>
<li><strong>Add user profiles</strong>: Store additional user data in DynamoDB</li>
<li><strong>Implement admin roles</strong>: Use Cognito groups to create admin users with special permissions</li>
<li><strong>Add MFA</strong>: Enable multi-factor authentication for extra security</li>
<li><strong>Social sign-in</strong>: Allow users to sign in with Google, Facebook, or other providers</li>
<li><strong>Password reset flow</strong>: Implement &ldquo;forgot password&rdquo; functionality</li>
<li><strong>Email customization</strong>: Customize the verification emails Cognito sends</li>
</ol>
<p>The foundation you&rsquo;ve built here is production-ready and can scale to millions of users!</p>
<p>GitHub Repository: <a href="https://github.com/lukelittle/adding-cognito-to-our-survey-app">https://github.com/lukelittle/adding-cognito-to-our-survey-app</a></p>
]]></content:encoded></item><item><title>15 Hours of Terraform in 3: Building with AWS Kiro</title><link>https://lukelittle.com/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/</link><pubDate>Tue, 20 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/</guid><description>Building a production-ready URL shortener with Kiro: spec-driven development, smart architecture decisions, and why AI coding assistants are game-changers</description><content:encoded><![CDATA[<p>At an AWS Road Show this fall, Darko Mesaros demoed a URL shortener he&rsquo;d built in Rust called <a href="https://github.com/darko-mesaros/krtk.rs">krtk.rs</a>. Something about watching a clean, fast URL shortener just <em>work</em> stuck with me. I&rsquo;ve built a few of these for demos since then, but I wanted to try something different this time: build one in Python with a retro 90s vibe, and let Kiro handle most of the heavy lifting.</p>
<p>Kiro is one of AWS&rsquo;s three frontier agents announced at re:Invent 2025—autonomous AI systems that maintain context and work independently for hours. While DevOps Agent handles incident response and Security Agent conducts penetration testing, Kiro is your AI developer that takes specifications and generates production-ready code.</p>
<p>This turned into a great experiment in spec-driven AI development. Here&rsquo;s what I learned about building with AI coding assistants.</p>
<h2 id="the-approach-spec-driven-development-with-kiro">The approach: Spec-driven development with Kiro</h2>
<p>I started by asking ChatGPT to help me write a comprehensive prompt for Kiro. The key insight: the better your specification, the better the AI output.</p>
<p>Here&rsquo;s what I asked for:</p>
<blockquote>
<p>You are a senior AWS serverless engineer. Take my existing project and refactor/extend it into a URL shortener with a retro 90s website UI.</p>
</blockquote>
<p>Then I got detailed with requirements:</p>
<ul>
<li>Frontend: A simple 90s-style static site where authenticated users can submit URLs and get short links back. Show their created links and click counts.</li>
<li>Auth: Amazon Cognito for create/delete/list operations. Public redirects don&rsquo;t need auth.</li>
<li>Data: DynamoDB table with <code>link</code> as partition key, plus <code>url</code>, <code>count</code>, <code>creator</code>, and <code>created</code> timestamp.</li>
<li>APIs: POST <code>/links</code> (create), DELETE <code>/links/{link}</code> (delete), GET <code>/links</code> (list), GET <code>/{link}</code> (public redirect)</li>
<li>Lambdas in Python: create_link, delete_link, list_links, redirect, plus analytics functions</li>
<li>Visit counting had to be decoupled: redirect fires an event to Kinesis, consumer Lambda batches events and updates DynamoDB counts atomically</li>
</ul>
<p>The decoupled visit counting was critical—I&rsquo;ve seen too many URL shorteners block redirects waiting for analytics writes. That&rsquo;s how you turn a 50ms redirect into a 150ms redirect.</p>
<h2 id="what-kiro-generated">What Kiro generated</h2>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/building-serverless-url-shortener-ai-assisted-kiro_hu_37feb7a46bfedf21.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/building-serverless-url-shortener-ai-assisted-kiro_hu_f2902b2bcf7808f4.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/building-serverless-url-shortener-ai-assisted-kiro_hu_37feb7a46bfedf21.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="1020" alt="15 Hours of Terraform in 3: Building with AWS Kiro" loading="lazy" decoding="async"></p>
<p>Kiro produced a comprehensive 15-section specification document covering functional requirements, API contracts, data models, security, and cost analysis. This spec became the foundation for everything else.</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-specification_hu_5a730904650ba966.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-specification_hu_c4d26c519ed844b9.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-specification_hu_5a730904650ba966.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="974" alt="Kiro&rsquo;s specification document" loading="lazy" decoding="async"></p>
<p>Then it translated that into concrete design decisions:</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-design-document_hu_aa00331bdeceb8c0.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-design-document_hu_becd8bb0926323d4.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-design-document_hu_aa00331bdeceb8c0.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="974" alt="Design document output" loading="lazy" decoding="async"></p>
<p>And broke the implementation into actionable tasks:</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-task-breakdown_hu_3d36241ec4a4f56c.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-task-breakdown_hu_28b90a52281f4235.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/kiro-task-breakdown_hu_3d36241ec4a4f56c.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="974" alt="Implementation task list" loading="lazy" decoding="async"></p>
<p>Finally, it generated the complete architecture:</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/url-shortener-architecture-diagram_hu_95b1771a544bdf2b.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/url-shortener-architecture-diagram_hu_52aa9f950cf2d952.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/url-shortener-architecture-diagram_hu_95b1771a544bdf2b.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="1113" alt="URL Shortener Architecture Diagram" loading="lazy" decoding="async"></p>
<p>This is production-grade serverless: S3 + CloudFront for the frontend, API Gateway for the REST API, Lambda for compute, DynamoDB for storage, Kinesis for event streaming.</p>
<p>The visit counting architecture is particularly clever:</p>
<ol>
<li>User hits a short link → redirect Lambda looks it up in DynamoDB</li>
<li>Lambda immediately returns a 301 redirect (fast!)</li>
<li><em>Then</em> it fires an event to Kinesis (fire-and-forget, non-blocking)</li>
<li>Kinesis batches up events</li>
<li>A consumer Lambda processes batches and updates DynamoDB counts atomically</li>
</ol>
<p>This approach delivers redirects under 200ms and reduces DynamoDB writes by 100x. For a link getting 1,000 clicks/minute, that&rsquo;s the difference between $75/month and $0.75/month just for counting.</p>
<h2 id="understanding-dynamodb-through-ai">Understanding DynamoDB through AI</h2>
<p>One of the valuable learning moments came when I asked Kiro about the DynamoDB schema:</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/dynamodb-schema-question_hu_a8fb4f666da05b02.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/dynamodb-schema-question_hu_ecddba7a53b585c5.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/dynamodb-schema-question_hu_a8fb4f666da05b02.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="974" alt="Asking about the DynamoDB schema" loading="lazy" decoding="async"></p>
<p>Kiro explained that DynamoDB only requires key attributes defined in Terraform—non-key attributes like <code>count</code> and <code>url</code> are dynamic. The GSI projection (<code>projection_type = &quot;ALL&quot;</code>) automatically includes everything. This kind of on-demand explanation is where AI assistants really shine.</p>
<h2 id="the-iterative-refinement-process">The iterative refinement process</h2>
<p>During deployment, I discovered Kiro had generated all the Lambda functions, DynamoDB tables, and API Gateway configs—but initially missed the S3 bucket and CloudFront distribution for the frontend.</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/forgot-s3-cloudfront_hu_19314c09d4527d3b.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/forgot-s3-cloudfront_hu_9b9d98679654ec0a.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/forgot-s3-cloudfront_hu_19314c09d4527d3b.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="973" alt="Kiro forgot S3 and CloudFront" loading="lazy" decoding="async"></p>
<p>This is normal in iterative development. I pointed it out, and Kiro immediately generated:</p>
<ul>
<li>S3 bucket with proper access controls</li>
<li>CloudFront distribution with Origin Access Control</li>
<li>Bucket policies</li>
<li>Output values for deployment</li>
</ul>
<p>This iterative feedback loop is exactly how development works—AI or not.</p>
<h2 id="configuration-refinements">Configuration refinements</h2>
<p>The first deployment revealed Kiro had configured CloudFront with an <code>/api/*</code> cache behavior that forwarded headers to S3. S3 doesn&rsquo;t support that pattern because:</p>
<ul>
<li>S3 serves static files (HTML, CSS, JS)</li>
<li>API Gateway serves API endpoints (<code>/links</code>, <code>/health</code>, etc.)</li>
<li>CloudFront should only cache static content</li>
</ul>
<p>Removing that behavior simplified the architecture and everything worked perfectly. AI-generated code can sometimes include extra patterns—simplifying is part of the refinement process.</p>
<h2 id="integration-patterns">Integration patterns</h2>
<p>The frontend initially used simulated JWT tokens while the backend had real Cognito validation configured. This kind of integration gap is common when components are generated separately.</p>
<p>Adding proper Cognito authentication was straightforward:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-javascript" data-lang="javascript"><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">poolData</span> <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">UserPoolId</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">COGNITO_USER_POOL_ID</span>,
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">ClientId</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">COGNITO_CLIENT_ID</span>
</span></span><span style="display:flex;"><span>};
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">userPool</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">new</span> <span style="color:#a6e22e">AmazonCognitoIdentity</span>.<span style="color:#a6e22e">CognitoUserPool</span>(<span style="color:#a6e22e">poolData</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#a6e22e">cognitoUser</span>.<span style="color:#a6e22e">authenticateUser</span>(<span style="color:#a6e22e">authenticationDetails</span>, {
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">onSuccess</span><span style="color:#f92672">:</span> <span style="color:#66d9ef">function</span> (<span style="color:#a6e22e">result</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">authToken</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">result</span>.<span style="color:#a6e22e">getIdToken</span>().<span style="color:#a6e22e">getJwtToken</span>();
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">showApp</span>();
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">loadUserLinks</span>();
</span></span><span style="display:flex;"><span>    },
</span></span><span style="display:flex;"><span>    <span style="color:#a6e22e">onFailure</span><span style="color:#f92672">:</span> <span style="color:#66d9ef">function</span>(<span style="color:#a6e22e">err</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">code</span> <span style="color:#f92672">===</span> <span style="color:#e6db74">&#39;UserNotConfirmedException&#39;</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">showError</span>(<span style="color:#e6db74">&#39;Check your email to confirm your account.&#39;</span>);
</span></span><span style="display:flex;"><span>        } <span style="color:#66d9ef">else</span> <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">code</span> <span style="color:#f92672">===</span> <span style="color:#e6db74">&#39;NotAuthorizedException&#39;</span>) {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">showError</span>(<span style="color:#e6db74">&#39;Invalid email or password.&#39;</span>);
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>});
</span></span></code></pre></div><p>Integration complete. System working end-to-end.</p>
<h2 id="what-worked-exceptionally-well">What worked exceptionally well</h2>
<p>Kiro generated production-quality code across multiple areas:</p>
<p><strong>Kinesis batching:</strong> The visit counting pipeline processed events exactly as designed, delivering that 100x cost reduction.</p>
<p><strong>DynamoDB schema:</strong> Properly implemented with dynamic non-key attributes and a well-designed GSI for querying by creator.</p>
<p><strong>Terraform structure:</strong> Clean, modular IaC with proper IAM roles using least-privilege permissions and good tagging for cost allocation.</p>
<p><strong>Lambda functions:</strong> All eight Python functions came with error handling, structured logging, and CloudWatch metrics built in.</p>
<p><strong>Cost modeling:</strong> Kiro accurately estimated ~$15/month for 100K redirects. For comparison, running this on containers would cost 10x more.</p>
<p>I even asked Kiro if the system was production-ready:</p>
<p><img src="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/production-ready-question_hu_68a8fdd16c67b9c6.webp" srcset="/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/production-ready-question_hu_5ab008686194c325.webp 750w, /posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/production-ready-question_hu_68a8fdd16c67b9c6.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="975" alt="Is it ready for production?" loading="lazy" decoding="async"></p>
<p>It correctly identified what was complete and what still needed work.</p>
<h2 id="why-spec-driven-ai-development-works">Why spec-driven AI development works</h2>
<p>This project demonstrated several key principles:</p>
<p><strong>Clear specifications produce better results:</strong> The comprehensive prompt I gave Kiro led to comprehensive, well-architected output. Garbage in, garbage out—quality specifications get quality code.</p>
<p><strong>AI excels at boilerplate:</strong> Kiro generated thousands of lines of Terraform, Lambda functions, and infrastructure configurations that would&rsquo;ve taken me 10-15 hours to write manually. This is where AI delivers massive productivity gains.</p>
<p><strong>Iteration is normal:</strong> Whether working with AI or human developers, you iterate. Point out gaps, refine implementations, simplify where needed. The feedback loop is fast with AI.</p>
<p><strong>Architecture decisions matter:</strong> The Kinesis batching decision came from my prompt. AI implemented it perfectly. The human provides the &ldquo;why,&rdquo; the AI handles the &ldquo;how.&rdquo;</p>
<p><strong>Learning opportunity:</strong> Using Kiro taught me DynamoDB patterns I hadn&rsquo;t used before. AI assistants can explain concepts while implementing them.</p>
<h2 id="the-final-system">The final system</h2>
<p>Everything deployed and works in production:</p>
<ul>
<li>S3 + CloudFront hosting the 90s-themed frontend</li>
<li>API Gateway + Lambda handling authenticated operations</li>
<li>Cognito managing users via backend endpoints</li>
<li>DynamoDB storing links with a creator-created GSI for queries</li>
<li>Kinesis streaming visit events for batch processing</li>
<li>CloudWatch logs, metrics, and alarms for observability</li>
</ul>
<p>Total cost: ~$15/month for 100K redirects. The code is open source: <a href="https://github.com/lukelittle/url-shortener">github.com/lukelittle/url-shortener</a></p>
<h2 id="the-value-of-ai-coding-assistants">The value of AI coding assistants</h2>
<p>Working with Kiro on this project showed me why AI coding assistants are becoming essential:</p>
<p><strong>10-15 hours of coding compressed into 2-3 hours of specification and refinement.</strong> That&rsquo;s a genuine productivity multiplier.</p>
<p><strong>Spec-driven development forces better architecture.</strong> Writing a comprehensive prompt made me think through requirements more carefully than I usually do for side projects.</p>
<p><strong>Learning while building.</strong> Kiro explained concepts as it generated code, turning implementation into education.</p>
<p><strong>Focus on what matters.</strong> Instead of writing boilerplate Terraform and Lambda handlers, I spent time on architecture decisions and integration patterns—the parts that actually require human judgment.</p>
<p>The 60/40 split I experienced (60% worked immediately, 40% needed refinement) is impressive for a first iteration. With better prompts and tighter feedback loops, that ratio keeps improving.</p>
<h2 id="would-i-use-ai-for-this-again">Would I use AI for this again?</h2>
<p>Absolutely. Kiro didn&rsquo;t replace my AWS knowledge—it amplified it. I still made the architectural decisions, understood the tradeoffs, and guided the implementation. But instead of spending days writing infrastructure code, I spent hours refining specifications and reviewing output.</p>
<p>This is the future: engineers focus on architecture, requirements, and integration patterns while AI handles implementation details. The tools get better every month.</p>
<p>The AI writes the code. You make sure it&rsquo;s the right code.</p>
]]></content:encoded></item><item><title>What October 20 Taught Me About DynamoDB (and What It Didn't)</title><link>https://lukelittle.com/posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/</link><pubDate>Sun, 18 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/</guid><description>The October 20 DNS failures broke DynamoDB applications that should have been resilient. Here&amp;#39;s why—and how to actually fix it.</description><content:encoded><![CDATA[<p>On October 20, 2025, DNS resolution failed in AWS us-east-1, and with it, a lot of DynamoDB applications went down.</p>
<p>Not because DynamoDB itself failed. The service was running. Data was there. Capacity was fine. But applications couldn&rsquo;t reach it because DNS queries for <code>dynamodb.us-east-1.amazonaws.com</code> stopped resolving correctly.</p>
<p>If you&rsquo;ve ever wondered what happens when the infrastructure layer beneath your supposedly resilient database becomes unreachable—October 20 was the answer. And it wasn&rsquo;t pretty.</p>
<p>I&rsquo;ve been thinking about that outage a lot lately, especially in the context of the survey application I built to teach students about serverless architecture. That app would have completely failed on October 20. The frontend would load from CloudFront, but every API call would hit a wall trying to reach DynamoDB in us-east-1.</p>
<p>So I decided to figure out what it would actually take to make that architecture survive a regional DNS failure. Not in theory—in practice, with real Terraform and honest tradeoffs.</p>
<p>Here&rsquo;s what I learned.</p>
<h2 id="why-highly-available-wasnt-enough">Why &ldquo;Highly Available&rdquo; Wasn&rsquo;t Enough</h2>
<p>After the outage, I went back and reviewed how I&rsquo;d configured DynamoDB for the survey app. On paper, it looked solid:</p>
<ul>
<li>On-demand capacity (no throttling to worry about)</li>
<li>SDK retries with exponential backoff</li>
<li>Adaptive capacity for hot partitions</li>
<li>CloudWatch alarms watching for errors</li>
</ul>
<p>This is pretty much the standard DynamoDB resilience checklist. And on October 20, none of it helped.</p>
<p>Because retries don&rsquo;t work if the endpoint can&rsquo;t be resolved.</p>
<p>The SDK would try to send a request, fail at DNS resolution, retry with backoff, fail again, retry longer, fail again—all hitting the same DNS wall. Exponential backoff just meant it took longer to give up.</p>
<p>Adaptive capacity is irrelevant when requests never reach the service. The alarms fired, but by the time anyone saw them, users had already moved on.</p>
<p>Here&rsquo;s the uncomfortable realization: most DynamoDB resilience advice assumes the service is reachable. It&rsquo;s optimized for throttling, hot keys, capacity planning. It doesn&rsquo;t address what happens when the layer beneath DynamoDB becomes unavailable.</p>
<p>And that&rsquo;s exactly what happened on October 20.</p>
<h2 id="global-tables-what-they-solve-and-what-they-very-much-dont">Global Tables: What They Solve (and What They Very Much Don&rsquo;t)</h2>
<p>After the outage, the obvious question was: would DynamoDB Global Tables have saved us?</p>
<p>Global Tables give you automated multi-region replication. Write to a table in us-east-1, and the data shows up in us-west-2 within a second or two. It&rsquo;s designed for disaster recovery and geographic distribution.</p>
<p>But here&rsquo;s what Global Tables do not give you:</p>
<ul>
<li>Automatic application failover</li>
<li>DNS independence</li>
<li>Transparent client redirection</li>
</ul>
<p>If your application is configured to talk to <code>dynamodb.us-east-1.amazonaws.com</code> and that endpoint becomes unreachable, Global Tables don&rsquo;t help. Your data sits in us-west-2, perfectly healthy and accessible—but your application never tries to use it.</p>
<p>This is where a lot of architects&rsquo; mental models break down. They think: &ldquo;I have Global Tables, so my data is replicated. I&rsquo;m resilient.&rdquo;</p>
<p>Not quite.</p>
<h2 id="the-hidden-assumption-that-broke-everything">The Hidden Assumption That Broke Everything</h2>
<p>Even with Global Tables configured, most applications still have this somewhere in the code:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span>dynamodb <span style="color:#f92672">=</span> boto3<span style="color:#f92672">.</span>resource(<span style="color:#e6db74">&#39;dynamodb&#39;</span>, region_name<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;us-east-1&#39;</span>)
</span></span></code></pre></div><p>Or this in their Terraform:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-hcl" data-lang="hcl"><span style="display:flex;"><span><span style="color:#66d9ef">provider</span> <span style="color:#e6db74">&#34;aws&#34;</span> {
</span></span><span style="display:flex;"><span>  region <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-east-1&#34;</span>
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>The application is hard-wired to us-east-1. It knows about one region. It sends all traffic there. If DNS in that region fails, the application fails—regardless of how many replica tables exist elsewhere.</p>
<p>This is the &ldquo;oh shit&rdquo; moment: <strong>you replicated your data across the globe, but your application never learned to look anywhere else.</strong></p>
<p>Global Tables solve the data availability problem. They don&rsquo;t solve the application failover problem. And on October 20, it was the second one that broke.</p>
<h2 id="availability-vs-consistency-suddenly-this-matters">Availability vs Consistency (Suddenly This Matters)</h2>
<p>DynamoDB has two consistency modes:</p>
<ul>
<li><strong>Strongly consistent reads</strong>: Always return the most recent write</li>
<li><strong>Eventually consistent reads</strong>: Might return slightly stale data, but stay available during replication lag</li>
</ul>
<p>In normal operations, this is mostly academic. But during a regional DNS failure, the tradeoff becomes very real.</p>
<p>If you fail over to a replica region and issue strongly consistent reads, those reads might succeed or fail depending on replication lag and table state. If you use eventually consistent reads, they&rsquo;ll work—but you might see data that&rsquo;s a few seconds behind.</p>
<p>For the survey app, the choice is obvious: show slightly stale vote counts. A result that&rsquo;s 10 seconds behind is infinitely better than no result at all. Users can still vote. The system degrades gracefully instead of falling over.</p>
<p>But you have to make that decision before the outage, not during it. You can&rsquo;t architect consistency tradeoffs while your dashboard is on fire.</p>
<h2 id="failure-domains-nobody-models">Failure Domains Nobody Models</h2>
<p>Most teams model failures like this:</p>
<ul>
<li>Availability Zone goes down</li>
<li>Service gets throttled</li>
<li>Account gets compromised</li>
</ul>
<p>DNS typically doesn&rsquo;t make the list. It&rsquo;s infrastructure—it just works.</p>
<p>Until it doesn&rsquo;t.</p>
<p>October 20 revealed that regional DNS resolution is a shared fate dependency. When it fails, everything that depends on it fails together. API Gateway, Lambda, DynamoDB, S3—if those services are addressed via DNS in the affected region, you can&rsquo;t reach them.</p>
<p>The survey app has this dependency everywhere: API Gateway in us-east-1 calls Lambda in us-east-1, which calls DynamoDB in us-east-1. Every single one of those calls requires DNS resolution. One DNS failure takes down the entire stack.</p>
<p>A &ldquo;global service&rdquo; like DynamoDB doesn&rsquo;t mean no regional failure modes. It means you need to understand which parts have regional dependencies—and DNS is absolutely one of them.</p>
<h2 id="warm-vs-cold-failover-pick-your-pain">Warm vs Cold Failover (Pick Your Pain)</h2>
<p>There are two ways to handle multi-region failover:</p>
<p><strong>Cold failover</strong> means your secondary region exists but isn&rsquo;t actively serving traffic. When the primary fails, you manually redirect traffic. This is cheaper—you&rsquo;re only paying for data replication—but recovery takes longer.</p>
<p><strong>Warm failover</strong> means your secondary region is live, running infrastructure, and ready to take over immediately. Both regions serve traffic. When one fails, users barely notice. This costs more because you&rsquo;re running duplicate infrastructure.</p>
<p>For the survey app, cold failover looks like:</p>
<ul>
<li>DynamoDB Global Table replicating votes to us-west-2</li>
<li>No Lambda, API Gateway, or CloudFront config for us-west-2</li>
<li>Manual <code>terraform apply</code> when things go wrong</li>
</ul>
<p>Warm failover looks like:</p>
<ul>
<li>Full stack deployed in both regions</li>
<li>Route 53 health checks watching both</li>
<li>Automatic traffic shifting when health checks fail</li>
</ul>
<p>DNS-level failures complicate this decision. With a service outage, cold failover&rsquo;s longer recovery might be acceptable. But when DNS fails, even logging into the console to trigger failover might be impacted.</p>
<p>Warm failover starts looking a lot more attractive when &ldquo;manually fail over&rdquo; might not be possible.</p>
<h2 id="how-traffic-actually-flips-the-part-everyone-handwaves">How Traffic Actually Flips (The Part Everyone Handwaves)</h2>
<p>Failover is a decision, not a default. Something has to decide when to switch regions and actually execute that switch.</p>
<p>For the survey app, there are a few options:</p>
<p><strong>Application-level region awareness</strong><br>
The frontend JavaScript knows about both regions and tries the secondary if the primary fails:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-javascript" data-lang="javascript"><span style="display:flex;"><span><span style="color:#66d9ef">const</span> <span style="color:#a6e22e">regions</span> <span style="color:#f92672">=</span> [
</span></span><span style="display:flex;"><span>  { <span style="color:#a6e22e">api</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;https://api-us-east-1.example.com&#39;</span>, <span style="color:#a6e22e">name</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;us-east-1&#39;</span> },
</span></span><span style="display:flex;"><span>  { <span style="color:#a6e22e">api</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;https://api-us-west-2.example.com&#39;</span>, <span style="color:#a6e22e">name</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;us-west-2&#39;</span> }
</span></span><span style="display:flex;"><span>];
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">async</span> <span style="color:#66d9ef">function</span> <span style="color:#a6e22e">submitVote</span>(<span style="color:#a6e22e">vote</span>) {
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">for</span> (<span style="color:#66d9ef">const</span> <span style="color:#a6e22e">region</span> <span style="color:#66d9ef">of</span> <span style="color:#a6e22e">regions</span>) {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">try</span> {
</span></span><span style="display:flex;"><span>      <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">response</span> <span style="color:#f92672">=</span> <span style="color:#66d9ef">await</span> <span style="color:#a6e22e">fetch</span>(<span style="color:#e6db74">`</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">region</span>.<span style="color:#a6e22e">api</span><span style="color:#e6db74">}</span><span style="color:#e6db74">/vote`</span>, {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">method</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;POST&#39;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">body</span><span style="color:#f92672">:</span> <span style="color:#a6e22e">JSON</span>.<span style="color:#a6e22e">stringify</span>({ <span style="color:#a6e22e">vote</span> })
</span></span><span style="display:flex;"><span>      });
</span></span><span style="display:flex;"><span>      <span style="color:#66d9ef">return</span> <span style="color:#a6e22e">response</span>;
</span></span><span style="display:flex;"><span>    } <span style="color:#66d9ef">catch</span> (<span style="color:#a6e22e">error</span>) {
</span></span><span style="display:flex;"><span>      <span style="color:#a6e22e">console</span>.<span style="color:#a6e22e">log</span>(<span style="color:#e6db74">`</span><span style="color:#e6db74">${</span><span style="color:#a6e22e">region</span>.<span style="color:#a6e22e">name</span><span style="color:#e6db74">}</span><span style="color:#e6db74"> failed, trying next`</span>);
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">throw</span> <span style="color:#66d9ef">new</span> Error(<span style="color:#e6db74">&#39;All regions failed&#39;</span>);
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>This works, but now your frontend code is aware of regional infrastructure. That complexity leaks all the way to the browser.</p>
<p>Here&rsquo;s what this application-level failover looks like in practice:</p>
<pre class="mermaid">sequenceDiagram
    autonumber
    participant Client
    participant Route53 as Route 53
    participant APIGW_East as API Gateway&lt;br/&gt;(us-east-1)
    participant Lambda_East as Lambda&lt;br/&gt;(us-east-1)
    participant DDB_East as DynamoDB&lt;br/&gt;(us-east-1)
    participant APIGW_West as API Gateway&lt;br/&gt;(us-west-2)
    participant Lambda_West as Lambda&lt;br/&gt;(us-west-2)
    participant DDB_West as DynamoDB&lt;br/&gt;(us-west-2)

    Note over Client,DDB_West: Scenario 1: Normal Operation (us-east-1 healthy)
    Client-&gt;&gt;Route53: Request vote submission
    Route53-&gt;&gt;APIGW_East: Route to us-east-1
    APIGW_East-&gt;&gt;Lambda_East: Invoke function
    Lambda_East-&gt;&gt;DDB_East: Write vote
    DDB_East--&gt;&gt;Lambda_East: Success
    Lambda_East--&gt;&gt;APIGW_East: 200 OK
    APIGW_East--&gt;&gt;Client: Vote recorded ✓

    Note over Client,DDB_West: Scenario 2: October 20 DNS Failure (no failover)
    Client-&gt;&gt;Route53: Request vote submission
    Route53-xAPIGW_East: DNS resolution fails ❌
    Note over Client: Request times out&lt;br/&gt;User sees error

    Note over Client,DDB_West: Scenario 3: DNS Failure with Application Failover
    Client-&gt;&gt;Route53: Request vote submission
    Route53-xAPIGW_East: DNS resolution fails ❌
    Note over Client: Client catches error,&lt;br/&gt;retries us-west-2
    Client-&gt;&gt;Route53: Retry request
    Route53-&gt;&gt;APIGW_West: Route to us-west-2
    APIGW_West-&gt;&gt;Lambda_West: Invoke function
    Lambda_West-&gt;&gt;DDB_West: Write vote (replica)
    DDB_West--&gt;&gt;Lambda_West: Success
    Lambda_West--&gt;&gt;APIGW_West: 200 OK
    APIGW_West--&gt;&gt;Client: Vote recorded ✓
    Note over DDB_East,DDB_West: Cross-region replication&lt;br/&gt;(&lt; 1 second)
</pre>

<p>The diagram illustrates three scenarios:</p>
<ol>
<li><strong>Normal operation</strong>: Everything works in us-east-1</li>
<li><strong>October 20 failure</strong>: DNS fails and users see errors</li>
<li><strong>With application failover</strong>: When us-east-1 fails, the client automatically retries with us-west-2 and succeeds</li>
</ol>
<p><strong>Feature flags or configuration toggles</strong><br>
An external service (like LaunchDarkly) controls which region receives traffic. During an outage, ops flips the flag. This centralizes the logic but adds another dependency—and another potential failure point.</p>
<p><strong>Route 53 health checks</strong><br>
DNS-based failover that automatically routes to healthy endpoints. This works for many scenarios, but if regional DNS is failing, Route 53 lookups might also be impacted.</p>
<p>The uncomfortable truth: there&rsquo;s no perfect solution. Each approach has tradeoffs around complexity, blast radius, and new failure modes.</p>
<h2 id="making-the-survey-app-actually-resilient">Making the Survey App Actually Resilient</h2>
<p><img src="/posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/surviving-dns-failures-dynamodb-multi-region_hu_330758da10979ec0.webp" srcset="/posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/surviving-dns-failures-dynamodb-multi-region_hu_aed0225e0a461084.webp 750w, /posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/surviving-dns-failures-dynamodb-multi-region_hu_330758da10979ec0.webp 1401w" sizes="(max-width: 800px) 100vw, 750px"
       width="1401" height="1000" alt="Surviving DNS Failures with Multi-Region DynamoDB" loading="lazy" decoding="async"></p>
<p>Let&rsquo;s revisit the serverless survey application. The original architecture:</p>
<ul>
<li>S3 + CloudFront serving the frontend</li>
<li>API Gateway exposing REST endpoints</li>
<li>Lambda functions (vote, results, reset)</li>
<li>DynamoDB storing votes</li>
<li>Everything in us-east-1</li>
</ul>
<p>On October 20, this would have completely failed. CloudFront would serve the static site from its global cache, but every API call would hit DNS failures trying to reach API Gateway in us-east-1.</p>
<p>Students would see the survey form but couldn&rsquo;t vote. The results page would spin forever. The reset function would be unreachable.</p>
<p>To make this resilient, I&rsquo;d implement warm failover with application-level region awareness.</p>
<p><strong>What keeps working during a DNS failure</strong>:</p>
<ul>
<li>Static frontend loads from CloudFront (it&rsquo;s global anyway)</li>
<li>Vote submissions succeed by failing over to us-west-2 API</li>
<li>Results page shows counts from the replica table</li>
<li>Eventually consistent reads mean slight lag is acceptable</li>
</ul>
<p><strong>What degrades gracefully</strong>:</p>
<ul>
<li>Latency goes up for users far from the secondary region</li>
<li>Global Table replication lag means new votes take longer to appear everywhere</li>
<li>No strong consistency guarantees during failover</li>
</ul>
<p><strong>What stops working</strong>:</p>
<ul>
<li>Nothing critical (that&rsquo;s the entire point)</li>
</ul>
<p>This is graceful degradation instead of complete failure.</p>
<h2 id="what-multi-region-dynamodb-actually-looks-like">What Multi-Region DynamoDB Actually Looks Like</h2>
<p>Here&rsquo;s the Terraform for making the survey app multi-region:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-hcl" data-lang="hcl"><span style="display:flex;"><span><span style="color:#75715e"># Primary region
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">provider</span> <span style="color:#e6db74">&#34;aws&#34;</span> {
</span></span><span style="display:flex;"><span>  alias  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;primary&#34;</span>
</span></span><span style="display:flex;"><span>  region <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-east-1&#34;</span>
</span></span><span style="display:flex;"><span>}<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"># Secondary region
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">provider</span> <span style="color:#e6db74">&#34;aws&#34;</span> {
</span></span><span style="display:flex;"><span>  alias  <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;secondary&#34;</span>
</span></span><span style="display:flex;"><span>  region <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-west-2&#34;</span>
</span></span><span style="display:flex;"><span>}<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"># DynamoDB table with Global Tables enabled
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_dynamodb_table&#34; &#34;survey_votes&#34;</span> {
</span></span><span style="display:flex;"><span>  provider         <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws</span>.<span style="color:#66d9ef">primary</span>
</span></span><span style="display:flex;"><span>  name             <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;survey-votes&#34;</span>
</span></span><span style="display:flex;"><span>  billing_mode     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;PAY_PER_REQUEST&#34;</span>
</span></span><span style="display:flex;"><span>  hash_key         <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;id&#34;</span>
</span></span><span style="display:flex;"><span>  stream_enabled   <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>
</span></span><span style="display:flex;"><span>  stream_view_type <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;NEW_AND_OLD_IMAGES&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">attribute</span> {
</span></span><span style="display:flex;"><span>    name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;id&#34;</span>
</span></span><span style="display:flex;"><span>    type <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;S&#34;</span>
</span></span><span style="display:flex;"><span>  }<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">  # This one line enables Global Tables
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>  <span style="color:#66d9ef">replica</span> {
</span></span><span style="display:flex;"><span>    region_name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-west-2&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  tags <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>    Environment <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;production&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"># Lambda in primary region
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_lambda_function&#34; &#34;vote_primary&#34;</span> {
</span></span><span style="display:flex;"><span>  provider      <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws</span>.<span style="color:#66d9ef">primary</span>
</span></span><span style="display:flex;"><span>  function_name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;survey-vote&#34;</span>
</span></span><span style="display:flex;"><span>  runtime       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;python3.11&#34;</span>
</span></span><span style="display:flex;"><span>  handler       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;vote.handler&#34;</span>
</span></span><span style="display:flex;"><span>  filename      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;lambda/vote.zip&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">environment</span> {
</span></span><span style="display:flex;"><span>    variables <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>      TABLE_NAME <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws_dynamodb_table</span>.<span style="color:#66d9ef">survey_votes</span>.<span style="color:#66d9ef">name</span>
</span></span><span style="display:flex;"><span>      REGION     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-east-1&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}<span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e">
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"># Lambda in secondary region (same code, different region)
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_lambda_function&#34; &#34;vote_secondary&#34;</span> {
</span></span><span style="display:flex;"><span>  provider      <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws</span>.<span style="color:#66d9ef">secondary</span>
</span></span><span style="display:flex;"><span>  function_name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;survey-vote&#34;</span>
</span></span><span style="display:flex;"><span>  runtime       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;python3.11&#34;</span>
</span></span><span style="display:flex;"><span>  handler       <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;vote.handler&#34;</span>
</span></span><span style="display:flex;"><span>  filename      <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;lambda/vote.zip&#34;</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">environment</span> {
</span></span><span style="display:flex;"><span>    variables <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>      TABLE_NAME <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws_dynamodb_table</span>.<span style="color:#66d9ef">survey_votes</span>.<span style="color:#66d9ef">name</span>
</span></span><span style="display:flex;"><span>      REGION     <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;us-west-2&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>The critical piece is the <code>replica</code> block. That tells DynamoDB to automatically replicate to us-west-2. AWS handles the replication—you don&rsquo;t write code for it.</p>
<p>But notice: you still need to deploy Lambda, API Gateway, and all the supporting infrastructure in both regions. Global Tables replicate data, not infrastructure.</p>
<h2 id="architecture-diagram-description">Architecture Diagram Description</h2>
<p>The architecture flows left to right:</p>
<p><strong>Far left: Client layer</strong></p>
<ul>
<li>User browsers, mobile apps, any HTTP client</li>
</ul>
<p><strong>Edge layer (global)</strong></p>
<ul>
<li>CloudFront distribution serving static files (HTML, CSS, JS)</li>
<li>Label: &ldquo;Global Edge Network&rdquo;</li>
</ul>
<p><strong>Middle: Application layer (two parallel stacks)</strong></p>
<p><strong>Primary region stack (us-east-1)</strong>:</p>
<ul>
<li>API Gateway endpoint</li>
<li>Three Lambda functions (vote, results, reset)</li>
<li>Connected to DynamoDB primary table</li>
<li>Status indicator: normally green (&ldquo;Healthy&rdquo;), red during October 20 (&ldquo;DNS Failed&rdquo;)</li>
</ul>
<p><strong>Secondary region stack (us-west-2)</strong>:</p>
<ul>
<li>Identical API Gateway endpoint</li>
<li>Identical Lambda functions</li>
<li>Connected to DynamoDB replica table</li>
<li>Status indicator: green (&ldquo;Healthy&rdquo;)</li>
</ul>
<p><strong>Far right: Data layer</strong></p>
<ul>
<li>Two DynamoDB tables shown side-by-side</li>
<li>Primary table (us-east-1)</li>
<li>Replica table (us-west-2)</li>
<li>Bi-directional arrows between them labeled &ldquo;&lt; 1s replication&rdquo;</li>
<li>Both showing identical data</li>
</ul>
<p><strong>Failover control (dashed line across diagram)</strong>:</p>
<ul>
<li>Route 53 health checks monitoring both regions</li>
<li>Decision diamond: &ldquo;Primary healthy?&rdquo;</li>
<li>Yes → route to us-east-1</li>
<li>No → route to us-west-2</li>
<li>Alternative path: client-side retry logic (try primary, fall back to secondary)</li>
</ul>
<p><strong>DNS dependency markers (red warning icons)</strong>:</p>
<ul>
<li>DNS resolution required for API Gateway in each region</li>
<li>DNS resolution required for DynamoDB endpoints in each region</li>
<li>These are the exact failure points October 20 exposed</li>
</ul>
<p><strong>Bottom: Observability signals</strong></p>
<ul>
<li>CloudWatch alarms in both regions</li>
<li>Metrics: API error rate, Lambda duration, DynamoDB throttles</li>
<li>Threshold shown: &ldquo;Error rate &gt; 5% for 2 minutes → consider failover&rdquo;</li>
</ul>
<p>The diagram makes one thing brutally clear: when DNS fails in us-east-1, the entire primary stack becomes unreachable—but us-west-2 keeps running. Data replication keeps them in sync. Application failover keeps users working.</p>
<h2 id="tradeoffs-nobody-wants-to-talk-about">Tradeoffs Nobody Wants to Talk About</h2>
<p>This architecture is more complex than single-region. You&rsquo;re maintaining infrastructure in two regions, managing replication, handling eventual consistency, and building failover logic.</p>
<p>For the survey app—which costs $0/month in a single region—running warm failover in two regions might cost $15/month. That&rsquo;s not a lot in absolute terms, but it&rsquo;s infinitely more expensive than free.</p>
<p>And here&rsquo;s an uncomfortable truth: most teams won&rsquo;t actually test their failover. They&rsquo;ll build it, deploy it, document it, assume it works, and never validate it until a real outage happens.</p>
<p>When October 20 comes, they&rsquo;ll discover:</p>
<ul>
<li>Their failover mechanism has a bug</li>
<li>Their health checks have false positives</li>
<li>Their SDK client caching interferes with region switching</li>
<li>Their observability doesn&rsquo;t show which region is actually serving traffic</li>
</ul>
<p>Multi-region resilience requires ongoing operational investment. It&rsquo;s not &ldquo;set and forget.&rdquo; You need runbooks, chaos engineering, regular failover drills, and teams who know how to operate it.</p>
<p>That&rsquo;s a lot of overhead for a student survey app.</p>
<h2 id="who-actually-needs-this">Who Actually Needs This</h2>
<p>Not every DynamoDB workload needs multi-region failover.</p>
<p><strong>You should build this if</strong>:</p>
<ul>
<li>Downtime directly costs revenue</li>
<li>You have strict uptime SLAs (99.95%+)</li>
<li>Users are globally distributed</li>
<li>Regulatory requirements mandate geographic redundancy</li>
<li>You&rsquo;ve done the math: outage cost &gt; infrastructure cost</li>
</ul>
<p><strong>You&rsquo;re probably over-engineering if</strong>:</p>
<ul>
<li>Your app is internal tooling</li>
<li>Downtime measured in hours is tolerable</li>
<li>You&rsquo;re optimizing for shipping speed over resilience</li>
<li>Your team doesn&rsquo;t have the operational maturity to manage this complexity</li>
</ul>
<p>The survey app I built for students? It absolutely doesn&rsquo;t need this. Downtime is annoying, not catastrophic. It&rsquo;s a teaching tool, not a production system.</p>
<p>But if you&rsquo;re running live event voting, mobile game leaderboards, or SaaS APIs that customers depend on—then yes, this makes sense.</p>
<p>The decision isn&rsquo;t technical. It&rsquo;s about risk tolerance, recovery time objectives, and operational burden.</p>
<p>On October 20, a lot of teams learned they&rsquo;d miscalculated that risk. They assumed regional DNS wouldn&rsquo;t fail.</p>
<p>It did.</p>
<p>Learn from that. Design for the failures that actually happen, not the ones that feel unlikely.</p>
]]></content:encoded></item><item><title>Data Pour with Nimish Donde: Resiliency, Data Gravity, and Building Cloud Platforms That Scale</title><link>https://lukelittle.com/posts/2026/01/data-pour-with-nimish-donde-resiliency-data-gravity-and-building-cloud-platforms-that-scale/</link><pubDate>Fri, 16 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/data-pour-with-nimish-donde-resiliency-data-gravity-and-building-cloud-platforms-that-scale/</guid><description>A conversation at Amélie&amp;#39;s French bakery with Nimish Donde about cloud transformation in banking, why resiliency is a mindset, and how to unlock data for AI.</description><content:encoded><![CDATA[<p>For this episode of Data Pour, I sat down with Nimish Donde—Head of Cloud Platform and Security Engineering at Truist—at Amélie&rsquo;s French bakery in Charlotte. It&rsquo;s a place that&rsquo;s been part of the city&rsquo;s fabric since 2008, growing from a single 24-hour location in NoDa (that I used to frequent during college) to four locations across Charlotte.</p>
<p>Like this bakery, Charlotte&rsquo;s tech scene has grown and evolved—and Nimish has been part of that transformation for the past 16 years.</p>
<h2 id="why-nimish">Why Nimish</h2>
<p>Nimish is one of my closest friends and mentors. We both worked on Ally&rsquo;s cloud transformation together—15 years for him at Ally, where he helped navigate a massive financial institution through one of the most significant technology shifts in banking.</p>
<p>Now he&rsquo;s at Truist, a much larger bank, tackling similar challenges but at an even greater scale. And what I&rsquo;ve always valued about Nimish is his ability to cut through the noise. He doesn&rsquo;t chase technology trends—he focuses on what actually delivers value to customers, builds trust with people, and creates platforms that work.</p>
<p>This conversation was long overdue. We grabbed French press coffee (the kind you order at Amélie&rsquo;s in those big presses), settled in, and talked about everything from resiliency and chaos engineering to Charlotte&rsquo;s evolution as a fintech hub.</p>
<h2 id="charlotte-the-second-largest-financial-hub-in-america">Charlotte: the second-largest financial hub in America</h2>
<p>One of the threads we explored early was Charlotte itself. It&rsquo;s the second-largest financial hub in the United States—home to Bank of America, Wells Fargo, Truist, Ally, and a growing number of financial enterprises setting up operations.</p>
<p>Over the past two decades, Nimish has watched the city transform from just a banking hub into a <strong>tech-financial hub</strong>—the birthplace of fintechs, a magnet for cloud talent, and a city where people now <em>want</em> to move for their careers, not just pass through on their way to New York or Silicon Valley.</p>
<p>Charlotte showed grit after the 2008 financial crisis. The city came back stronger, the tech scene expanded, and the talent pipeline from local universities (UNC Charlotte, Wake Forest, UNC, USC) started feeding directly into these financial institutions.</p>
<p>Even Amélie&rsquo;s mirrors that resilience. This location shut down during the pandemic but came back in 2023—booming again, just like the city itself.</p>
<h2 id="cloud-platforms-as-products-not-projects">Cloud platforms as products, not projects</h2>
<p>Nimish&rsquo;s role at Truist breaks down into three core tenets:</p>
<ol>
<li><strong>Developer experience and enablement</strong> — Lower the barrier to entry so more teams can build on the platform</li>
<li><strong>Resiliency and reliability</strong> — Ensure customers (internal developers) get maximum value</li>
<li><strong>AI and data gravity</strong> — Enable the platform to accelerate AI journeys</li>
</ol>
<p>The key insight here: <strong>cloud platforms are products, not operational systems</strong>.</p>
<p>Too many organizations treat cloud as infrastructure to manage. Nimish treats it as a product with customers—and those customers are the developers building applications that serve end users. Everything flows from that mindset.</p>
<h2 id="resiliency-is-a-mindset-not-a-metric">Resiliency is a mindset, not a metric</h2>
<p>We spent a significant chunk of the conversation on resiliency—something both of us are obsessed with.</p>
<p>Nimish&rsquo;s definition: <strong>resiliency is a mindset, not a destination</strong>.</p>
<p>Gone are the days when you measured platform quality by uptime percentages. Resiliency today is about the <em>experiences</em> you build for customers. It&rsquo;s about how you respond under stress. It&rsquo;s about building systems that anticipate failure, learn from it, and continuously improve.</p>
<p>He referenced the SRE evolution and Google SRE&rsquo;s famous motto: &ldquo;Hope is not a strategy.&rdquo; That philosophy shaped how Nimish approaches resiliency today.</p>
<h3 id="chaos-engineering-as-a-first-class-citizen">Chaos engineering as a first-class citizen</h3>
<p>One of the most interesting parts of our discussion was around <strong>chaos engineering</strong>.</p>
<p>Regulators—especially in Europe with legislation like DORA—are no longer accepting tabletop exercises as proof of resiliency. They want evidence. They want to see that under stress, your systems still deliver value.</p>
<p>Chaos engineering isn&rsquo;t just a nice-to-have anymore. It&rsquo;s becoming a compliance requirement. It&rsquo;s how you prove resiliency in production. And for banks, that shift is profound.</p>
<p>Nimish believes chaos engineering will become a first-class citizen in the audit process—just like security compliance scores today. You&rsquo;ll need to validate your platforms&rsquo; resilience continuously, not just once a year.</p>
<h2 id="the-8020-rule-technology-is-the-easy-20">The 80/20 rule: technology is the easy 20%</h2>
<p>A theme that came up repeatedly: <strong>technology is the easy part</strong>.</p>
<p>Whether it&rsquo;s on-prem systems or cloud platforms, the binary is only 20% of the challenge. The other 80%? People.</p>
<p>Bringing people along on transformation. Building influence. Convincing risk partners, internal audit, regulators, and developers that this is the right model. Creating feedback loops. Enabling self-service while embedding guardrails.</p>
<p>Nimish has mastered that 80%. And he&rsquo;s clear about it: if you can&rsquo;t influence people, your platform will fail—even if the technology is perfect.</p>
<h2 id="data-gravity-and-ai-you-cant-skip-the-foundation">Data gravity and AI: you can&rsquo;t skip the foundation</h2>
<p>We talked extensively about <strong>data gravity</strong>—the idea that data needs a central home, properly cataloged and governed, before AI can deliver real value.</p>
<p>Nimish&rsquo;s take: <strong>garbage in, garbage out</strong>.</p>
<p>You can&rsquo;t unlock AI value if your data is in disparate locations, inconsistent, or inaccessible. Executive sponsorship is critical. You need investment in data warehousing, enrichment, cataloging, and lineage <em>before</em> you start training large language models.</p>
<p>His advice for enterprises: <strong>eat your own dog food</strong>.</p>
<p>If you&rsquo;re the cloud platform team, build your own data warehouse. Use the platform you&rsquo;re providing to developers. Give them transparency into their usage. Lower the barrier to entry by shining a light on dark spaces where people wouldn&rsquo;t normally pay attention.</p>
<p>That transparency accelerates the entire organization&rsquo;s AI journey.</p>
<h2 id="guardrails-as-code-the-tiered-cake-model">Guardrails as code: the tiered cake model</h2>
<p>Nimish shared his <strong>tiered cake model</strong> for building secure, compliant platforms—an analogy that resonated deeply (especially sitting in a French bakery):</p>
<ol>
<li><strong>Base layer</strong>: Landing zones with preventative and detective controls (policies, SCPs, IAM boundaries)</li>
<li><strong>Second layer</strong>: Data visibility—a warehouse that tells the story of what&rsquo;s happening on your platform</li>
<li><strong>Third layer</strong>: Proactive controls—CI/CD, Infrastructure as Code, reusable templates, inner source</li>
<li><strong>Cherry on top</strong>: Human interaction—communities of practice, office hours, self-service tools</li>
</ol>
<p>Each layer has guardrails baked in. And when packaged together, you can hand the whole thing to auditors and regulators with full transparency on ingredients, processes, and governance.</p>
<p>The result? Developers move fast. Compliance is embedded. And the enterprise is protected.</p>
<h2 id="advice-for-young-engineers-be-resilient-adaptable-and-comfortable-with-ambiguity">Advice for young engineers: be resilient, adaptable, and comfortable with ambiguity</h2>
<p>I asked Nimish what he&rsquo;d tell a young engineer aspiring to leadership.</p>
<p>His answer: <strong>three things matter most</strong>:</p>
<ol>
<li><strong>Resilience</strong> — Bounce back. Learn from failure. Keep going.</li>
<li><strong>Adaptability</strong> — The tech changes constantly. Your ability to evolve matters more than what you know today.</li>
<li><strong>Comfort with ambiguity</strong> — Leadership isn&rsquo;t clean. You won&rsquo;t always have the answer. Learn to navigate uncertainty.</li>
</ol>
<p>And most importantly: <strong>have empathy</strong>. Understand people. Use that as an accelerator.</p>
<p>Technology skills get you in the door. Soft skills keep you there.</p>
<h2 id="why-this-conversation-mattered">Why this conversation mattered</h2>
<p>Nimish and I have known each other for years, but this was the first time we sat down and recorded a full conversation about the work we&rsquo;ve been doing. It felt less like an interview and more like two friends catching up—talking through transformation, resiliency, Charlotte&rsquo;s evolution, and where the industry is heading.</p>
<p>I&rsquo;m grateful he made time to do this. If you work in cloud, fintech, or platform engineering—or if you&rsquo;re just curious about what it takes to build systems that scale in highly regulated environments—this episode is worth your time.</p>
<h2 id="watch-the-full-episode">Watch the full episode</h2>
<p>This blog is just the teaser. If you want the full story—including our takes on agentic AI, the growth of Charlotte&rsquo;s microbrewery scene (which somehow correlates with the tech scene), and Nimish&rsquo;s upcoming hike to Machu Picchu—watch the episode.</p>
<p>And if you&rsquo;re ever in Charlotte, grab a salted caramel brownie at Amélie&rsquo;s. Nimish has been a fan for 15 years. I&rsquo;m partial to the pistachio macaron.</p>
<p>Watch the full Data Pour episode here:<br>
<a href="https://youtu.be/3xuco4R8EHE">https://youtu.be/3xuco4R8EHE</a></p>
]]></content:encoded></item><item><title>Your AI On-Call Engineer: Inside AWS DevOps Agent</title><link>https://lukelittle.com/posts/2026/01/your-ai-on-call-engineer-inside-aws-devops-agent/</link><pubDate>Mon, 12 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/your-ai-on-call-engineer-inside-aws-devops-agent/</guid><description>How AWS&amp;#39;s frontier agents are changing incident response forever - an autonomous DevOps engineer that works while you sleep</description><content:encoded><![CDATA[<p>At re:Invent 2025, AWS CEO Matt Garman announced something that made me stop and actually pay attention during a keynote—which doesn&rsquo;t happen often.</p>
<p>He introduced <strong>frontier agents</strong>: AI systems that don&rsquo;t just help you write code or answer questions. They work autonomously for hours or days, maintaining context, investigating problems, and making decisions without you holding their hand.</p>
<p>Three agents got announced:</p>
<ul>
<li><strong>Kiro</strong> - your AI developer</li>
<li><strong>AWS Security Agent</strong> - your AI security engineer</li>
<li><strong>AWS DevOps Agent</strong> - your AI operations engineer</li>
</ul>
<p>This isn&rsquo;t another coding assistant that autocompletes your Lambda functions. This is AWS betting that AI agents can handle the kind of multi-hour incident investigations that currently wake up humans at 2 AM.</p>
<p>Let&rsquo;s be real: AI-assisted incident response isn&rsquo;t new. PagerDuty, Datadog, Dynatrace, and a dozen startups have been doing &ldquo;pull operational data into an LLM and suggest fixes&rdquo; for years. What makes AWS DevOps Agent different is the depth of integration into the AWS control plane and the architectural pattern it represents.</p>
<p>You can watch the frontier agents announcement teaser here: <a href="https://www.youtube.com/watch?v=fMQfzwS0prQ">https://www.youtube.com/watch?v=fMQfzwS0prQ</a></p>
<p>But I wanted to dig deeper and figure out what this actually means for teams running production systems on AWS.</p>
<h2 id="what-makes-an-agent-frontier-class">What makes an agent &ldquo;frontier-class&rdquo;</h2>
<p>AWS uses the term &ldquo;frontier agent&rdquo; to mean something specific. It&rsquo;s not just GPT-4 with AWS API access.</p>
<p><strong>1. Autonomous goal-directed behavior</strong></p>
<p>Traditional AI: &ldquo;Hey ChatGPT, what might cause high Lambda errors?&rdquo;<br>
Frontier agent: &ldquo;Investigate this Lambda error spike&rdquo; → agent figures out how</p>
<p>You give it an objective, it decomposes the problem, forms hypotheses, collects evidence, and executes—without asking you for step-by-step guidance.</p>
<p><strong>2. Multi-agent coordination</strong></p>
<p>DevOps Agent doesn&rsquo;t work alone. When investigating an incident, it spawns specialized sub-agents—one analyzing logs, another reconstructing the deployment timeline, a third mapping topology. These agents run concurrently, investigating multiple hypotheses simultaneously and coordinating across AWS accounts. It&rsquo;s less &ldquo;AI assistant&rdquo; and more &ldquo;AI team.&rdquo;</p>
<p><strong>3. Long-running independent operation</strong></p>
<p>Here&rsquo;s the paradigm shift: it works for hours without constant human intervention.</p>
<p>Traditional AI forgets everything when you close the chat. Frontier agents maintain persistent context, remember your infrastructure, learn from past incidents, and pick up where they left off after restarts.</p>
<p>When your Lambda error alarm goes off at 2 AM, DevOps Agent can investigate for 30 minutes, form a hypothesis, collect evidence, and have a diagnosis ready by the time you wake up and check Slack.</p>
<h2 id="how-it-actually-works">How it actually works</h2>
<p>DevOps Agent integrates with your existing monitoring tools—it doesn&rsquo;t replace them.</p>
<p>On the observability side, DevOps Agent integrates natively with CloudWatch and can pull data from Datadog, Dynatrace, New Relic, and Splunk. If you&rsquo;re using custom monitoring tools, you can build integrations via Model Context Protocol (MCP) servers—AWS&rsquo;s standard for extending agent capabilities.</p>
<p>For incident coordination, there&rsquo;s built-in support for ServiceNow and PagerDuty, plus Slack for real-time updates. Pretty much any tool with webhooks can be integrated into the workflow.</p>
<p>DevOps Agent can be triggered three ways: automatically when a CloudWatch alarm fires (fully autonomous response), manually through the web UI when you want to investigate something specific, or on a schedule for proactive analysis—think nightly scans looking for anomalies before they become incidents.</p>
<p>When an alert fires—say, Lambda errors spike at 2 AM—here&rsquo;s what happens:</p>
<pre class="mermaid">graph TB
    Alert[CloudWatch Alarm Fires] --&gt; Orchestrator[Investigation Orchestrator]
    
    Orchestrator --&gt; Topo[Topology Sub-Agent&lt;br/&gt;Maps dependencies]
    Orchestrator --&gt; Telem[Telemetry Sub-Agent&lt;br/&gt;Analyzes metrics/logs]
    Orchestrator --&gt; Deploy[Deployment Sub-Agent&lt;br/&gt;Checks recent changes]
    
    Topo --&gt; RCA[Root Cause Analysis]
    Telem --&gt; RCA
    Deploy --&gt; RCA
    
    RCA --&gt; Slack[Post to Slack #incidents]
    RCA --&gt; Ticket[Create ServiceNow ticket]
</pre>

<p>The clever part is the <strong>application topology map</strong>. DevOps Agent builds and maintains an intelligent map of your entire system—which Lambda functions call which APIs, which services depend on which databases, when each component was last deployed and by whom. It tracks cross-account dependencies and even external dependencies like third-party APIs, SaaS integrations, and CDNs.</p>
<p>When an incident happens, this topology becomes invaluable. The agent can immediately identify blast radius (what&rsquo;s affected by this outage?), trace dependency chains (if the API is down, what upstream services caused it? what downstream services are impacted?), and correlate timing (there was a deploy 15 minutes ago—is that when this started?).</p>
<h2 id="the-investigation-loop">The investigation loop</h2>
<p>Once triggered, DevOps Agent enters an iterative loop:</p>
<ol>
<li><strong>Generate hypotheses</strong> based on alert type, topology, recent changes</li>
<li><strong>Collect evidence</strong> by querying logs, metrics, traces, configs</li>
<li><strong>Correlate patterns</strong> across time, services, accounts</li>
<li><strong>Assess confidence</strong> in each hypothesis</li>
<li><strong>Recommend mitigation</strong> or continue investigating</li>
<li><strong>Learn from outcome</strong> to improve future investigations</li>
</ol>
<p>This keeps going until it reaches high confidence in the root cause or exhausts reasonable paths.</p>
<p>What separates this from dumb rule-based systems: it doesn&rsquo;t just pattern-match. It reasons about your infrastructure.</p>
<h2 id="agent-spaces-and-iam-permission-boundaries">Agent Spaces and IAM permission boundaries</h2>
<p>Everything starts with an <strong>Agent Space</strong>—the workspace where the agent operates and the IAM permission boundary defining what it can access.</p>
<p>You can structure Agent Spaces multiple ways:</p>
<ul>
<li><strong>Per-application:</strong> One space per critical app</li>
<li><strong>Per-team:</strong> One space per on-call team</li>
<li><strong>Centralized:</strong> One space in monitoring account observing everything</li>
</ul>
<p>Here&rsquo;s what makes this not just &ldquo;magic AI with root access&rdquo;: DevOps Agent uses explicit, auditable IAM trust relationships.</p>
<p>The Agent Space role trust policy:</p>
<ul>
<li><strong>Principal:</strong> <code>aidevops.amazonaws.com</code> (not some opaque service)</li>
<li><strong>Conditions:</strong> SourceAccount and SourceArn bound to your specific AgentSpace</li>
<li><strong>Permissions:</strong> Standard IAM policies you control</li>
</ul>
<p>You can audit exactly what DevOps Agent accessed, when, and why. It&rsquo;s not a black box.</p>
<h3 id="multi-account-setup-the-real-production-pattern">Multi-account setup (the real production pattern)</h3>
<p>Production incidents rarely happen in a single AWS account. You have workload accounts, shared services accounts, security accounts, monitoring accounts.</p>
<p>DevOps Agent supports this natively via External Account Associations:</p>
<pre class="mermaid">graph TD
    subgraph Monitor[&#34;Monitoring Account&#34;]
        AgentSpace[DevOps Agent Space]
    end
    
    subgraph Workload1[&#34;Workload Account 1&#34;]
        Role1[IAM Role&lt;br/&gt;ReadOnly + Logs]
    end
    
    subgraph Workload2[&#34;Workload Account 2&#34;]
        Role2[IAM Role&lt;br/&gt;ReadOnly + Logs]
    end
    
    AgentSpace --&gt;|Cross-account trust| Role1
    AgentSpace --&gt;|Cross-account trust| Role2
</pre>

<p>Create your Agent Space in a central monitoring account, associate it with workload accounts via cross-account IAM roles, and let it investigate incidents spanning account boundaries.</p>
<p>This is how you do AWS at scale.</p>
<h2 id="deploying-it-terraform-example">Deploying it (Terraform example)</h2>
<p>AWS provides Terraform resources (<code>aws_devopsagent_agentspace</code>, <code>aws_devopsagent_association</code>) and CDK constructs.</p>
<p>Basic Terraform setup:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-hcl" data-lang="hcl"><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_devopsagent_agentspace&#34; &#34;main&#34;</span> {
</span></span><span style="display:flex;"><span>  name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;production-monitoring&#34;</span>
</span></span><span style="display:flex;"><span>  
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">agent_role</span> {
</span></span><span style="display:flex;"><span>    create_role <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>
</span></span><span style="display:flex;"><span>    role_name   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;DevOpsAgentSpaceRole&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>  
</span></span><span style="display:flex;"><span>  enable_web_app <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span><span style="color:#75715e">  # Optional UI
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_devopsagent_association&#34; &#34;workload&#34;</span> {
</span></span><span style="display:flex;"><span>  agent_space_id <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws_devopsagent_agentspace</span>.<span style="color:#66d9ef">main</span>.<span style="color:#66d9ef">id</span>
</span></span><span style="display:flex;"><span>  
</span></span><span style="display:flex;"><span>  <span style="color:#66d9ef">external_account</span> {
</span></span><span style="display:flex;"><span>    account_id <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;987654321098&#34;</span>
</span></span><span style="display:flex;"><span>    role_arn   <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;arn:aws:iam::987654321098:role/DevOpsAgentWorkloadRole&#34;</span>
</span></span><span style="display:flex;"><span>  }
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>In each workload account, create a role that trusts your Agent Space:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-hcl" data-lang="hcl"><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_iam_role&#34; &#34;devops_agent_workload&#34;</span> {
</span></span><span style="display:flex;"><span>  name <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;DevOpsAgentWorkloadRole&#34;</span>
</span></span><span style="display:flex;"><span>  
</span></span><span style="display:flex;"><span>  assume_role_policy <span style="color:#f92672">=</span> <span style="color:#66d9ef">jsonencode</span>({
</span></span><span style="display:flex;"><span>    Principal <span style="color:#f92672">=</span> {
</span></span><span style="display:flex;"><span>      AWS <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;arn:aws:iam::123456789012:role/DevOpsAgentSpaceRole&#34;</span>
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>  })
</span></span><span style="display:flex;"><span>}
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">resource</span> <span style="color:#e6db74">&#34;aws_iam_role_policy_attachment&#34; &#34;read&#34;</span> {
</span></span><span style="display:flex;"><span>  role       <span style="color:#f92672">=</span> <span style="color:#66d9ef">aws_iam_role</span>.<span style="color:#66d9ef">devops_agent_workload</span>.<span style="color:#66d9ef">name</span>
</span></span><span style="display:flex;"><span>  policy_arn <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;arn:aws:iam::aws:policy/ReadOnlyAccess&#34;</span>
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>Connect your monitoring tools (Data API keys, GitHub tokens) through the AWS Console.</p>
<p><strong>Important:</strong> AWS explicitly says Terraform resources may change before GA. Pin your provider versions.</p>
<h2 id="testing-it">Testing it</h2>
<p>AWS provides test scenarios. I recommend running these before connecting production systems.</p>
<p><strong>Test 1: Lambda error investigation</strong></p>
<p>Deploy a Lambda that intentionally throws errors:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#f92672">import</span> random
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">lambda_handler</span>(event, context):
</span></span><span style="display:flex;"><span>    errors <span style="color:#f92672">=</span> [
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;Simulated database timeout&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;Test API rate limit&#34;</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#34;Validation error&#34;</span>
</span></span><span style="display:flex;"><span>    ]
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">raise</span> <span style="color:#a6e22e">Exception</span>(<span style="color:#e6db74">f</span><span style="color:#e6db74">&#34;Test: </span><span style="color:#e6db74">{</span>random<span style="color:#f92672">.</span>choice(errors)<span style="color:#e6db74">}</span><span style="color:#e6db74">&#34;</span>)
</span></span></code></pre></div><p>Create a CloudWatch alarm, trigger it, watch DevOps Agent:</p>
<ul>
<li>Detect the spike</li>
<li>Analyze logs</li>
<li>Check deployment timeline</li>
<li>Identify root cause</li>
<li>Recommend fixes</li>
</ul>
<p><strong>Test 2: EC2 CPU spike</strong></p>
<p>Deploy an EC2 instance, run a CPU stress test, trigger an alarm, watch it correlate with recent changes and recommend auto-scaling.</p>
<h2 id="whats-not-ready-yet-the-honest-limitations">What&rsquo;s not ready yet (the honest limitations)</h2>
<h3 id="1-us-east-1-only">1. us-east-1 only</h3>
<p>DevOps Agent is currently only available in us-east-1.</p>
<p>If you have data residency requirements (GDPR, finance, healthcare), this is a blocker. Cross-region investigations require routing everything through us-east-1.</p>
<p>Mitigation: Deploy Agent Space in us-east-1, use cross-account associations to observe other regions. AWS will probably expand regions post-GA.</p>
<h3 id="2-investigation-vs-action">2. Investigation vs action</h3>
<p>It&rsquo;s unclear whether DevOps Agent can <em>execute</em> remediation or just <em>recommend</em> it.</p>
<p>The documentation emphasizes &ldquo;investigations,&rdquo; &ldquo;recommendations,&rdquo; &ldquo;mitigation suggestions&rdquo;—not &ldquo;auto-rollback&rdquo; or &ldquo;auto-scale.&rdquo;</p>
<p>My read: GA will probably support both:</p>
<ul>
<li><strong>Investigation-only mode (default):</strong> analyze → recommend → human executes</li>
<li><strong>Action mode (opt-in):</strong> execute pre-approved actions within guardrails</li>
</ul>
<p>For regulated industries, you&rsquo;ll live in investigation-only mode. For fast-moving startups, action mode might be tempting.</p>
<h3 id="3-integration-maturity">3. Integration maturity</h3>
<p>Integrations exist for CloudWatch, Datadog, Dynatrace, New Relic, Splunk, GitHub, GitLab, ServiceNow, PagerDuty—but they&rsquo;re first-generation.</p>
<p>Missing:</p>
<ul>
<li>OpenTelemetry native support</li>
<li>ArgoCD, Flux, Spinnaker</li>
<li>Opsgenie, Incident.io</li>
<li>AppDynamics, Elastic APM</li>
</ul>
<p>Good news: Model Context Protocol (MCP) support means you can build custom integrations.</p>
<h3 id="4-learning-curve">4. Learning curve</h3>
<p>DevOps Agent builds its topology map over time. Early investigations might be less accurate.</p>
<p>Mitigation:</p>
<ul>
<li>Run test investigations to let it learn</li>
<li>Tag resources consistently</li>
<li>Document dependencies explicitly</li>
</ul>
<h2 id="should-you-actually-use-this">Should you actually use this?</h2>
<p><strong>Use it if:</strong></p>
<ul>
<li>You&rsquo;re heavily invested in AWS</li>
<li>Your team is drowning in incident response toil</li>
<li>You have operational maturity (monitoring, tagging, CI/CD)</li>
<li>You&rsquo;re comfortable with preview-phase tech</li>
</ul>
<p><strong>Wait if:</strong></p>
<ul>
<li>You need multi-region support now</li>
<li>You require deterministic pricing</li>
<li>Your incident response is already highly optimized</li>
<li>You need production SLAs (preview = no SLAs)</li>
</ul>
<p><strong>Key insight:</strong> DevOps Agent amplifies good practices and exposes bad ones. If your infrastructure is poorly tagged, deployments aren&rsquo;t tracked, and metrics are scattered, it&rsquo;ll struggle. But if you have solid foundations, it can be transformative.</p>
<h2 id="my-honest-take">My honest take</h2>
<p>This is the future of operations. Not because AI replaces engineers, but because it handles undifferentiated heavy lifting.</p>
<p>The question isn&rsquo;t whether agentic operations are coming—they&rsquo;re here. The question is whether you&rsquo;ll be ready when GA drops.</p>
<p>If you&rsquo;re experimenting with this or have questions, <a href="https://www.linkedin.com/in/lucaslittle/">reach out on LinkedIn</a>. The technology is moving fast, and we&rsquo;re all figuring it out together.</p>
<p><strong>Resources:</strong></p>
<ul>
<li><a href="https://docs.aws.amazon.com/devopsagent/latest/userguide/">AWS DevOps Agent User Guide</a></li>
<li><a href="https://github.com/aws-samples/sample-aws-devops-agent-terraform">Terraform Sample Repo</a></li>
<li><a href="https://aws.amazon.com/ai/frontier-agents">AWS Frontier Agents Overview</a></li>
</ul>
]]></content:encoded></item><item><title>Data Pour with Dr. Mohamed Shehab: Build, Build, Build</title><link>https://lukelittle.com/posts/2026/01/data-pour-with-dr.-mohamed-shehab-build-build-build/</link><pubDate>Fri, 09 Jan 2026 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2026/01/data-pour-with-dr.-mohamed-shehab-build-build-build/</guid><description>A conversation at UNC Charlotte&amp;#39;s PORTAL about why students love his mobile dev course, embracing AI in education, and his advice for breaking into tech: build.</description><content:encoded><![CDATA[<p>For this episode, I visited the PORTAL building at UNC Charlotte to sit down with Dr. Mohamed Shehab—a professor whose mobile development course keeps showing up in conversations with students as one of the most impactful experiences of their degree.</p>
<p>I&rsquo;ve heard it repeatedly from students I&rsquo;ve hired: &ldquo;Dr. Shehab&rsquo;s course changed how I think about building software.&rdquo; When you hear that kind of feedback consistently, you have to ask what he&rsquo;s doing differently.</p>
<h2 id="why-students-remember-this-course">Why students remember this course</h2>
<p>I originally connected with Dr. Shehab after noticing the pattern. Current students, recent grads, people years into their careers—they all talked about his mobile dev class the way you talk about a course that actually stuck.</p>
<p>The answer is pretty straightforward: it&rsquo;s hands-on, industry-aligned, and constantly evolving.</p>
<p>Dr. Shehab started teaching mobile development in 2008, right when the iPhone App Store launched. He got lucky with timing, but what kept the course relevant was the approach: flip the lecture material to video, use class time for hands-on work, teach students to ship working apps—not just pass exams.</p>
<p>The tech changes constantly. Updates every few months. So the course has to change with it. Students learn how to build, deploy, and debug real applications. They encounter the same decision-making tradeoffs developers face in production.</p>
<p>That&rsquo;s the alignment. Students walk out with job-ready experience, not just academic knowledge.</p>
<h2 id="embracing-ai-and-changing-how-you-evaluate">Embracing AI (and changing how you evaluate)</h2>
<p>We spent a lot of time talking about AI in the classroom. Dr. Shehab&rsquo;s take is pragmatic: the technology is here, students are using it, developers are using it—you have to adapt.</p>
<p>But it creates real challenges for educators. How do you evaluate student understanding when they can generate code with a chatbot?</p>
<p>Dr. Shehab&rsquo;s approach: change how you evaluate. Don&rsquo;t just check if the app runs—have students demo it and explain how it works. Ask questions. Make them walk through their code. If they used GPT to build it but can understand and articulate the solution, that&rsquo;s acceptable. If they can&rsquo;t explain it, they get half credit.</p>
<p>He&rsquo;s also experimenting with GitHub Copilot in the development environment—teaching students to use AI as an assistive tool, not a replacement for understanding.</p>
<p>The goal isn&rsquo;t to block AI. It&rsquo;s to teach students how to use it effectively while still building foundational skills.</p>
<h2 id="build-build-build">Build, build, build</h2>
<p>I asked what advice he gives students trying to break into the job market—especially when entry-level roles feel harder to land and AI is changing expectations.</p>
<p>His answer was clear: <strong>build things</strong>.</p>
<p>Don&rsquo;t just do assignments. Build real projects. Put them on GitHub. Deploy them. Show your work.</p>
<p>The biggest challenge right now is for early talent. If you&rsquo;re graduating and want to stand out, you need a strong portfolio. You need to show up to employers with something you&rsquo;ve built—not expecting them to hold your hand and train you from scratch.</p>
<p>Communication matters too. AI can do a lot of things, but it can&rsquo;t be you. Being a good collaborator, explaining things clearly, presenting confidently—those skills still matter.</p>
<p>And get involved outside the classroom. Go to meetups. Join student orgs. Use the maker spaces. Network. Ask questions at recruiting events—even if you think the question is stupid. Have presence.</p>
<p>Dr. Shehab&rsquo;s motto: <strong>build, build, build</strong>. No hand-waving. Features built = points. No features = no points.</p>
<h2 id="the-maker-space-movement">The maker space movement</h2>
<p>One of the more interesting parts of our conversation was the emphasis on maker spaces. UNC Charlotte has invested heavily—3D printers, CNC machines, fabrication labs. Students can access them for free.</p>
<p>Dr. Shehab&rsquo;s point: the university isn&rsquo;t going to teach you everything in the classroom. A lot of learning happens outside—in maker spaces, student orgs, side projects. Open yourself to other experiences. Print something. Build something. See what&rsquo;s possible.</p>
<p>When companies interview students and see that breadth of experience—someone who&rsquo;s not just focused on one narrow skill but has built, experimented, and solved real problems—that opens doors.</p>
<h2 id="building-stronger-industry-partnerships">Building stronger industry partnerships</h2>
<p>We also talked about what industry can do to be better partners with academia.</p>
<p>More collaboration. More joint projects. More events on campus exposing students to the ecosystem. More feedback to faculty on what skills actually matter in the market.</p>
<p>Dr. Shehab mentioned that UNC Charlotte is exploring focused talent pipelines—where companies have input on curriculum and certificate programs. Containerization certificates. Visualization certificates. AI certificates. Programs designed not just for traditional students but for employees or potential employees.</p>
<p>The challenge is cultural. A lot of academics are traditional: teach a class, do research, go home. But exposing faculty to industry—and integrating those relationships—can boost relevance for everyone.</p>
<h2 id="why-i-wanted-him-on-the-show">Why I wanted him on the show</h2>
<p>I&rsquo;m incredibly grateful to Dr. Shehab for making time to be part of The Data Pour. Seeing the level of impact he&rsquo;s had on students is what motivated me to get more involved at UNC Charlotte—to help create stronger pathways between academia and industry.</p>
<p>I&rsquo;m excited to continue working together in the future.</p>
<h2 id="watch-the-episode">Watch the episode</h2>
<p>If you want the full conversation—including the parts on the PORTAL incubator, startup culture, and what Dr. Shehab sees as the future of CS education—watch the YouTube episode here:</p>
<p><a href="https://www.youtube.com/watch?v=vZOze-a_JTU&amp;t=564s">https://www.youtube.com/watch?v=vZOze-a_JTU&amp;t=564s</a></p>
]]></content:encoded></item><item><title>Pokémon Surveys, Serverless Architecture, and Teaching Students to Build on AWS</title><link>https://lukelittle.com/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/</link><pubDate>Sat, 27 Dec 2025 21:30:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/</guid><description>Creating a small survey app to teach the cloud and explain its power.</description><content:encoded><![CDATA[<p>Back in November, I was preparing for the Cracking the Cloud presentation at UNC Charlotte. I needed a way to explain how the cloud fundamentally changed what&rsquo;s possible on the internet—not through abstract concepts, but through something students could immediately relate to.</p>
<p>That&rsquo;s when I remembered <strong>Thomas Game Docs</strong>.</p>
<p>If you&rsquo;ve never heard of her: she&rsquo;s a YouTuber who makes incredibly well-produced video essays about video games. And she sometimes runs surveys asking her audience things like &ldquo;Who&rsquo;s the LEAST popular Pokémon?&rdquo; or &ldquo;Who&rsquo;s the LEAST popular Animal Crossing villager?&rdquo;</p>
<p>These aren&rsquo;t small surveys. They get millions of responses.</p>
<p>I helped with the backend for the Pokémon survey—a Flask app on Heroku. But the technology stack wasn&rsquo;t the interesting part. What mattered was that hosting something like this <strong>doesn&rsquo;t require infrastructure expertise anymore</strong>.</p>
<p>She didn&rsquo;t need to buy servers, configure databases, or hire a DevOps team.</p>
<p>Twenty years ago, hosting a survey that could handle 50,000 votes meant:</p>
<ul>
<li>Buy or rent physical servers</li>
<li>Set up database infrastructure</li>
<li>Configure load balancers</li>
<li>Plan for capacity (and hope you got it right)</li>
<li>Deal with outages, scaling issues, and hardware failures</li>
</ul>
<p>All of that would cost thousands of dollars—and that&rsquo;s before you wrote a single line of code.</p>
<p>Today? You build it with Lambda, API Gateway, and DynamoDB. You deploy it with Terraform. And unless traffic gets truly ridiculous, <strong>it costs you basically nothing</strong>.</p>
<p>The cloud didn&rsquo;t just make infrastructure cheaper. It made building things <strong>accessible</strong>.</p>
<p>That&rsquo;s what I wanted students to understand. Not that AWS has a lot of services. But that those services remove the barriers that used to keep people from building.</p>
<hr>
<h2 id="the-demo-a-survey-students-could-actually-participate-in">The demo: a survey students could actually participate in</h2>
<p><img src="/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/pokemon-surveys-and-cloud-infrastructure_hu_f14e3c0c47ba9455.webp" srcset="/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/pokemon-surveys-and-cloud-infrastructure_hu_eb463a26a2e45749.webp 750w, /posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/pokemon-surveys-and-cloud-infrastructure_hu_f14e3c0c47ba9455.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="842" alt="Pokémon Surveys and Cloud Infrastructure" loading="lazy" decoding="async"></p>
<p>To drive the point home, I didn&rsquo;t just talk about Thomas Game Docs surveys.</p>
<p>I had the students take one.</p>
<p>At the start of the presentation, I pulled up a simple survey asking about their exposure to AWS:</p>
<ul>
<li>Have you used AWS before?</li>
<li>Have you deployed something to the cloud?</li>
</ul>
<p>They voted. They saw the results update in real-time. And then I showed them <strong>exactly how it worked</strong>—with no servers running, no databases to manage, and no ongoing costs to worry about.</p>
<p>That survey? It&rsquo;s the same repo I&rsquo;m writing about now: <a href="https://github.com/lukelittle/cracking-the-cloud">cracking-the-cloud</a>.</p>
<h2 id="the-architecture-small-but-real">The architecture (small, but real)</h2>
<p><img src="/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/diagram_hu_67a3240d858eed62.webp" srcset="/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/diagram_hu_811848a8f1ebf8fa.webp 750w, /posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/diagram_hu_67a3240d858eed62.webp 1500w" sizes="(max-width: 800px) 100vw, 750px"
       width="1500" height="932" alt="Architecture Diagram" loading="lazy" decoding="async"></p>
<p>Here&rsquo;s the final shape of the system:</p>
<ul>
<li><strong>S3</strong> hosts the static frontend (HTML, CSS, JS)</li>
<li><strong>CloudFront</strong> sits in front for HTTPS, caching, and global delivery</li>
<li><strong>API Gateway</strong> exposes a REST API</li>
<li><strong>Lambda</strong> handles business logic (vote, results, reset)</li>
<li><strong>DynamoDB</strong> stores votes</li>
<li><strong>IAM</strong> wires permissions together</li>
<li><strong>Terraform</strong> defines everything</li>
</ul>
<p>No servers. No containers. No databases to patch. No stateful nonsense.</p>
<p>Just managed services doing exactly what they&rsquo;re good at.</p>
<p>The request flow looks like this:</p>
<p>User clicks a button → JavaScript calls the API → API Gateway invokes Lambda → Lambda writes to DynamoDB → Response goes back to the browser.</p>
<p>Simple. Explicit. Observable.</p>
<h2 id="why-static-frontend--api-on-purpose">Why static frontend + API (on purpose)</h2>
<p>I didn&rsquo;t use React.<br>
I didn&rsquo;t use Next.js.<br>
I didn&rsquo;t use server-side rendering.</p>
<p>Not because those tools are bad—they&rsquo;re not. But because I wanted you to see what&rsquo;s actually happening.</p>
<p>When you open the frontend code, you can immediately see:</p>
<ul>
<li>where the API URL lives</li>
<li>how a POST request is formed</li>
<li>what the response looks like</li>
<li>how the browser handles the data</li>
</ul>
<p>No build steps. No transpilation. No abstractions hiding what&rsquo;s really going on.</p>
<p>Once you understand how a browser talks to an API using vanilla JavaScript, <strong>then</strong> you can add React, TypeScript, and all the modern tooling. But you&rsquo;ll know what those tools are doing for you—not just that they work.</p>
<p>The frontend&rsquo;s job here is to show you the fundamentals, not teach you the latest framework.</p>
<h2 id="the-lambdas-three-on-purpose">The Lambdas (three, on purpose)</h2>
<p>There are three Lambda functions:</p>
<ol>
<li><strong>Vote</strong> (<code>backend/vote.py</code>) – Processes vote submissions</li>
<li><strong>Results</strong> (<code>backend/results.py</code>) – Retrieves vote counts</li>
<li><strong>Reset</strong> (<code>backend/reset.py</code>) – Clears all data</li>
</ol>
<p>Could this be one Lambda with a switch statement?<br>
Absolutely.</p>
<p>Did I do that?<br>
Absolutely not.</p>
<p>Each function has:</p>
<ul>
<li>one responsibility</li>
<li>one API route</li>
<li>one IAM policy</li>
</ul>
<p>This lets students see how permissions map to behavior.</p>
<ul>
<li>The <strong>vote</strong> function can write, but not delete</li>
<li>The <strong>results</strong> function can read, but not write</li>
<li>The <strong>reset</strong> function can delete, but nothing else</li>
</ul>
<p>You don&rsquo;t need a lecture on least privilege when the code makes it obvious.</p>
<h2 id="how-voting-actually-works">How voting actually works</h2>
<p>When a student clicks a vote button, here&rsquo;s the journey that request takes through the serverless stack:</p>
<pre class="mermaid">sequenceDiagram
    participant User as 👤 User Browser
    participant S3 as 🪣 S3 + CloudFront
    participant APIG as 🌐 API Gateway
    participant Lambda as ⚡ vote.py
    participant DDB as 🗄️ DynamoDB
    
    User-&gt;&gt;S3: GET /vote.html
    S3--&gt;&gt;User: HTML + JavaScript
    
    Note over User: User clicks vote button&lt;br/&gt;sessionId generated (UUID)&lt;br/&gt;stored in sessionStorage
    
    User-&gt;&gt;APIG: POST /vote&lt;br/&gt;{ &#34;vote&#34;: &#34;aws&#34;, &#34;sessionId&#34;: &#34;abc123&#34; }
    APIG-&gt;&gt;Lambda: Invoke vote function
    
    Note over Lambda: Validate sessionId exists&lt;br/&gt;Validate vote in [&#39;no&#39;, &#39;aws&#39;, &#39;other&#39;]
    
    Lambda-&gt;&gt;DDB: PutItem&lt;br/&gt;{ id: &#34;abc123&#34;, vote: &#34;aws&#34; }
    Note over DDB: Overwrites if sessionId&lt;br/&gt;already voted&lt;br/&gt;(allows vote changes)
    DDB--&gt;&gt;Lambda: Success
    
    Lambda--&gt;&gt;APIG: 200 OK&lt;br/&gt;{ &#34;message&#34;: &#34;Vote recorded&#34; }
    APIG--&gt;&gt;User: Response
    
    Note over User: JavaScript displays&lt;br/&gt;&#34;Vote recorded!&#34; message
</pre>

<p>The critical piece here is the <strong>sessionId</strong>. It&rsquo;s a random UUID generated in the browser and stored in <code>sessionStorage</code>—which means it persists for the current tab but disappears when you close the browser.</p>
<p>This gives us:</p>
<ul>
<li><strong>One vote per browser session</strong> – you can&rsquo;t spam-click the vote button</li>
<li><strong>Vote changes allowed</strong> – if you vote &ldquo;No experience&rdquo; and change your mind, the second vote overwrites the first (DynamoDB&rsquo;s <code>PutItem</code> does this automatically)</li>
<li><strong>Privacy by default</strong> – no accounts, no tracking, no persistent identifiers</li>
<li><strong>Simple anti-spam</strong> – good enough for a teaching demo</li>
</ul>
<p>Could someone bypass this by opening incognito windows? Yes. Is that fine for a teaching app? Also yes. The point is showing how to prevent duplicate votes, not building a production election system.</p>
<p>Here&rsquo;s what <code>vote.py</code> actually looks like:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Parse the incoming request</span>
</span></span><span style="display:flex;"><span>    body <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(event<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;body&#39;</span>, <span style="color:#e6db74">&#39;</span><span style="color:#e6db74">{}</span><span style="color:#e6db74">&#39;</span>))
</span></span><span style="display:flex;"><span>    vote_option <span style="color:#f92672">=</span> body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;vote&#39;</span>)
</span></span><span style="display:flex;"><span>    session_id <span style="color:#f92672">=</span> body<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;sessionId&#39;</span>)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Validate inputs</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> <span style="color:#f92672">not</span> session_id:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">400</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;message&#39;</span>: <span style="color:#e6db74">&#39;Session ID is required&#39;</span>})
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> vote_option <span style="color:#f92672">not</span> <span style="color:#f92672">in</span> [<span style="color:#e6db74">&#39;no&#39;</span>, <span style="color:#e6db74">&#39;aws&#39;</span>, <span style="color:#e6db74">&#39;other&#39;</span>]:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">400</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;message&#39;</span>: <span style="color:#e6db74">&#39;Invalid vote option&#39;</span>})
</span></span><span style="display:flex;"><span>        }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Store the vote</span>
</span></span><span style="display:flex;"><span>    table<span style="color:#f92672">.</span>put_item(Item<span style="color:#f92672">=</span>{
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;id&#39;</span>: session_id,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;vote&#39;</span>: vote_option
</span></span><span style="display:flex;"><span>    })
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;headers&#39;</span>: {<span style="color:#e6db74">&#39;Access-Control-Allow-Origin&#39;</span>: <span style="color:#e6db74">&#39;*&#39;</span>},
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;message&#39;</span>: <span style="color:#e6db74">&#39;Vote recorded successfully&#39;</span>})
</span></span><span style="display:flex;"><span>    }
</span></span></code></pre></div><p>Twenty lines of code. No ORM. No database migrations. No connection pooling. Just write to DynamoDB and return a response.</p>
<h2 id="how-results-actually-work">How results actually work</h2>
<p>The results page is where students first encounter the concept of scanning a database:</p>
<pre class="mermaid">sequenceDiagram
    participant User as 👤 User Browser
    participant S3 as 🪣 S3 + CloudFront
    participant APIG as 🌐 API Gateway
    participant Lambda as ⚡ results.py
    participant DDB as 🗄️ DynamoDB
    
    User-&gt;&gt;S3: GET /results.html
    S3--&gt;&gt;User: HTML + JavaScript + Chart.js
    
    Note over User: Page loads&lt;br/&gt;JavaScript calls API
    
    User-&gt;&gt;APIG: GET /results
    APIG-&gt;&gt;Lambda: Invoke results function
    
    Lambda-&gt;&gt;DDB: Scan table&lt;br/&gt;ProjectionExpression=&#39;vote&#39;
    Note over DDB: Returns all vote values&lt;br/&gt;[&#39;aws&#39;, &#39;no&#39;, &#39;aws&#39;, &#39;other&#39;, ...]
    
    DDB--&gt;&gt;Lambda: Page 1 of results
    
    Note over Lambda: Check for LastEvaluatedKey&lt;br/&gt;(pagination if table &gt; 1MB)
    
    loop While LastEvaluatedKey exists
        Lambda-&gt;&gt;DDB: Scan with ExclusiveStartKey
        DDB--&gt;&gt;Lambda: Next page of results
    end
    
    Note over Lambda: Count votes using Counter&lt;br/&gt;{ &#39;no&#39;: 15, &#39;aws&#39;: 42, &#39;other&#39;: 8 }
    
    Lambda--&gt;&gt;APIG: 200 OK&lt;br/&gt;{ &#34;no&#34;: 15, &#34;aws&#34;: 42, &#34;other&#34;: 8 }
    APIG--&gt;&gt;User: Response
    
    Note over User: Chart.js renders&lt;br/&gt;vote counts as bar chart
</pre>

<p>The interesting part here is <strong>pagination</strong>. DynamoDB&rsquo;s <code>Scan</code> operation returns a maximum of 1MB of data per request. If your table is larger than that, you get a <code>LastEvaluatedKey</code> in the response, which you use to fetch the next page.</p>
<p>Here&rsquo;s what that looks like in code:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># First scan</span>
</span></span><span style="display:flex;"><span>    response <span style="color:#f92672">=</span> table<span style="color:#f92672">.</span>scan(ProjectionExpression<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;vote&#39;</span>)
</span></span><span style="display:flex;"><span>    items <span style="color:#f92672">=</span> response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;Items&#39;</span>, [])
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Keep scanning if there&#39;s more data</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">while</span> <span style="color:#e6db74">&#39;LastEvaluatedKey&#39;</span> <span style="color:#f92672">in</span> response:
</span></span><span style="display:flex;"><span>        response <span style="color:#f92672">=</span> table<span style="color:#f92672">.</span>scan(
</span></span><span style="display:flex;"><span>            ProjectionExpression<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;vote&#39;</span>,
</span></span><span style="display:flex;"><span>            ExclusiveStartKey<span style="color:#f92672">=</span>response[<span style="color:#e6db74">&#39;LastEvaluatedKey&#39;</span>]
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        items<span style="color:#f92672">.</span>extend(response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;Items&#39;</span>, []))
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Extract just the vote values</span>
</span></span><span style="display:flex;"><span>    votes <span style="color:#f92672">=</span> [item[<span style="color:#e6db74">&#39;vote&#39;</span>] <span style="color:#66d9ef">for</span> item <span style="color:#f92672">in</span> items]
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Count them</span>
</span></span><span style="display:flex;"><span>    vote_counts <span style="color:#f92672">=</span> Counter(votes)
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Return with defaults for zero-vote options</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;headers&#39;</span>: {<span style="color:#e6db74">&#39;Access-Control-Allow-Origin&#39;</span>: <span style="color:#e6db74">&#39;*&#39;</span>},
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;no&#39;</span>: vote_counts<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;no&#39;</span>, <span style="color:#ae81ff">0</span>),
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;aws&#39;</span>: vote_counts<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;aws&#39;</span>, <span style="color:#ae81ff">0</span>),
</span></span><span style="display:flex;"><span>            <span style="color:#e6db74">&#39;other&#39;</span>: vote_counts<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;other&#39;</span>, <span style="color:#ae81ff">0</span>)
</span></span><span style="display:flex;"><span>        })
</span></span><span style="display:flex;"><span>    }
</span></span></code></pre></div><p>Students immediately see three concepts:</p>
<ul>
<li><strong>Scanning costs</strong> – you&rsquo;re reading the entire table, which is fine for 100 votes but would be expensive for 10 million</li>
<li><strong>Pagination handling</strong> – real-world data doesn&rsquo;t fit in one response</li>
<li><strong>Aggregation happens in code</strong> – DynamoDB doesn&rsquo;t have <code>COUNT(*) GROUP BY vote</code>, so you pull the data and count it yourself</li>
</ul>
<p>This naturally leads to questions like &ldquo;how would you make this more efficient?&rdquo; which is exactly where you want students&rsquo; brains to go.</p>
<h2 id="how-reset-actually-works">How reset actually works</h2>
<p>The reset function is the most dangerous one in the app—and also the most instructive:</p>
<pre class="mermaid">sequenceDiagram
    participant User as 👤 User Browser
    participant S3 as 🪣 S3 + CloudFront
    participant APIG as 🌐 API Gateway
    participant Lambda as ⚡ reset.py
    participant DDB as 🗄️ DynamoDB
    
    User-&gt;&gt;S3: GET /reset.html
    S3--&gt;&gt;User: HTML + JavaScript
    
    Note over User: User clicks&lt;br/&gt;&#34;Reset All Votes&#34; button&lt;br/&gt;(⚠️ Destructive operation)
    
    User-&gt;&gt;APIG: POST /reset
    APIG-&gt;&gt;Lambda: Invoke reset function
    
    Lambda-&gt;&gt;DDB: Scan table&lt;br/&gt;ProjectionExpression=&#39;id&#39;
    Note over DDB: Only return IDs&lt;br/&gt;(need keys to delete)
    DDB--&gt;&gt;Lambda: All item IDs
    
    loop While LastEvaluatedKey exists
        Lambda-&gt;&gt;DDB: Scan with ExclusiveStartKey
        DDB--&gt;&gt;Lambda: More IDs
    end
    
    Note over Lambda: Batch delete in groups of 25&lt;br/&gt;(DynamoDB batch write limit)
    
    loop For each batch of 25 items
        Lambda-&gt;&gt;DDB: BatchWriteItem&lt;br/&gt;Delete items
        DDB--&gt;&gt;Lambda: Batch delete success
    end
    
    Note over Lambda: All items deleted&lt;br/&gt;Table is now empty
    
    Lambda--&gt;&gt;APIG: 200 OK&lt;br/&gt;{ &#34;message&#34;: &#34;Survey reset&#34; }
    APIG--&gt;&gt;User: Response
    
    Note over User: &#34;All votes deleted!&#34;
</pre>

<p>This introduces batch operations:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">handler</span>(event, context):
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Scan for all IDs (we only need keys to delete)</span>
</span></span><span style="display:flex;"><span>    scan_response <span style="color:#f92672">=</span> table<span style="color:#f92672">.</span>scan(ProjectionExpression<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;id&#39;</span>)
</span></span><span style="display:flex;"><span>    items <span style="color:#f92672">=</span> scan_response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;Items&#39;</span>, [])
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Handle pagination</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">while</span> <span style="color:#e6db74">&#39;LastEvaluatedKey&#39;</span> <span style="color:#f92672">in</span> scan_response:
</span></span><span style="display:flex;"><span>        scan_response <span style="color:#f92672">=</span> table<span style="color:#f92672">.</span>scan(
</span></span><span style="display:flex;"><span>            ProjectionExpression<span style="color:#f92672">=</span><span style="color:#e6db74">&#39;id&#39;</span>,
</span></span><span style="display:flex;"><span>            ExclusiveStartKey<span style="color:#f92672">=</span>scan_response[<span style="color:#e6db74">&#39;LastEvaluatedKey&#39;</span>]
</span></span><span style="display:flex;"><span>        )
</span></span><span style="display:flex;"><span>        items<span style="color:#f92672">.</span>extend(scan_response<span style="color:#f92672">.</span>get(<span style="color:#e6db74">&#39;Items&#39;</span>, []))
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e"># Delete all items using batch writer</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">if</span> items:
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">with</span> table<span style="color:#f92672">.</span>batch_writer() <span style="color:#66d9ef">as</span> batch:
</span></span><span style="display:flex;"><span>            <span style="color:#66d9ef">for</span> item <span style="color:#f92672">in</span> items:
</span></span><span style="display:flex;"><span>                batch<span style="color:#f92672">.</span>delete_item(Key<span style="color:#f92672">=</span>{<span style="color:#e6db74">&#39;id&#39;</span>: item[<span style="color:#e6db74">&#39;id&#39;</span>]})
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;statusCode&#39;</span>: <span style="color:#ae81ff">200</span>,
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;headers&#39;</span>: {<span style="color:#e6db74">&#39;Access-Control-Allow-Origin&#39;</span>: <span style="color:#e6db74">&#39;*&#39;</span>},
</span></span><span style="display:flex;"><span>        <span style="color:#e6db74">&#39;body&#39;</span>: json<span style="color:#f92672">.</span>dumps({<span style="color:#e6db74">&#39;message&#39;</span>: <span style="color:#e6db74">&#39;Survey reset successfully&#39;</span>})
</span></span><span style="display:flex;"><span>    }
</span></span></code></pre></div><p>The <code>batch_writer()</code> context manager is doing a lot of hidden work:</p>
<ul>
<li>Groups deletes into batches of 25 (DynamoDB&rsquo;s limit)</li>
<li>Automatically retries failed operations</li>
<li>Handles throttling gracefully</li>
<li>Only commits when the context exits</li>
</ul>
<p>Students don&rsquo;t need to know all of that on day one, but they can see that deleting 100 items doesn&rsquo;t require 100 API calls.</p>
<p><strong>Important note:</strong> In a real app, you&rsquo;d absolutely add authentication and authorization here. This function is intentionally unprotected for teaching purposes—it demonstrates the mechanics of batch operations without the complexity of auth flows.</p>
<h2 id="dynamodb-intentionally-unsexy">DynamoDB (intentionally unsexy)</h2>
<p>The DynamoDB table is <strong>boring by design</strong>.</p>
<ul>
<li>Partition key (<code>id</code>)</li>
<li>Simple attributes (<code>vote</code>)</li>
<li>No GSIs</li>
<li>No streams</li>
<li>No TTL magic</li>
</ul>
<p>Why?</p>
<p>Because the lesson isn&rsquo;t &ldquo;DynamoDB is infinite and weird.&rdquo;<br>
The lesson is:</p>
<p><strong>You can persist state without running a database.</strong></p>
<p>Once students are comfortable, <em>then</em> you add:</p>
<ul>
<li>secondary indexes</li>
<li>conditional writes</li>
<li>access patterns</li>
<li>cost modeling</li>
</ul>
<p>But not on day one.</p>
<h2 id="terraform-as-the-real-curriculum">Terraform as the real curriculum</h2>
<p>Here&rsquo;s the quiet truth: <strong>The Terraform is the most important part of this project.</strong></p>
<p>Students don&rsquo;t learn AWS by clicking around the console. They learn AWS by reading infrastructure definitions and realizing: <em>&ldquo;Oh—that&rsquo;s what connects to that.&rdquo;</em></p>
<p>This repo forces them to see how API Gateway connects to Lambda (<code>aws_api_gateway_integration</code>), how Lambda permissions work (<code>aws_lambda_permission</code>), how CloudFront talks to S3 (<code>origin_access_identity</code>), how outputs become frontend configuration (<code>cloudfront_domain</code>).</p>
<p>They can delete everything and recreate it in minutes. That alone teaches more than most cloud courses.</p>
<hr>
<h2 id="how-does-this-compare-to-the-heroku-version">How does this compare to the Heroku version?</h2>
<p>Remember the Flask app I built for Thomas Game Docs&rsquo; Pokémon survey?</p>
<p>I monitored that deployment closely. Every time she dropped an announcement on social media—Twitter, YouTube community posts, etc.—I watched the app response time blow up. We were constantly aware that we were one viral tweet away from needing to manually scale the Heroku dyno or upgrade the database.</p>
<p>With this serverless version? <strong>There wouldn&rsquo;t have been a hiccup.</strong></p>
<p>Lambda would have spun up as many concurrent executions as needed. API Gateway would have handled the traffic without breaking a sweat. DynamoDB would have throttled gracefully and auto-scaled. CloudFront would have cached the static assets globally.</p>
<p>No monitoring dashboards.<br>
No capacity planning.<br>
No &ldquo;should we upgrade now or wait?&rdquo; decisions.<br>
No watching metrics at 2 AM when a post goes viral.</p>
<p>The infrastructure would have scaled to meet demand and then scaled back down when traffic dropped. And the bill would have stayed under $5 for the entire campaign.</p>
<p>That&rsquo;s the difference between &ldquo;serverless&rdquo; and &ldquo;server-you-manage-less.&rdquo;</p>
<h2 id="costs-because-someone-always-asks">Costs (because someone always asks)</h2>
<ul>
<li><strong>S3</strong>: pennies</li>
<li><strong>CloudFront</strong>: free tier</li>
<li><strong>Lambda</strong>: free tier</li>
<li><strong>API Gateway</strong>: free tier</li>
<li><strong>DynamoDB</strong>: free tier</li>
</ul>
<p><strong>Total monthly cost for light usage</strong>: effectively $0.</p>
<p>Which matters, because students shouldn&rsquo;t need a credit card panic attack to learn cloud fundamentals.</p>
<h2 id="if-you-want-to-fork-it">If you want to fork it</h2>
<p>The entire project is open source and designed to be broken, modified, and rebuilt:</p>
<p><strong><a href="https://github.com/lukelittle/cracking-the-cloud">github.com/lukelittle/cracking-the-cloud</a></strong></p>
<hr>
<h2 id="how-this-project-can-be-used-to-teach-cloud">How this project can be used to teach cloud</h2>
<p>The beauty of this baseline is that every extension becomes a teaching moment. The answer to nearly every &ldquo;Can I add&hellip;?&rdquo; question is: <strong>yes</strong>.</p>
<p><strong>&ldquo;Can I add another survey question?&rdquo;</strong> Yes. Modify the DynamoDB schema and update the frontend. Students learn about schema evolution and backwards compatibility.</p>
<p><strong>&ldquo;Can I add authentication?&rdquo;</strong> Yes. Add Cognito, modify the Lambda to verify JWT tokens, update the frontend to handle login flows. Students learn about identity providers, token validation, and authorization.</p>
<p><strong>&ldquo;Can I track who voted when?&rdquo;</strong> Yes. Add timestamps to DynamoDB items, maybe stream changes to S3 for analytics. Students learn about audit trails and data retention.</p>
<p><strong>&ldquo;Can I swap DynamoDB for RDS?&rdquo;</strong> Yes. But now you need VPCs, security groups, connection pooling, and Lambda cold start considerations. Students learn why DynamoDB was the right choice for this use case.</p>
<p><strong>&ldquo;Can I add email notifications?&rdquo;</strong> Yes. Give a Lambda permission to use SES, trigger it from DynamoDB Streams. Students learn about event-driven architecture and service integration.</p>
<p><strong>&ldquo;Can I add a CI/CD pipeline?&rdquo;</strong> Yes. Add GitHub Actions, IAM roles with OIDC, and S3 sync logic. Students learn about deployment automation and security best practices.</p>
<p>Because the baseline is so small, every addition is visible. Every new service has a before-and-after moment. This is where the app stops being a demo and starts being a scaffold—students can extend it in any direction and immediately see what changes.</p>
<p>Change the question. Add auth. Add metrics. Rip it apart. That&rsquo;s how you build instincts. Not by memorizing services, but by wiring them together and watching what happens.</p>
]]></content:encoded></item><item><title>Data Pour with Lucas Ward: Consulting, People Leadership, and Staying Curious in the AI Era</title><link>https://lukelittle.com/posts/2025/12/data-pour-with-lucas-ward-consulting-people-leadership-and-staying-curious-in-the-ai-era/</link><pubDate>Tue, 16 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/data-pour-with-lucas-ward-consulting-people-leadership-and-staying-curious-in-the-ai-era/</guid><description>A conversation with Lucas Ward at Charlotte Beer Garden about leading people in consulting, enterprise vs midsize engineering, and AI readiness.</description><content:encoded><![CDATA[<p>For this episode of Data Pour, I sat down with Lucas Ward—one of my SWE managers and a senior technical manager at Ippon—to talk about what it actually looks like to lead people in a consulting practice while the industry is shifting under our feet.</p>
<p>We filmed at the Charlotte Beer Garden, which feels like the most Charlotte setting possible: more taps than you can count, plus the steady soundtrack of loud cars and motorcycles rolling by. We both went with an OMB beer (Mecktoberfest) because… you kind of have to.</p>
<p>If you want the full conversation, watch the YouTube episode—this post is the background and the highlights.</p>
<h2 id="why-i-wanted-lucas-on-the-show">Why I wanted Lucas on the show</h2>
<p>Lucas isn&rsquo;t just a strong engineer—he&rsquo;s the kind of leader who can translate. In consulting, that&rsquo;s the job: you&rsquo;re constantly moving between client expectations, delivery realities, and the growth of the people on your team.</p>
<p>And Lucas has lived both sides.</p>
<p>He broke into tech the hard way—after years in restaurants—then got his start at a small Charlotte business where he had to wear every hat imaginable: shipping code, running systems in production, taking late-night calls when something broke, and learning by doing because there wasn&rsquo;t anyone else to do it.</p>
<p>That &ldquo;wear every hat&rdquo; foundation shows up in how he leads today.</p>
<h2 id="people-leadership-in-consulting-is-a-different-sport">People leadership in consulting is a different sport</h2>
<p>One of the most real parts of our conversation was the dynamic of managing people when you aren&rsquo;t sitting next to them every day.</p>
<p>In a traditional org, you see the work constantly. In consulting, your direct report might be on another client, another team, another tech lead. Sometimes you&rsquo;re getting signal through check-ins and feedback loops rather than direct observation—which makes coaching and performance conversations trickier (and honestly, more important to handle thoughtfully).</p>
<p>Lucas talked about what it means to guide technical folks who are laser-focused on hard skills—while also helping them build the soft skills that determine whether they actually grow: communication, delivery, stakeholder management, and the ability to explain why the work matters.</p>
<p>Technical skills get you a seat at the table. The rest is what keeps you there.</p>
<h2 id="consultant--solve-the-problem-and-explain-the-value">&ldquo;Consultant&rdquo; = solve the problem and explain the value</h2>
<p>Lucas gave a definition I loved:</p>
<p>A consultant is someone who can make the square peg fit in the round hole—and then explain how they did it to the people who care about the peg fitting.</p>
<p>That&rsquo;s the work. Not just building. Translating. Connecting technical decisions to business outcomes. Helping stakeholders understand what&rsquo;s happening and why it matters.</p>
<h2 id="enterprise-vs-midsize-dont-write-off-the-breadth-engineers">Enterprise vs midsize: don&rsquo;t write off the &ldquo;breadth&rdquo; engineers</h2>
<p>We also dug into how midsize clients operate differently than large enterprises—and what each can learn from the other.</p>
<p>Enterprises have specialization and depth. Midsize companies often have engineers with ridiculous breadth because they&rsquo;ve had to take products through the entire lifecycle—build, deploy, operate, support.</p>
<p>Lucas&rsquo;s advice to enterprise hiring managers was simple: don&rsquo;t dismiss candidates from smaller companies. They may not have the same &ldquo;one deep slice&rdquo; experience, but they often bring systems thinking and end-to-end ownership that&rsquo;s hard to teach.</p>
<h2 id="ai-readiness-everybody-wants-ai-but-not-everyone-is-ready">AI readiness: everybody wants AI, but not everyone is ready</h2>
<p>This came up a lot:</p>
<p>Large enterprises are thinking about data strategy, fine-tuning, governance, and ROI.</p>
<p>Midsize companies often want AI, but their data is siloed, inconsistent, and not operationally prepared for &ldquo;real&rdquo; AI use cases.</p>
<p>And Lucas made a key point: AI readiness is incremental—and most of what you do to become &ldquo;AI ready&rdquo; is just good engineering anyway. Clean data, lineage, security controls, sustainable platforms, well-architected foundations. AI just forces the conversation.</p>
<p>Also: not every problem is an LLM problem. Traditional ML is having a comeback for a reason—it&rsquo;s often cheaper, easier to operationalize, and more effective for specific tasks.</p>
<h2 id="career-advice-lucas-would-give-and-what-hed-tell-you">Career advice Lucas would give (and what he&rsquo;d tell you)</h2>
<p>Lucas&rsquo;s advice for people trying to break into tech right now:</p>
<p>Stick to it.</p>
<p>Don&rsquo;t let the AI noise psych you out. There are plenty of companies nowhere near &ldquo;AI-first&rdquo; that still need great engineers. Take the shot, even if it&rsquo;s not the shiny job. Keep side projects going. Show your work. Stay learning. Don&rsquo;t give up.</p>
<p>And the advice he&rsquo;d give himself 10 years ago was even better:</p>
<p>If you&rsquo;re stuck, challenge your assumptions. Step back. Look elsewhere. There are more solutions than you think—most people just burn time spinning in one lane.</p>
<h2 id="watch-the-episode">Watch the episode</h2>
<p>This one is for anyone navigating leadership, consulting, or just trying to stay sane while AI changes the shape of the industry.</p>
<p>Watch the full Data Pour episode with Lucas Ward here: <a href="https://www.youtube.com/watch?v=732vznyGndQ&amp;t=1011s">https://www.youtube.com/watch?v=732vznyGndQ&amp;t=1011s</a></p>
]]></content:encoded></item><item><title>AWS re:Invent 2025: Democratizing AI (Again)</title><link>https://lukelittle.com/posts/2025/12/aws-reinvent-2025-democratizing-ai-again/</link><pubDate>Mon, 15 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/aws-reinvent-2025-democratizing-ai-again/</guid><description>My take on AWS re:Invent 2025 keynotes—AWS is democratizing AI the same way they democratized infrastructure 15 years ago.</description><content:encoded><![CDATA[<p>AWS re:Invent 2025 keynotes felt like AWS repeating the same move they pulled 15 years ago—democratizing something that used to be gated behind massive budgets and specialized teams. This time it&rsquo;s AI.</p>
<p>I published a full write-up on Ippon&rsquo;s blog breaking down the core thread running through the keynote: scale as the prerequisite for &ldquo;access,&rdquo; abundance via custom silicon (Trainium), and the shift from &ldquo;AI as a feature&rdquo; to &ldquo;AI as a production capability.&rdquo;</p>
<p>The big idea that stuck with me: it&rsquo;s not just models—AWS is packaging hard-won operational and security expertise into agents (DevOps + Security) so more teams can build safely without needing a 30-person specialist squad.</p>
<p>AWS has clearly gone all in on agentic AI, and it&rsquo;s worth watching the keynote to understand where the industry is going.</p>
<p>If you&rsquo;re trying to separate signal from noise after re:Invent, this is the lens I&rsquo;d use—and the question I&rsquo;d ask: now that the barriers keep dropping, what are you actually going to build?</p>
<p>Read the full article here:
<a href="https://blog.ippon.tech/takeaways-from-aws-reinvent-keynotes-2025">https://blog.ippon.tech/takeaways-from-aws-reinvent-keynotes-2025</a></p>
]]></content:encoded></item><item><title>Data Pour with James Barney: from Big Data to AI</title><link>https://lukelittle.com/posts/2025/12/data-pour-with-james-barney-from-big-data-to-ai/</link><pubDate>Mon, 15 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/data-pour-with-james-barney-from-big-data-to-ai/</guid><description>A conversation with James Barney at People&amp;#39;s Market in Myers Park, Charlotte—from big data infrastructure to enterprise AI realities.</description><content:encoded><![CDATA[<p>We filmed this episode of Data Pour at People&rsquo;s Market in Myers Park, Charlotte—grabbed drinks, hit record, and got into the kind of conversation that happens when two people who grew up in &ldquo;big data&rdquo; start comparing notes on where AI is actually headed.</p>
<p>Funny thing is—this bottle shop closed down a week after we filmed. Perfect metaphor for tech, honestly: everything feels stable until it isn&rsquo;t.</p>
<h2 id="why-james">Why James</h2>
<p>James Barney is one of my closest friends and mentors. I&rsquo;ve known him for a long time, and he&rsquo;s been the person I go to when I want an AI take that isn&rsquo;t hype and isn&rsquo;t fear—just reality.</p>
<p>He&rsquo;s spent the last decade-plus inside big enterprises, helping teams evaluate what&rsquo;s real, what scales, and what&rsquo;s just noise. Right now, he&rsquo;s helping lead AI initiatives at a major insurance company—meaning his day job is basically: everyone wants the next big thing… which parts of this are going to stick?</p>
<h2 id="our-shared-roots-fintech-big-data-and-invisible-systems">Our shared roots: fintech, big data, and invisible systems</h2>
<p>We both started our tech careers in big data, and a lot of our growth came from the same kind of pressure cooker—fintech needs.</p>
<p>If you&rsquo;ve worked in regulated financial environments, you know the deal: huge volumes of data, constant scrutiny, and systems that have to be accurate, auditable, and fast. That world forces you to get good at the unsexy stuff—digesting, categorizing, transforming, and moving data reliably.</p>
<p>In the episode, we talk about the era when we were building Kafka infrastructure before half the managed conveniences existed. It was messy, duct-tape engineering—but it&rsquo;s also the foundation for how so many modern &ldquo;invisible&rdquo; systems work today.</p>
<h2 id="the-point-ai-is-new-but-the-patterns-arent">The point: AI is new… but the patterns aren&rsquo;t</h2>
<p>One of the threads we pull on: AI feels new to everyone, but the underlying enterprise challenges rhyme with what we already lived through with cloud and big data.</p>
<p>The model is only part of the story. The real value shows up when you add:</p>
<ul>
<li>clean, governed data</li>
<li>real context</li>
<li>integrations into systems that matter</li>
<li>tools that can retrieve or act (not just &ldquo;generate text&rdquo;)</li>
</ul>
<p>We also get into why copilots can feel underwhelming in enterprises: the public tools people use every day are &ldquo;fully layered.&rdquo; Inside a company, you&rsquo;re often starting with the plain base layer—and you have to build the rest.</p>
<h2 id="watch-the-full-episode">Watch the full episode</h2>
<p>This blog is just the background and the framing—if you want the full story (and the full vibe), watch the YouTube episode. It&rsquo;s a real conversation between two people who&rsquo;ve been in the trenches, trying to describe what&rsquo;s actually happening in AI without turning it into marketing copy or doomposting.</p>
<p>Link to the episode:
<a href="https://youtu.be/ibtZckIQ_JI">https://youtu.be/ibtZckIQ_JI</a></p>
]]></content:encoded></item><item><title>I Tried to Deploy a Simple Website on AWS. It Became a Full-Blown Side Quest.</title><link>https://lukelittle.com/posts/2025/12/i-tried-to-deploy-a-simple-website-on-aws.-it-became-a-full-blown-side-quest./</link><pubDate>Mon, 08 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/i-tried-to-deploy-a-simple-website-on-aws.-it-became-a-full-blown-side-quest./</guid><description>I tried to launch a Hugo site on AWS and ended up doing four hours of DNS archaeology and CloudFront forensics.</description><content:encoded><![CDATA[<p>Recently I came across another engineer&rsquo;s personal blog — clean layout, good typography, that &ldquo;I actually finish my side projects&rdquo; energy — and it pushed me to finally build one of my own.</p>
<p>I picked Hugo because I like Go, and because using Jekyll in 2025 feels like opting into pain. I briefly considered Ghost, remembered it either requires paying Ghost or hosting Ghost, and closed the tab. And since I&rsquo;m &ldquo;the AWS guy,&rdquo; it felt morally necessary to deploy the whole thing on AWS. Maybe I&rsquo;d even use Kiro if I felt extra fancy.</p>
<p>For context: <strong>Hugo</strong> and <strong>Jekyll</strong> are both static site generators—tools that convert Markdown files into HTML at build time. Jekyll, written in Ruby, was the OG choice for GitHub Pages and still powers thousands of blogs. But it&rsquo;s slow, requires managing Ruby dependencies, and feels dated. Hugo, written in Go, is blazingly fast (builds in milliseconds), has a single binary with zero dependencies, and handles large sites without breaking a sweat. Both produce the same outcome—static HTML you can throw on a CDN—but Hugo does it in a fraction of the time and with far less friction.</p>
<p>So I wrote some Terraform, vibe-coded a theme, deployed it, and immediately remembered that I don&rsquo;t build personal websites very often.</p>
<h2 id="the-architecture">The Architecture</h2>
<p>Before diving into the problems I hit, here&rsquo;s what the final architecture looks like.</p>
<p>It&rsquo;s a classic serverless static site setup: <strong>Route 53</strong> handles DNS, pointing the domain to a <strong>CloudFront</strong> distribution. CloudFront sits in front of an <strong>S3 bucket</strong> that stores all the static files—HTML, CSS, JavaScript, images. The bucket is private; CloudFront is the only thing allowed to read from it via an Origin Access Identity.</p>
<p>Here&rsquo;s where it gets slightly more interesting: attached to CloudFront is a <strong>CloudFront Function</strong>—a lightweight JavaScript function that runs at the edge, modifying incoming requests before they hit the origin. This function handles two things: redirecting <code>www.lukelittle.com</code> to the apex domain, and rewriting clean URLs (like <code>/posts/my-article</code>) to their actual paths (<code>/posts/my-article/index.html</code>).</p>
<p>SSL certificates come from <strong>AWS Certificate Manager</strong> and are automatically attached to CloudFront. The whole thing is defined in <strong>Terraform</strong>, and deployment happens via <strong>GitHub Actions</strong>—push to main, Hugo builds the site, the output gets synced to S3, and CloudFront&rsquo;s cache gets invalidated.</p>
<p>The request flow is straightforward: user hits the domain → Route 53 resolves it to CloudFront → CloudFront invokes the edge function → function rewrites the request if needed → CloudFront fetches from S3 (or serves from cache) → content gets delivered globally from the nearest edge location.</p>
<p>Clean. Simple. Serverless. Costs about $0.50/month for the Route 53 hosted zone. Everything else fits in free tier.</p>
<p>Now, the problems.</p>
<h2 id="first-realization-www-doesnt-redirect-itself">First realization: <code>www</code> doesn&rsquo;t redirect itself</h2>
<p>I wanted <code>www.lukelittle.com</code> → <code>lukelittle.com</code>.<br>
Simple enough — except CloudFront doesn&rsquo;t have <code>.htaccess</code> or Apache-style rewrite configs.</p>
<p>The fix is a <strong>CloudFront Function</strong> on <code>viewer-request</code>:</p>
<pre class="mermaid">sequenceDiagram
    participant User as 👤 User Browser
    participant CF as ☁️ CloudFront
    participant Func as ⚡ CloudFront Function
    
    User-&gt;&gt;CF: GET https://www.lukelittle.com/
    CF-&gt;&gt;Func: Viewer Request Event
    Note over Func: Check host header&lt;br/&gt;host === &#39;www.lukelittle.com&#39;
    Func--&gt;&gt;CF: 301 Redirect
    Note over Func: Location: https://lukelittle.com/
    CF--&gt;&gt;User: 301 Moved Permanently
    User-&gt;&gt;CF: GET https://lukelittle.com/
    Note over User: Browser follows redirect
    CF--&gt;&gt;User: 200 OK (homepage)
</pre>

<p>Here&rsquo;s the actual CloudFront Function code:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-javascript" data-lang="javascript"><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">handler</span>(<span style="color:#a6e22e">event</span>) {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">request</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">event</span>.<span style="color:#a6e22e">request</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">host</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">headers</span>.<span style="color:#a6e22e">host</span>.<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e">// Redirect www to apex domain
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>    <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">host</span> <span style="color:#f92672">===</span> <span style="color:#e6db74">&#39;www.lukelittle.com&#39;</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">statusCode</span><span style="color:#f92672">:</span> <span style="color:#ae81ff">301</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">statusDescription</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;Moved Permanently&#39;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">headers</span><span style="color:#f92672">:</span> {
</span></span><span style="display:flex;"><span>                <span style="color:#a6e22e">location</span><span style="color:#f92672">:</span> { <span style="color:#a6e22e">value</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;https://lukelittle.com&#39;</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">uri</span> }
</span></span><span style="display:flex;"><span>            }
</span></span><span style="display:flex;"><span>        };
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#a6e22e">request</span>;
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>CloudFront Functions are perfect for this because they run at the edge, cost almost nothing (you get 2 million free invocations per month), and don&rsquo;t require Lambda, bucket changes, or origin rewrites. They execute in under a millisecond, which means your redirect happens before the user even realizes they typed <code>www</code>.</p>
<p>That part was easy. The next part was not.</p>
<h2 id="second-realization-cloudfront-does-not-assume-indexhtml">Second realization: CloudFront does <em>not</em> assume <code>index.html</code></h2>
<p>Hugo outputs directories like:</p>
<pre tabindex="0"><code>/posts/
/posts/index.html
</code></pre><p>Apache and Nginx automatically serve <code>index.html</code> when you access <code>/posts/</code>.<br>
CloudFront does not. It will happily 404 unless you rewrite the URI yourself.</p>
<p>Here&rsquo;s what&rsquo;s happening under the hood:</p>
<pre class="mermaid">sequenceDiagram
    participant User as 👤 User Browser
    participant CF as ☁️ CloudFront
    participant Func as ⚡ CloudFront Function
    participant S3 as 🪣 S3 Bucket
    
    User-&gt;&gt;CF: GET https://lukelittle.com/posts/cracking-the-cloud
    CF-&gt;&gt;Func: Viewer Request Event
    Note over Func: URI: /posts/cracking-the-cloud&lt;br/&gt;No extension detected&lt;br/&gt;!uri.includes(&#39;.&#39;)
    Func-&gt;&gt;Func: Append /index.html
    Note over Func: New URI:&lt;br/&gt;/posts/cracking-the-cloud/index.html
    Func--&gt;&gt;CF: Modified Request
    CF-&gt;&gt;S3: GetObject&lt;br/&gt;/posts/cracking-the-cloud/index.html
    S3--&gt;&gt;CF: HTML Content
    CF--&gt;&gt;User: 200 OK (post content)
    Note over CF: Cache for 1 hour
</pre>

<p>So I added this logic to the same CloudFront Function:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-javascript" data-lang="javascript"><span style="display:flex;"><span><span style="color:#66d9ef">function</span> <span style="color:#a6e22e">handler</span>(<span style="color:#a6e22e">event</span>) {
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">request</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">event</span>.<span style="color:#a6e22e">request</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">uri</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">uri</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">host</span> <span style="color:#f92672">=</span> <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">headers</span>.<span style="color:#a6e22e">host</span>.<span style="color:#a6e22e">value</span>;
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e">// Redirect www to apex domain
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>    <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">host</span> <span style="color:#f92672">===</span> <span style="color:#e6db74">&#39;www.lukelittle.com&#39;</span>) {
</span></span><span style="display:flex;"><span>        <span style="color:#66d9ef">return</span> {
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">statusCode</span><span style="color:#f92672">:</span> <span style="color:#ae81ff">301</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">statusDescription</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;Moved Permanently&#39;</span>,
</span></span><span style="display:flex;"><span>            <span style="color:#a6e22e">headers</span><span style="color:#f92672">:</span> {
</span></span><span style="display:flex;"><span>                <span style="color:#a6e22e">location</span><span style="color:#f92672">:</span> { <span style="color:#a6e22e">value</span><span style="color:#f92672">:</span> <span style="color:#e6db74">&#39;https://lukelittle.com&#39;</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">uri</span> }
</span></span><span style="display:flex;"><span>            }
</span></span><span style="display:flex;"><span>        };
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#75715e">// Append index.html for clean URLs
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span>    <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">uri</span>.<span style="color:#a6e22e">endsWith</span>(<span style="color:#e6db74">&#39;/&#39;</span>)) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">uri</span> <span style="color:#f92672">+=</span> <span style="color:#e6db74">&#39;index.html&#39;</span>;
</span></span><span style="display:flex;"><span>    } <span style="color:#66d9ef">else</span> <span style="color:#66d9ef">if</span> (<span style="color:#f92672">!</span><span style="color:#a6e22e">uri</span>.<span style="color:#a6e22e">includes</span>(<span style="color:#e6db74">&#39;.&#39;</span>)) {
</span></span><span style="display:flex;"><span>        <span style="color:#a6e22e">request</span>.<span style="color:#a6e22e">uri</span> <span style="color:#f92672">+=</span> <span style="color:#e6db74">&#39;/index.html&#39;</span>;
</span></span><span style="display:flex;"><span>    }
</span></span><span style="display:flex;"><span>    
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">return</span> <span style="color:#a6e22e">request</span>;
</span></span><span style="display:flex;"><span>}
</span></span></code></pre></div><p>Now CloudFront behaves like a normal web server circa 2008. Hugo pages immediately started working.</p>
<p>For a moment.</p>
<h2 id="where-everything-went-off-the-rails">Where everything went off the rails</h2>
<p>I tried to add <code>www.lukelittle.com</code> as an alternate domain name on the CloudFront distribution.</p>
<p>CloudFront refused.<br>
The error claimed it was <strong>already associated with another distribution</strong>.</p>
<p>It wasn&rsquo;t — at least not in any AWS account I currently have access to.</p>
<p>So I did what any rational engineer does:</p>
<ul>
<li>Googled</li>
<li>Re-Googled</li>
<li>Asked ChatGPT</li>
<li>Deleted the distribution</li>
<li>Recreated the distribution</li>
<li>Repeated the cycle</li>
<li>Began questioning my past life choices</li>
</ul>
<p>After four hours, I accepted defeat and temporarily upgraded my support plan.</p>
<p>The answer was unexpected:</p>
<p><code>www.lukelittle.com</code> <em>was</em> attached to a CloudFront distribution — in <strong>another AWS account</strong>.</p>
<p>Which account?<br>
I have no clue.<br>
Possibilities include:</p>
<ul>
<li>some forgotten sandbox from 2017</li>
<li>a leftover test account</li>
<li>an old attempt at this blog I completely wiped from memory</li>
<li>a parallel universe</li>
</ul>
<p>Support asked me to add a TXT record to prove I owned the domain. I added it in Route 53, they cleared the stale binding, and immediately everything began working exactly as expected.</p>
<hr>
<h2 id="how-deployment-actually-works">How deployment actually works</h2>
<p>Once I got the infrastructure sorted, I needed a deployment pipeline. GitHub Actions + OIDC federation makes this dead simple:</p>
<pre class="mermaid">sequenceDiagram
    participant Dev as 👨‍💻 Developer
    participant GH as GitHub
    participant GHA as GitHub Actions
    participant IAM as 🔑 IAM Role
    participant S3 as 🪣 S3 Bucket
    participant CFront as ☁️ CloudFront
    
    Dev-&gt;&gt;GH: git push origin main
    GH-&gt;&gt;GHA: Trigger workflow
    Note over GHA: hugo --minify
    GHA-&gt;&gt;GHA: Build static site
    
    GHA-&gt;&gt;IAM: AssumeRoleWithWebIdentity
    Note over IAM: OIDC Federation&lt;br/&gt;Verify GitHub token
    IAM--&gt;&gt;GHA: Temporary credentials
    
    GHA-&gt;&gt;S3: aws s3 sync ./public s3://bucket/
    Note over S3: Upload HTML, CSS, JS&lt;br/&gt;--delete flag removes old files
    S3--&gt;&gt;GHA: Sync complete
    
    GHA-&gt;&gt;CFront: CreateInvalidation --paths &#34;/*&#34;
    Note over CFront: Clear edge cache&lt;br/&gt;Force fresh content
    CFront--&gt;&gt;GHA: Invalidation ID
    
    Note over Dev,CFront: Deployment complete ✅&lt;br/&gt;New content live globally
</pre>

<p>Here&rsquo;s what makes this beautiful:</p>
<p><strong>No long-lived credentials.</strong> GitHub Actions uses OIDC to assume an IAM role, gets temporary credentials that expire in an hour, and those credentials only work for this specific repo. If someone compromises the GitHub Actions environment, they get access for 60 minutes max—and only to deploy this blog. Not exactly a treasure trove.</p>
<p><strong>Hugo builds in ~200ms.</strong> Static site generators are fast when your entire site fits in memory. No database queries, no server-side rendering, just Markdown → HTML and done.</p>
<p><strong>S3 sync is smart.</strong> It only uploads files that changed. New post? Upload one file. Tweak CSS? Upload one file. CloudFront invalidation clears the edge cache, so every visitor gets fresh content within seconds.</p>
<p><strong>Global deployment in under a minute.</strong> Push to main → build → sync → invalidate → live. The entire pipeline runs faster than most people can brew coffee.</p>
<h2 id="and-now-the-site-actually-exists">And now the site actually exists</h2>
<p>Despite writing constantly — deep dives, rants, slides, Data Pour episodes — I&rsquo;ve never had a single place to put any of it. Everything has been scattered across GitHub repos, Slack threads, LinkedIn posts, and random folders.</p>
<p>This site fixes that.</p>
<p>Hugo prerenders everything.<br>
CloudFront serves it globally.<br>
There&rsquo;s no backend.<br>
No patching.<br>
No maintenance.<br>
Just HTML, a CDN, and vibes.</p>
<p>The whole stack:</p>
<ul>
<li><strong>Hugo</strong> for static site generation</li>
<li><strong>S3</strong> for object storage</li>
<li><strong>CloudFront</strong> for CDN + edge functions</li>
<li><strong>Route 53</strong> for DNS</li>
<li><strong>ACM</strong> for SSL certificates</li>
<li><strong>GitHub Actions</strong> for CI/CD</li>
<li><strong>Terraform</strong> for infrastructure as code</li>
</ul>
<p>Total monthly cost: ~$0.50 for Route 53 hosted zone. Everything else fits in free tier.</p>
<h2 id="the-unexpected-benefit">The unexpected benefit</h2>
<p>Honestly, getting stuck for a few hours was probably good for me. I had to slow down, re-read documentation, and remember exactly how CloudFront, ACM, and Route 53 interact — instead of relying on half-remembered muscle memory.</p>
<p>Building this site reminded me why I got into infrastructure in the first place: you can take a bunch of managed services, wire them together thoughtfully, and end up with something that just works. No servers to patch, no databases to tune, no midnight pages about memory leaks.</p>
<p>Just a website that loads fast, costs nothing, and requires zero maintenance.</p>
<p>Not the night I planned, but not wasted either.</p>
<p>Anyway — the blog is live now. Hopefully the next update doesn&rsquo;t require another round of DNS archaeology or CloudFront forensics.</p>
<p><strong>Want to see the code?</strong> The entire infrastructure setup is open source:<br>
<a href="https://github.com/lukelittle/homepage">github.com/lukelittle/homepage</a></p>
]]></content:encoded></item><item><title>AI, Uncertainty, and the Rise of the Renaissance Developer</title><link>https://lukelittle.com/posts/2025/12/ai-uncertainty-and-the-rise-of-the-renaissance-developer/</link><pubDate>Sun, 07 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/ai-uncertainty-and-the-rise-of-the-renaissance-developer/</guid><description>Thoughts on AI, the future of development, and what it means to be a builder in an era of rapid technological change.</description><content:encoded><![CDATA[<p>When I talk to students — especially STEM students — one question keeps coming up: &ldquo;Is AI going to take my job?&rdquo; They ask it jokingly, but you can tell they&rsquo;re serious. What they want is assurance. They want to know there&rsquo;s a light at the end of the tunnel. That the late nights, the debt, the effort, and the hope they&rsquo;ve poured into their degree will amount to something real.</p>
<p>And I rarely give a quick answer, because the truth is uncomfortable: maybe. It&rsquo;s not the certainty anyone wants, but it&rsquo;s honest.</p>
<p>For a long time, I danced around that answer. I hedged. I emphasized the optimistic parts, hoping they would land softer. But last week, during his final re:Invent keynote, Werner Vogels showed us how to talk about this moment with clarity instead of avoidance.</p>
<p>He walked the audience through the history of software development — from punch cards and COBOL to structured programming, object orientation, distributed systems, the cloud, and now agentic AI. In every era, developers were convinced the sky was falling. And in some ways, it was. Punch-card clerks disappeared. Mainframe operators disappeared. Entire roles vanished because something faster, cheaper, and more scalable came along.</p>
<p>But one role never disappeared: the builder — the person who understands systems, solves real-world problems, thinks in abstractions, and brings ideas to life.</p>
<p>That&rsquo;s the part people forget. Technology doesn&rsquo;t just destroy. Technology democratizes. It opens doors. And it does so in waves.</p>
<p>Every wave has displaced someone. Every wave has taken something away. But every wave has also created something bigger for the people willing to evolve. Builders adapted. They reskilled. They moved up the stack. They went from pushing paper to writing code, from writing code to designing systems. The pattern is always the same: a wave arrives, jobs shift, and the builders who lean in rise with it.</p>
<p>I&rsquo;ve lived through enough of these waves to see the pattern clearly.</p>
<p>The first wave I experienced was the World Wide Web. Before the web, we logged into CompuServe, posted on message boards, browsed clunky online catalogs. And then came the browser — Netscape, Mosaic, whatever arrived on your screen first. Suddenly anyone with a little HTML could put a page online for the entire world to see. The web was so small in those days that there were literal lists of websites. That&rsquo;s how new it was.</p>
<p>And that was the first time technology felt like something you could create with, not just consume. The web democratized publishing and expression. It changed how we interacted with the world — Amazon instead of bookstores, eBay instead of classifieds — and it changed what it meant to be a builder. You didn&rsquo;t need a printing press or a distribution network. You needed curiosity and a willingness to learn a markup language.</p>
<p>Then came the cloud. Suddenly anyone with a credit card and a vision could become a global service provider. You didn&rsquo;t need a server room. You didn&rsquo;t need procurement approvals. You didn&rsquo;t need capital. The cloud collapsed the distance between idea and impact. Entire companies were born in dorm rooms because the playing field had leveled.</p>
<p>And now we&rsquo;re in the third wave: agentic AI. Tools that can generate code, test it, reason about it, orchestrate systems, and solve problems alongside you — without you manually writing every line. You don&rsquo;t need specialized hardware or a custom model. You can build on top of general-purpose models, wrap your logic behind a Model Context Protocol, and stand up entire services in hours. Once again, the gap between imagination and execution has narrowed.</p>
<p>This is what Vogels calls the Renaissance Developer — someone who cultivates depth and breadth. Someone who goes deep enough to solve hard problems but understands enough of adjacent systems to see the whole picture. Someone who stays curious, thinks in systems, communicates clearly, takes ownership, and adapts when the next wave comes.</p>
<p>Because there will always be a next wave.</p>
<p>So when students ask whether AI will take their job, I finally know how to answer. The question is slightly off. The better question is: What can I build with AI?</p>
<p>If you define yourself by a tool — a language, a framework, a stack — then yes, you&rsquo;re in trouble. Tools come and go. They always have. The developers who clung to punch cards got left behind. The ones who refused to learn the web got left behind. The ones who ignored the cloud got left behind. And the ones who dismiss AI will get left behind too.</p>
<p>But if you define yourself by what you create — if you define yourself as a builder — then AI isn&rsquo;t a threat. AI is another wave of democratization. AI is a force multiplier. AI is the next thing that expands what&rsquo;s possible for people willing to evolve.</p>
<p>The path forward isn&rsquo;t about predicting the next dominant framework or clinging to one narrow specialty. The path forward is about building. That&rsquo;s the one thing AI cannot replace: the human who knows what should exist and has the drive to bring it to life.</p>
<p>So will AI take your job?
Maybe.
Probably, if the job is defined purely by a tool.</p>
<p>But will AI take away your ability to create, adapt, imagine, build?
No. It won&rsquo;t. It can&rsquo;t.</p>
<p>If you lean into what it means to be a Renaissance Developer, then AI doesn&rsquo;t close the door on your future.
AI blows it wide open.</p>
]]></content:encoded></item><item><title>Cracking the Cloud: Preparing Students to Build in a Tougher Job Market</title><link>https://lukelittle.com/posts/2025/12/cracking-the-cloud-preparing-students-to-build-in-a-tougher-job-market/</link><pubDate>Sat, 06 Dec 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/12/cracking-the-cloud-preparing-students-to-build-in-a-tougher-job-market/</guid><description>How students can leverage cloud platforms to gain real experience and stand out in an increasingly competitive job market.</description><content:encoded><![CDATA[<p>I was invited to speak at UNC Charlotte recently to the computer science programs, and I approached the session, Cracking the Cloud, with a very specific goal: to give students a realistic, actionable way to stand out in a job market that&rsquo;s becoming more competitive every year. Companies are slowing early-career hiring. AI is reshaping workflows. Expectations for junior talent are rising, not shrinking. Students can sense this shift, but many don&rsquo;t know what to do with that reality. That&rsquo;s where the conversation begins.</p>
<p>The message I shared with them is simple: even as the job market tightens, students today have an unprecedented advantage that previous generations didn&rsquo;t — they can build real, meaningful things from their dorm rooms. Modern cloud platforms have removed the resource barriers that once kept students from gaining hands-on experience. They don&rsquo;t need servers, budget approvals, or someone in authority to greenlight their ideas. With nothing more than a laptop and the AWS Free Tier, they can deploy applications, experiment with architectures, analyze data, and work with AI. The tools that used to be locked behind enterprise doors are now available to anyone willing to try.</p>
<p>This democratization of technology matters because it gives students something incredibly valuable: the ability to build experience before they have experience. When the market tightens, that distinction becomes decisive. Employers want graduates who can contribute, who understand how modern systems work, and who have touched real tools — not just read about them in a textbook. Students who start building early walk into interviews with confidence and momentum. Students who wait often find themselves starting from behind.</p>
<p>That&rsquo;s why I encourage students to pursue an early cloud certification. It&rsquo;s not about collecting badges. It&rsquo;s about giving them a foundation that unlocks everything else. A certification gets them past HR filters, but more importantly, it teaches them the vocabulary and mental models they need to become actual builders. Once they have that baseline, the cloud stops feeling abstract. They can dive in, deploy something real, break things, fix them, and learn through doing. That&rsquo;s where real growth happens.</p>
<p>To help them get started, I shared a simple 30-60-90 plan:</p>
<p>In the first 30 days, learn the fundamentals. The AWS Cloud Practitioner exam is a gentle entry point that teaches students how the cloud works and how the pieces fit together.</p>
<p>By 60 days, build something small. A website, an API, a basic data pipeline — anything that forces them to make architectural decisions and use real services.</p>
<p>By 90 days, connect with people. Attend meetups, talk with alumni, share what they built, and start forming the relationships that will carry them into internships and full-time roles.</p>
<p>This plan works because it&rsquo;s simple, it&rsquo;s doable, and it converts uncertainty into direction. Students don&rsquo;t need to wait for permission. They just need a starting point.</p>
<h2 id="what-comes-next">What Comes Next</h2>
<p>The response to the talk made one thing clear: students are hungry for guidance, community, and hands-on support. So naturally, I bought CrackingTheCloud.com, and I intend to do something meaningful with it. The vision is bigger than a single event or a single presentation. I want to build a space — both digital and physical — where early-career engineers can learn from industry, explore real tools, and develop confidence long before graduation.</p>
<p>Over the coming months, I plan to start hosting meetups in Charlotte and Richmond, bringing together students, engineers, hiring managers, and community leaders. The goal isn&rsquo;t to create another generic networking group. It&rsquo;s to create an environment where students can ask real questions, see real demos, get unfiltered career advice, and make connections that matter. When students can sit across from practitioners, hear how the industry works, and get direct feedback on their projects, the gap between classroom theory and real-world engineering gets much smaller.</p>
<p>And this matters — not just for students, but for the entire ecosystem. When we invest in early talent, we improve the pipeline for every company in the region. We reduce onboarding burden. We strengthen local tech communities. And we help students step into the industry with confidence, clarity, and purpose. The cloud may have democratized the tools, but it&rsquo;s up to us to democratize the guidance.</p>
<p>Cracking the Cloud started as a talk, but it won&rsquo;t end there. There&rsquo;s a real opportunity to build something lasting — a community that helps students navigate a challenging market, access modern tools, and become the kind of builders the industry needs. And if the energy from UNC Charlotte is any indication, this is only the beginning.</p>
]]></content:encoded></item><item><title>Taking the Mic on The Data Pour</title><link>https://lukelittle.com/posts/2025/11/taking-the-mic-on-the-data-pour/</link><pubDate>Fri, 07 Nov 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/11/taking-the-mic-on-the-data-pour/</guid><description>My first appearance on The Data Pour, filmed at Wooden Robot Brewery, where I talked about my career journey, cloud platforms, and the future of AI.</description><content:encoded><![CDATA[<p>I recently stepped into the role of host for Ippon&rsquo;s Data Pour series, but before I officially took over, our CTO, Andy Lamora, sat me down at Wooden Robot Brewery in Charlotte and interviewed me on camera. It was a great atmosphere — good beer, good weather — and then a handful of cameras appeared, all pointed directly at me. I&rsquo;m still getting used to public speaking, so the entire setup felt a little awkward. Beer helps, but only so much.</p>
<p>The conversation itself was wide-ranging in the best way. Andy and I talked through my background and how I somehow ended up moving from baking to banking to cloud consulting. It&rsquo;s not a linear story, but that&rsquo;s the point — careers rarely follow a straight line, and most of the interesting stuff happens in the detours. We dug into some of the pivots that shaped my path, the things that pulled me deeper into cloud and DevOps work, and what I&rsquo;ve learned hopping between industries and roles along the way.</p>
<p>We eventually moved into the topics I spend most of my time thinking about: cloud platforms, data, and where AI fits into all of it. I shared some opinions on agentic systems, the future of operational automation, and why building a strong platform data foundation matters long before you start layering in intelligent tooling. It wasn&rsquo;t a rehearsed script — just two people talking through where the industry is heading, what excites us, and what still makes us pause.</p>
<p>Filming at Wooden Robot added an unexpected layer to the experience. Between the smell of beer brewing, the coffee bar behind us, and the general energy of the place, the whole conversation felt more natural than anything you&rsquo;d get in a studio. That&rsquo;s the vibe I want to bring into Season 4: real conversations, in real environments, with people who actually build and think about this stuff every day.</p>
<p>If you want to watch me awkwardly navigate my first time on camera — before I technically became the host — the episode is here:
<a href="https://www.youtube.com/watch?v=ZyoRAQsS9c0">https://www.youtube.com/watch?v=ZyoRAQsS9c0</a></p>
<p><em>Season 4 has since wrapped. Every episode is on the <a href="/talks/">Talks</a> page.</em></p>
]]></content:encoded></item><item><title>Learning from the October 20 AWS Outage: Questions Every Team Should Ask</title><link>https://lukelittle.com/posts/2025/10/learning-from-the-october-20-aws-outage-questions-every-team-should-ask/</link><pubDate>Mon, 20 Oct 2025 09:00:00 -0500</pubDate><guid>https://lukelittle.com/posts/2025/10/learning-from-the-october-20-aws-outage-questions-every-team-should-ask/</guid><description>What happened in the October 20 AWS outage, how to explain the impact to leadership, and the resilience questions every team should ask.</description><content:encoded><![CDATA[<p>On the morning of October 20, AWS us-east-1 services were degraded—in particular, DNS services for DynamoDB. Most of us didn&rsquo;t find out from monitoring alerts or dashboards. We found out because the apps on our phones stopped working.</p>
<p>That&rsquo;s the reality of modern infrastructure incidents: they often surface as user-facing failures long before the official root cause analysis lands in your inbox.</p>
<p>I wrote about this outage for <a href="https://blog.ippon.tech/explaining-the-october-20-aws-outage-to-leadership/">Ippon Technologies</a>, focusing on three critical aspects that go beyond just understanding what broke:</p>
<h2 id="what-actually-happened">What Actually Happened</h2>
<p>This wasn&rsquo;t a full region going dark—it was a DNS problem at a foundational layer. DNS (Domain Name System) is the internet&rsquo;s address book. When DNS breaks, everything that depends on it breaks too.</p>
<p>It&rsquo;s like someone removing all the street signs in a city overnight—your services are still there, but nothing can find its way.</p>
<p>For many teams, it meant increased error rates, intermittent failures, and retry storms. Not catastrophic downtime, but the kind of disruption that floods support tickets and frustrates users.</p>
<h2 id="having-the-leadership-conversation">Having the Leadership Conversation</h2>
<p>This is where many engineers struggle. You know the issue was upstream. You know it&rsquo;s AWS&rsquo;s infrastructure. But leadership doesn&rsquo;t care about the cloud provider—they care about impact and what you&rsquo;re doing about it.</p>
<p>The key is framing it as a <strong>dependency visibility problem</strong>, not a blame game:</p>
<blockquote>
<p>&ldquo;We were affected by a regional outage in AWS&rsquo;s us-east-1 region due to DNS issues. This exposed areas where we&rsquo;re overly dependent on single-region infrastructure. We&rsquo;re using this to map our regional dependencies, prioritize applications by criticality, and identify where we need fallback logic and multi-region routing.&rdquo;</p>
</blockquote>
<p>That&rsquo;s ownership. That&rsquo;s a path forward.</p>
<h2 id="the-hard-questions-you-need-to-answer">The Hard Questions You Need to Answer</h2>
<p>The article explores four critical questions every team should be able to answer confidently:</p>
<ol>
<li><strong>Which of your apps run in us-east-1?</strong></li>
<li><strong>Which rely on DynamoDB?</strong></li>
<li><strong>Which of those are Tier 1 or customer-facing?</strong></li>
<li><strong>Which of those have active-active failover across regions?</strong></li>
</ol>
<p>Most teams can&rsquo;t answer these questions. Not because they&rsquo;re negligent—but because cloud estates grow organically. Services get deployed. Teams change. Documentation drifts. Before you know it, you&rsquo;re running critical workloads on infrastructure patterns that no one fully understands anymore.</p>
<h2 id="making-resilience-visible">Making Resilience Visible</h2>
<p>The full article dives into:</p>
<ul>
<li><strong>Structured risk assessment</strong> using tools like AWS Resilience Hub to define applications and assess risk against RTO/RPO targets</li>
<li><strong>Chaos engineering</strong> with AWS Fault Injection Service (FIS) to validate that resilience isn&rsquo;t just theoretical</li>
<li><strong>Cultural shifts</strong> to prioritize resilience alongside feature delivery</li>
<li><strong>Practical next steps</strong> for mapping dependencies and building observability around failure modes</li>
</ul>
<h2 id="why-this-matters">Why This Matters</h2>
<p>Today&rsquo;s outage wasn&rsquo;t catastrophic—but it was loud enough to get everyone&rsquo;s attention. It revealed real architectural risks that often go unnoticed until they become outages.</p>
<p>Outages like this are reminders, not just disruptions. They&rsquo;re opportunities to begin conversations across architecture, risk, and engineering teams about what resilience really means for your organization.</p>
<p>Not in terms of making everything indestructible, but in making risk visible, decisions intentional, and recovery predictable.</p>
<p>👉 <strong><a href="https://blog.ippon.tech/explaining-the-october-20-aws-outage-to-leadership/">Read the full article on the Ippon blog</a></strong> for detailed guidance on communicating with leadership, assessing your infrastructure, and building measurable resilience practices.</p>
<p>The goal isn&rsquo;t perfection—it&rsquo;s visibility, intention, and readiness for when the next incident inevitably arrives.</p>
]]></content:encoded></item></channel></rss>