[{"content":"Everyone in financial services talks about resilience. We have DR plans, architecture diagrams, dashboards, and increasingly, tools like AWS Resilience Hub. On paper, it all looks good. In practice, most resilience programs don\u0026rsquo;t fail because of missing tooling — they fail because they never move beyond isolated assessments.\nA team runs a Resilience Hub assessment. They get a score. Maybe they even fix a few findings. And then nothing happens. No aggregation. No cadence. No program. Just a snapshot that lives in one account, attached to one application, reviewed once.\nI\u0026rsquo;ve seen this pattern enough times in financial services to know it\u0026rsquo;s not a tooling problem. It\u0026rsquo;s an organizational one. And if you want resilience that actually means something, you have to confront it directly.\nResilience Doesn\u0026rsquo;t Live in One Account Modern financial systems don\u0026rsquo;t live in one place. They span line-of-business accounts, shared platform layers, data environments, and third-party dependencies. The critical user journeys — payments, trades, claims — cut across all of them.\nYet Resilience Hub is typically deployed against one app, in one account, with one assessment. That\u0026rsquo;s not resilience. That\u0026rsquo;s local optimization pretending to be a program.\nAs outlined in the Building Resilient Financial Services whitepaper, resilience must be measurable and demonstrable across interconnected services, not evaluated in isolation. A payment flow that touches API Gateway in one account, a processing Lambda in a shared services account, and DynamoDB in a data account doesn\u0026rsquo;t care that each of those components passed its individual resilience check. What matters is whether the flow as a whole can survive failure.\nThis is the gap Resilience Hub doesn\u0026rsquo;t close on its own. It was designed to evaluate applications, not programs. If you want the latter, you have to build it.\nThe Pattern That Actually Works: Hub-and-Spoke If you want to scale resilience, you need to treat it like a distributed measurement system. Each account owns its applications and runs its own assessments. But the program — the aggregation, the scoring, the trend analysis — lives centrally.\nThis isn\u0026rsquo;t complicated architecture. It\u0026rsquo;s a standard IAM role in each spoke account that a central account can assume, a scheduled Lambda or Step Functions workflow that pulls assessment data across accounts and normalizes the structure, a data store in S3 or DynamoDB that tracks history (not just current state), and a dashboard layer — QuickSight or whatever your organization already uses — tied to critical services and business impact.\nThe hard part isn\u0026rsquo;t building it. The hard part is getting the organizational agreement that resilience data should flow centrally in the first place.\nResilience Hub Is a Data Source, Not a Dashboard This is where I think most teams go wrong. They use Resilience Hub as a dashboard — open the console, look at the score, maybe screenshot it for a quarterly review. That\u0026rsquo;s underutilizing it significantly.\nResilience Hub is a modeling engine, a policy evaluator, and a scoring system. It is not your enterprise view, your governance layer, or your program. The real value emerges when you stop looking at the UI and start treating it as a data source.\nThe API surface is small and that\u0026rsquo;s a good thing:\nclient.list_apps() client.list_app_assessments() client.describe_app_assessment() From those three calls, you can build enterprise-wide resilience scoring, trend analysis over time, drift detection across environments, and audit evidence tied to real systems. The UI tells you what happened once. The API lets you prove improvement over time — and that\u0026rsquo;s what regulators actually care about.\nWhere This Connects to Program Cadence This is the part most organizations completely miss. Resilience is not a tool, a dashboard, or a quarterly checkbox. It\u0026rsquo;s a cadence.\nI keep coming back to this when working with financial services clients: the difference between organizations that have a resilience program and organizations that have a resilience report is whether they review on a regular cadence and actually act on what they find.\nCadence What actually happens Weekly Pull Resilience Hub data, detect drift Monthly Review resilience posture across services Quarterly Validate against real scenarios Continuous Track improvement and regressions Monthly reviews. Quarterly scenario testing. Continuous improvement loops. If you\u0026rsquo;re not doing these, you don\u0026rsquo;t have a program — you have a document that says you do.\nFrom Assessment to Evidence Regulators don\u0026rsquo;t care about your architecture diagrams. They care about whether services stay within tolerance, whether failures are tested, and whether you improve over time. These aren\u0026rsquo;t unreasonable asks, but they require evidence that most organizations can\u0026rsquo;t produce because they never built the infrastructure to collect it.\nA multi-account Resilience Hub strategy gives you centralized visibility across distributed systems, consistent measurement aligned to business services, and — critically — evidence you can actually defend when someone asks how you know your systems are resilient.\nMost organizations never make this leap. They run assessments, fix findings, and move on. But resilience doesn\u0026rsquo;t scale that way. It scales when you treat it as a system — measured continuously, aggregated centrally, and reviewed on a cadence.\nAWS Resilience Hub gives you the raw signal. What you build around it determines whether you have a tool or a capability.\n","permalink":"https://lukelittle.com/posts/2026/04/scaling-resilience-with-aws-resilience-hub-a-multi-account-reality-check/","summary":"\u003cp\u003eEveryone in financial services talks about resilience. We have DR plans, architecture diagrams, dashboards, and increasingly, tools like AWS Resilience Hub. On paper, it all looks good. In practice, most resilience programs don\u0026rsquo;t fail because of missing tooling — they fail because they never move beyond isolated assessments.\u003c/p\u003e\n\u003cp\u003eA team runs a Resilience Hub assessment. They get a score. Maybe they even fix a few findings. And then nothing happens. No aggregation. No cadence. No program. Just a snapshot that lives in one account, attached to one application, reviewed once.\u003c/p\u003e","title":"Scaling Resilience with AWS Resilience Hub: A Multi-Account Reality Check"},{"content":"Richmond AWS User Group: An Evening of AI-Powered Development The Richmond AWS User Group hosted another successful meetup on March 5th, featuring a fascinating presentation on Kiro, AWS\u0026rsquo;s cutting-edge agentic coding tool. The event brought together local cloud enthusiasts, developers, and students for an evening of learning, networking, and pizza.\nKiro: Revolutionizing Software Development Dinesh Balaaji Prabakaran from Amazon Web Services delivered an impressive talk on Kiro, showcasing how this innovative tool is transforming the software development landscape. Kiro represents a significant advancement in AI-assisted development, helping developers bridge the gap between ideas and implementation through intelligent code generation, problem reasoning, and accelerated development workflows.\nWhat makes Kiro particularly noteworthy is its ability to understand complex requirements and generate functional code that aligns with best practices. Throughout the presentation, you could see attendees\u0026rsquo; expressions shift from curiosity to amazement as Dinesh demonstrated Kiro\u0026rsquo;s capabilities.\nA Strong Follow-Up to Our Previous Kiro Session This meetup built upon our previous session where I had the opportunity to build something live with Kiro, giving attendees a first-hand look at its practical applications. Having these consecutive sessions focused on Kiro allowed our community to develop a deeper understanding of the tool and its potential impact on development practices.\nEngaged Community Discussions The Q\u0026amp;A session reflected our diverse community, with questions ranging from practical implementation details to more complex discussions about guardrails and ethical considerations. The varied perspectives from industry practitioners, academics, and students created a rich dialogue about both the technical capabilities and the broader implications of agentic coding tools.\nCommunity Perks AWS generously provided swag and Kiro credits for attendees, giving everyone the opportunity to experiment with the tool themselves. These practical resources ensure that the knowledge shared during the meetup can be immediately applied, extending the learning experience beyond the event itself.\nGrowing the Richmond Cloud Community The Richmond AWS User Group continues to serve as a valuable hub for cloud professionals, students, and enthusiasts in the area. These regular gatherings provide opportunities for networking, learning, and staying current with the rapidly evolving AWS ecosystem.\nWhether you\u0026rsquo;re an experienced builder, a student exploring career options, a solutions architect, or an engineer looking to expand your cloud expertise, the Richmond AWS User Group welcomes you to join our community of learners.\nLooking Ahead As AI-assisted development tools like Kiro become increasingly integrated into modern development workflows, our community remains committed to exploring these technologies together. We\u0026rsquo;re already planning future sessions that will dive deeper into specific use cases and advanced features.\nStay connected with the Richmond AWS User Group for announcements about upcoming events, and join us as we continue to explore the cutting edge of cloud technology together.\n","permalink":"https://lukelittle.com/posts/2026/03/richmond-aws-user-group-presentation-on-kiro/","summary":"\u003ch2 id=\"richmond-aws-user-group-an-evening-of-ai-powered-development\"\u003eRichmond AWS User Group: An Evening of AI-Powered Development\u003c/h2\u003e\n\u003cp\u003eThe Richmond AWS User Group hosted another successful meetup on March 5th, featuring a fascinating presentation on Kiro, AWS\u0026rsquo;s cutting-edge agentic coding tool. The event brought together local cloud enthusiasts, developers, and students for an evening of learning, networking, and pizza.\u003c/p\u003e\n\u003ch2 id=\"kiro-revolutionizing-software-development\"\u003eKiro: Revolutionizing Software Development\u003c/h2\u003e\n\u003cp\u003eDinesh Balaaji Prabakaran from Amazon Web Services delivered an impressive talk on Kiro, showcasing how this innovative tool is transforming the software development landscape. Kiro represents a significant advancement in AI-assisted development, helping developers bridge the gap between ideas and implementation through intelligent code generation, problem reasoning, and accelerated development workflows.\u003c/p\u003e","title":"Richmond AWS User Group Presentation on Kiro"},{"content":"I had the privilege of speaking at the Department of Computer Science at Virginia Commonwealth University\u0026rsquo;s College of Engineering on March 3rd, presenting \u0026ldquo;Cracking the Cloud: How AWS Certifications Can Launch Your Career\u0026rdquo; to a group of engaged Computer Science seniors.\nThe Presentation My talk focused on how early-career technologists can differentiate themselves in an increasingly competitive field. Three themes anchored the session: strategic certifications — particularly AWS certifications as a way to validate cloud skills before you have the job title to back them up — self-directed projects that demonstrate practical problem-solving, and building genuine relationships in the industry. That last one always gets a laugh because I tell students it sometimes means getting away from the keyboard entirely and, as I like to say, \u0026ldquo;doing a little grass touching.\u0026rdquo;\nI shared my own path — leveraging AWS certifications early to get into high-visibility projects as industries were actively migrating to cloud. While I had to update a few dates from previous versions of the deck, the core message has only gotten more urgent.\nAWS: The Connective Tissue of AI and Data One thread I came back to throughout the talk is that there has never been a better time to learn AWS than right now. AI and data technologies are advancing at a pace that makes prior tech cycles look like dress rehearsals. AWS isn\u0026rsquo;t just infrastructure anymore — it\u0026rsquo;s the connective tissue that binds these technologies together and makes them deployable at enterprise scale.\nFrom Amazon Bedrock for foundation model integration to SageMaker for MLOps, AWS is where the theoretical promise of AI meets production reality. For students entering the workforce today, cloud architecture fluency is no longer a specialty — it\u0026rsquo;s table stakes. The engineers who understand how to wire these services together, govern them properly, and make them perform under real-world constraints are the ones who will lead.\nVCU\u0026rsquo;s CS Program: Engineering Discipline Meets Modern Stack What made this talk particularly meaningful was the audience. VCU\u0026rsquo;s BS in Computer Science, housed within the College of Engineering, puts a serious emphasis on applied problem-solving. By the time students reach their senior year, they\u0026rsquo;re deep into a two-semester capstone sequence — Senior Design Studio I and II — where they spend six dedicated lab hours per week building real, sponsored projects. These aren\u0026rsquo;t toy applications. Students are working with faculty advisors and external partners, going through the same design-implement-test-validate loop you\u0026rsquo;d find in a professional engineering org.\nThe program also offers concentrations in areas like cybersecurity, software engineering, and AI, and has interdisciplinary ties to VCU\u0026rsquo;s biomedical engineering and health informatics programs. That breadth creates students who think across problem domains rather than staying narrowly within their lane — which matters enormously when you\u0026rsquo;re trying to architect cloud solutions for a complex enterprise.\nIt doesn\u0026rsquo;t hurt that VCU is a genuinely charming place to spend four years. The Monroe Park campus sits inside the Fan District — an 85-block Victorian neighborhood that claims the largest concentration of intact late-19th-century row houses in the country. Walk a few blocks in any direction and you\u0026rsquo;re into independent coffee shops, tree-lined streets, and houses that have been standing since before anyone knew what a computer was. It\u0026rsquo;s the kind of environment that makes you want to linger after a talk, and we did.\nStudents also benefit from Richmond\u0026rsquo;s growing tech ecosystem. The city is no longer just a financial services hub; it\u0026rsquo;s attracting a broader range of companies looking for engineering talent, and VCU\u0026rsquo;s career services and engineering partnerships reflect that.\nVCU vs. UNC Charlotte: Format Shapes the Room Comparing this presentation to my previous talks at UNC Charlotte\u0026rsquo;s College of Computing and Informatics was instructive — not because one program is better than the other, but because the format difference fundamentally changes who\u0026rsquo;s in the room and why.\nAt VCU, I spoke to roughly 150 students — the entire senior cohort. Everyone was at the same inflection point in their academic journey, which created a particular kind of focused energy. There was no self-selection happening; this was a scheduled session for the class, which means the full range of interests and backgrounds was represented.\nMy sessions at UNC Charlotte were structured differently. Those presentations were voluntary — students who wanted to learn more about cloud careers showed up because they specifically sought it out. Around 50 students attended, and as you\u0026rsquo;d expect from a self-selected group, the baseline curiosity about cloud was higher walking in the door. UNC Charlotte\u0026rsquo;s College of Computing and Informatics is a large program in its own right, with concentrations spanning AI/Robotics/Gaming, Data Science, and Cybersecurity. Charlotte\u0026rsquo;s position as a major fintech hub also shapes how those students think about technology; there\u0026rsquo;s an implicit finance-and-data lens that I don\u0026rsquo;t have to establish from scratch, which helps when you\u0026rsquo;re talking about cloud architecture in financial services.\nBoth groups were genuinely engaged. The difference was in the questions. The UNC Charlotte sessions had a sharper cloud-specific focus from the start — which makes sense given who opted in. At VCU, the questions were broader, and about a fifth of them centered on mainframes. I was impressed by the VCU students trying to mentally map architectures to problems.\nThe Recruiter\u0026rsquo;s Perspective I was fortunate to have Lindsey Rodriguez join us remotely to share the view from the hiring side. Lindsey\u0026rsquo;s perspective grounded the session in a way that purely technical talks can miss. Certifications matter in initial résumé screens — they signal intentionality and give a recruiter something concrete to anchor on when differentiating candidates with similar academic backgrounds. But what matters downstream is demonstrating practical problem-solving and the ability to translate technical complexity for non-technical stakeholders. That combination — technical credibility plus communication — is what gets early-career engineers into high-visibility roles quickly.\nAcknowledgments None of this would have happened without Khawlah Harahsheh, Laura Lemza, and Rebecca Kurihine, whose organization and hospitality made the day run smoothly. Thanks also to Devin Veasna for capturing photos of the session.\nI\u0026rsquo;ve already connected with several students from the day and I\u0026rsquo;m looking forward to watching them build things. There\u0026rsquo;s a particular kind of energy in a room full of people who are technically sharp and genuinely curious about the industry they\u0026rsquo;re about to enter. Sessions like this remind me why community engagement matters beyond the metrics — it\u0026rsquo;s an investment in people who are going to be shaping the stack for the next twenty years.\n","permalink":"https://lukelittle.com/posts/2026/03/speaking-at-vcu-cracking-the-cloud-and-the-evolution-of-aws-learning/","summary":"\u003cp\u003eI had the privilege of speaking at the Department of Computer Science at Virginia Commonwealth University\u0026rsquo;s College of Engineering on March 3rd, presenting \u0026ldquo;Cracking the Cloud: How AWS Certifications Can Launch Your Career\u0026rdquo; to a group of engaged Computer Science seniors.\u003c/p\u003e\n\u003ch2 id=\"the-presentation\"\u003eThe Presentation\u003c/h2\u003e\n\u003cp\u003eMy talk focused on how early-career technologists can differentiate themselves in an increasingly competitive field. Three themes anchored the session: strategic certifications — particularly AWS certifications as a way to validate cloud skills before you have the job title to back them up — self-directed projects that demonstrate practical problem-solving, and building genuine relationships in the industry. That last one always gets a laugh because I tell students it sometimes means getting away from the keyboard entirely and, as I like to say, \u0026ldquo;doing a little grass touching.\u0026rdquo;\u003c/p\u003e","title":"Speaking at VCU: Cracking the Cloud and the Evolution of AWS Learning"},{"content":"Everyone wants to ship AI into production. Almost no one wants to own what happens when it goes wrong. I\u0026rsquo;ve been in enough rooms with financial services clients to know how this plays out. A team builds something impressive on Bedrock — a RAG-powered knowledge assistant, an internal compliance copilot, a customer-facing chatbot. The demo looks great. Then someone in Legal raises their hand. What happens if it leaks a customer\u0026rsquo;s SSN? What if it makes a recommendation that sounds like investment advice? What if a clever user tricks it into ignoring your system prompt?\nThe AI project stalls. Not because the technology isn\u0026rsquo;t ready — because the governance layer isn\u0026rsquo;t there.\nAWS Bedrock Guardrails is that governance layer. And in regulated environments like banking, insurance, and healthcare, it\u0026rsquo;s not optional — it\u0026rsquo;s the prerequisite for going to production. This post walks through what Guardrails actually does, how it works under the hood, why it matters specifically in financial services, and how to implement it with code you can actually deploy.\nWhat Guardrails Solves Let\u0026rsquo;s be direct about the problem. Large language models have four failure modes that matter most in regulated industries:\nHarmful content generation: Even well-prompted models can produce hate speech, violent content, or guidance on misconduct if pushed in the right direction — especially in customer-facing contexts where you can\u0026rsquo;t predict every input.\nPrompt injection and jailbreaks: Sophisticated users will attempt to override your system prompt, bypass your application logic, or extract information from your context window that they shouldn\u0026rsquo;t have. This isn\u0026rsquo;t theoretical — it\u0026rsquo;s the first thing a red team tests.\nPII leakage: In a RAG system where your model has access to customer records, there\u0026rsquo;s a real risk of the model surfacing one customer\u0026rsquo;s information in another customer\u0026rsquo;s session, or including SSNs and account numbers in a response that gets logged, cached, or screenshotted.\nHallucination: For a general-purpose chatbot, hallucination is annoying. For a compliance assistant answering questions about regulatory requirements, it\u0026rsquo;s a liability.\nBedrock Guardrails addresses all four — as a managed layer that sits between your application and the foundation model, evaluating both input and output independently.\nHow It Works The core mental model is simple: guardrails wrap the model invocation, not the model itself. You define a set of policies once, attach them to your Bedrock calls, and every prompt and every response gets evaluated against those policies before anything reaches the end user.\nThere are two evaluation passes:\nInput evaluation runs before the prompt reaches the foundation model. If the user\u0026rsquo;s message violates a policy, the model never sees it — you get a blocked message back immediately. Output evaluation runs after the model generates a response. The model might have produced something that passes input filters but fails on output — hallucinated content that contradicts your source documents, or a response that inadvertently includes PII from the retrieved context. If either pass blocks, you get back a configurable message. The model response is never surfaced to the user.\nThe Six Policy Types 1. Content Filters Detect and filter harmful content across six categories: Hate, Insults, Sexual, Violence, Misconduct, and Prompt Attack. Each category has an adjustable filter strength — Low, Medium, or High — so you can calibrate based on your use case. A customer service chatbot for a brokerage doesn\u0026rsquo;t need the same thresholds as an internal developer tool.\nAWS extended content filtering to code-related content in 2025, which matters for any application where users can submit or request code. Harmful content in comments, variable names, and string literals is now caught at the same level as prose.\n2. Prompt Attack Detection This sits inside content filters but deserves its own callout. Jailbreaks and prompt injections are the most common adversarial inputs your application will face once it\u0026rsquo;s live. Guardrails detects both and gives you the option to block or log them — useful for incident response when your security team wants to know who tried what.\n3. Denied Topics Define topics that are off-limits in the context of your application. For a retail banking chatbot, this might be investment advice (FINRA), cryptocurrency recommendations, or competitor product comparisons. You describe the topic in plain language; AWS uses that description to classify user inputs and model responses.\nThis is one of the more powerful policy types for financial services, because it lets you draw a hard line around regulatory risk without having to enumerate every possible phrasing of a question.\n4. Sensitive Information Filters (PII Redaction) Bedrock Guardrails uses probabilistic ML detection to identify PII in both inputs and outputs. Predefined entity types include: SSN, Date of Birth, phone numbers, email addresses, credit card numbers, driver\u0026rsquo;s license numbers, bank account numbers, and more.\nFor anything not on the predefined list — like account routing numbers in a proprietary format, or internal employee IDs — you can add custom regex patterns.\nWhen PII is detected, you have two options: block the entire message, or mask the sensitive fields and allow the rest through. Masking is useful for logging and audit scenarios where you want to retain the conversation structure without storing raw PII.\n5. Contextual Grounding Checks This is the hallucination filter, and it\u0026rsquo;s the most technically interesting policy type for RAG applications.\nContextual grounding checks compare the model\u0026rsquo;s response against two things: the source documents retrieved from your knowledge base, and the user\u0026rsquo;s original query. It generates two scores:\nGrounding score: How factually consistent is the response with the source material? Relevance score: Does the response actually answer what was asked? You set a threshold between 0 and 0.99 for each. A response below either threshold gets blocked. AWS recommends starting around 0.7 for both and adjusting based on testing. In practice, this means if your compliance knowledge base says \u0026ldquo;employees must complete annual AML training,\u0026rdquo; and the model responds \u0026ldquo;employees should complete AML training within 90 days of hire\u0026rdquo; — that\u0026rsquo;s a grounding failure. The content is plausible; it\u0026rsquo;s just not what your source says. Contextual grounding catches it.\n6. Automated Reasoning Checks This is the newest and most powerful capability for factual accuracy. Where contextual grounding uses ML scoring, Automated Reasoning uses formal logic — encoding your organization\u0026rsquo;s policies as structured logical rules, then verifying model responses against those rules mathematically.\nThe practical implication: Automated Reasoning doesn\u0026rsquo;t just score a response; it can explain why a response is incorrect and what correction would make it valid. For HR policy bots, compliance Q\u0026amp;A systems, and any use case where you need to be able to show your work to an auditor, this is the capability that changes the conversation.\nImplementation Let\u0026rsquo;s make this concrete. Here\u0026rsquo;s a Terraform module that creates a guardrail configured for a financial services knowledge assistant:\nresource \u0026#34;aws_bedrock_guardrail\u0026#34; \u0026#34;finserv_assistant\u0026#34; { name = \u0026#34;finserv-knowledge-assistant\u0026#34; description = \u0026#34;Guardrails for retail banking knowledge assistant\u0026#34; blocked_input_messaging = \u0026#34;I\u0026#39;m not able to help with that request. Please contact your relationship manager for assistance.\u0026#34; blocked_outputs_messaging = \u0026#34;I wasn\u0026#39;t able to generate a response that meets our quality standards. Please rephrase your question.\u0026#34; # PII: block sensitive data in inputs, mask in outputs sensitive_information_policy_config { pii_entities_config { type = \u0026#34;SSN\u0026#34; action = \u0026#34;BLOCK\u0026#34; } pii_entities_config { type = \u0026#34;US_BANK_ACCOUNT_NUMBER\u0026#34; action = \u0026#34;ANONYMIZE\u0026#34; } pii_entities_config { type = \u0026#34;CREDIT_DEBIT_CARD_NUMBER\u0026#34; action = \u0026#34;ANONYMIZE\u0026#34; } pii_entities_config { type = \u0026#34;US_PASSPORT_NUMBER\u0026#34; action = \u0026#34;BLOCK\u0026#34; } # Custom regex for internal account IDs regexes_config { name = \u0026#34;internal-account-id\u0026#34; description = \u0026#34;Internal account reference numbers\u0026#34; pattern = \u0026#34;ACC-[0-9]{8}\u0026#34; action = \u0026#34;ANONYMIZE\u0026#34; } } # Block investment advice topics — FINRA risk mitigation topic_policy_config { topics_config { name = \u0026#34;investment-advice\u0026#34; definition = \u0026#34;Specific recommendations to buy, sell, or hold financial securities, stocks, bonds, mutual funds, ETFs, or other investment products.\u0026#34; type = \u0026#34;DENY\u0026#34; examples = [ \u0026#34;Should I buy more Apple stock?\u0026#34;, \u0026#34;What funds should I put my 401k into?\u0026#34;, \u0026#34;Is now a good time to sell my bonds?\u0026#34; ] } topics_config { name = \u0026#34;competitor-products\u0026#34; definition = \u0026#34;Comparisons or recommendations involving competing financial institutions or their products.\u0026#34; type = \u0026#34;DENY\u0026#34; } } # Content filters — calibrated for customer-facing context content_policy_config { filters_config { type = \u0026#34;HATE\u0026#34; input_strength = \u0026#34;HIGH\u0026#34; output_strength = \u0026#34;HIGH\u0026#34; } filters_config { type = \u0026#34;INSULTS\u0026#34; input_strength = \u0026#34;MEDIUM\u0026#34; output_strength = \u0026#34;MEDIUM\u0026#34; } filters_config { type = \u0026#34;PROMPT_ATTACK\u0026#34; input_strength = \u0026#34;HIGH\u0026#34; output_strength = \u0026#34;NONE\u0026#34; } filters_config { type = \u0026#34;VIOLENCE\u0026#34; input_strength = \u0026#34;HIGH\u0026#34; output_strength = \u0026#34;HIGH\u0026#34; } } # Contextual grounding — prevent hallucination in RAG responses contextual_grounding_policy_config { filters_config { type = \u0026#34;GROUNDING\u0026#34; threshold = 0.75 } filters_config { type = \u0026#34;RELEVANCE\u0026#34; threshold = 0.70 } } # Encrypt guardrail configuration with customer-managed key kms_key_arn = aws_kms_key.bedrock_guardrail_key.arn tags = { Environment = \u0026#34;production\u0026#34; Application = \u0026#34;finserv-knowledge-assistant\u0026#34; Compliance = \u0026#34;FFIEC\u0026#34; } } Now attach it to your Bedrock invocation:\nimport boto3 import json bedrock = boto3.client(\u0026#34;bedrock-runtime\u0026#34;, region_name=\u0026#34;us-east-1\u0026#34;) def invoke_with_guardrails(prompt: str, source_documents: list[str]) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34; Invoke a Bedrock model with Guardrails applied to both input and output. source_documents: list of text chunks retrieved from your knowledge base \u0026#34;\u0026#34;\u0026#34; guardrail_id = \u0026#34;your-guardrail-id\u0026#34; guardrail_version = \u0026#34;DRAFT\u0026#34; # Use a pinned version in production # Format the request with grounding source for contextual checks grounding_source = \u0026#34;\\n\\n\u0026#34;.join(source_documents) response = bedrock.invoke_model( modelId=\u0026#34;anthropic.claude-3-5-sonnet-20241022-v2:0\u0026#34;, guardrailIdentifier=guardrail_id, guardrailVersion=guardrail_version, body=json.dumps({ \u0026#34;anthropic_version\u0026#34;: \u0026#34;bedrock-2023-05-31\u0026#34;, \u0026#34;max_tokens\u0026#34;: 1024, \u0026#34;system\u0026#34;: f\u0026#39;\u0026#39;\u0026#39;You are a helpful assistant for retail banking customers. Answer questions based only on the provided documentation. If the answer is not in the documentation, say so clearly. Documentation: {grounding_source}\u0026#39;\u0026#39;\u0026#39;, \u0026#34;messages\u0026#34;: [ { \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: [ { \u0026#34;type\u0026#34;: \u0026#34;text\u0026#34;, \u0026#34;text\u0026#34;: prompt, \u0026#34;guardContent\u0026#34;: { \u0026#34;text\u0026#34;: { \u0026#34;qualifiers\u0026#34;: [\u0026#34;query\u0026#34;] } } } ] } ] }), contentType=\u0026#34;application/json\u0026#34;, accept=\u0026#34;application/json\u0026#34; ) result = json.loads(response[\u0026#34;body\u0026#34;].read()) # Check if guardrails intervened if response.get(\u0026#34;amazon-bedrock-guardrailAction\u0026#34;) == \u0026#34;INTERVENED\u0026#34;: guardrail_trace = response.get(\u0026#34;amazon-bedrock-trace\u0026#34;, {}) return { \u0026#34;blocked\u0026#34;: True, \u0026#34;reason\u0026#34;: guardrail_trace.get(\u0026#34;guardrail\u0026#34;, {}).get(\u0026#34;actionReason\u0026#34;, \u0026#34;Policy violation\u0026#34;), \u0026#34;response\u0026#34;: None } return { \u0026#34;blocked\u0026#34;: False, \u0026#34;reason\u0026#34;: None, \u0026#34;response\u0026#34;: result[\u0026#34;content\u0026#34;][0][\u0026#34;text\u0026#34;] } One thing worth calling out: the guardContent qualifier on the user message tells Guardrails which part of the prompt to evaluate for the relevance check. Without it, Guardrails would try to evaluate your entire system prompt (including the retrieved documents) as if it were the user query — which produces noisy results.\nThe Audit Trail A guardrail that blocks requests is only half the picture. The other half is knowing what it blocked, when, and why.\nEvery Guardrail invocation emits metrics to Amazon CloudWatch:\nGuardrailInvocations — total count GuardrailInterventions — how many were blocked GuardrailIntervention[PolicyType] — breakdowns by policy Set up a CloudWatch alarm on GuardrailInterventions spiking above your baseline and you have an early warning system for adversarial use or misconfigured prompts.\nFor a more complete audit trail — the kind that satisfies an FFIEC examiner or a SOC 2 auditor — route blocked events through EventBridge to an S3 bucket and query them with Athena. The pattern looks like this:\n# Lambda function triggered by EventBridge rule on Bedrock Guardrail events import boto3 import json from datetime import datetime s3 = boto3.client(\u0026#34;s3\u0026#34;) AUDIT_BUCKET = \u0026#34;your-ai-audit-logs-bucket\u0026#34; def handler(event, context): \u0026#34;\u0026#34;\u0026#34;Log guardrail interventions to immutable S3 audit trail.\u0026#34;\u0026#34;\u0026#34; audit_record = { \u0026#34;timestamp\u0026#34;: datetime.utcnow().isoformat(), \u0026#34;guardrail_id\u0026#34;: event.get(\u0026#34;guardrailId\u0026#34;), \u0026#34;policy_triggered\u0026#34;: event.get(\u0026#34;policyType\u0026#34;), \u0026#34;action\u0026#34;: event.get(\u0026#34;action\u0026#34;), # BLOCKED or ANONYMIZED \u0026#34;session_id\u0026#34;: event.get(\u0026#34;sessionId\u0026#34;), # Do NOT log the raw prompt — PII may not be fully redacted at this point \u0026#34;prompt_token_count\u0026#34;: event.get(\u0026#34;inputTokenCount\u0026#34;), \u0026#34;region\u0026#34;: event.get(\u0026#34;awsRegion\u0026#34;), \u0026#34;model_id\u0026#34;: event.get(\u0026#34;modelId\u0026#34;) } key = f\u0026#34;guardrail-interventions/{datetime.utcnow().strftime(\u0026#39;%Y/%m/%d\u0026#39;)}/{context.aws_request_id}.json\u0026#34; s3.put_object( Bucket=AUDIT_BUCKET, Key=key, Body=json.dumps(audit_record), ServerSideEncryption=\u0026#34;aws:kms\u0026#34;, BucketKeyEnabled=True ) return {\u0026#34;statusCode\u0026#34;: 200} This gives you an immutable, KMS-encrypted log of every guardrail intervention — queryable by date, policy type, model, or session ID without ever storing the raw prompt content.\nWhat This Actually Changes for Banks I keep coming back to one question in these conversations: what does it take to get an AI project from a successful proof-of-concept to a production system a compliance officer will sign off on?\nThe answer usually involves four things: data isolation, access controls, auditability, and behavioral controls. The first three are solved problems on AWS — VPC endpoints, IAM, CloudTrail. The fourth one — actually constraining what the model says and does — has historically required custom application logic that\u0026rsquo;s brittle, hard to test, and invisible to your governance team.\nBedrock Guardrails changes that. It gives you behavioral controls that are:\nCentralized. One guardrail definition applied consistently across every invocation, every session, every user. Versioned. You can pin a guardrail version to your production deployment and test changes in a draft version before promoting. Auditable. Every intervention is observable through CloudWatch metrics and loggable through EventBridge. Model-agnostic. The ApplyGuardrail API works independently of the foundation model — you can apply your guardrail to Claude, Titan, Llama, and even third-party models outside of Bedrock through the standalone API. That last point matters more than it sounds. Most banks aren\u0026rsquo;t going to standardize on a single foundation model. As the model landscape evolves, your safety policies shouldn\u0026rsquo;t have to be rewritten every time you swap out the underlying model.\nGetting Started The fastest way to get a guardrail running is through the AWS console — there\u0026rsquo;s a test playground in the Guardrails UI where you can paste prompts and verify your policies before deploying. Start there, calibrate your contextual grounding thresholds against real examples from your knowledge base, then export the configuration to Terraform or CloudFormation for your production deployment. A few things to validate before go-live:\nTest your PII detection against real data samples (anonymized). The predefined entity types work well for standard formats; you\u0026rsquo;ll discover gaps quickly with actual data. Set your contextual grounding thresholds conservatively at first (0.7/0.7) and monitor your block rate. Too many false positives means end users get frustrated; too few means you\u0026rsquo;re letting hallucinations through. Verify your denied topics by trying to phrase a restricted question a dozen different ways. The topic detection is robust, but your definition matters — vague definitions lead to both over-blocking and under-blocking. If you\u0026rsquo;re building in a regulated environment and you\u0026rsquo;re not running Guardrails, you\u0026rsquo;re carrying a liability that grows every day your AI system is in production. The capability exists. The question is whether you implement it before something goes wrong, or after.\n","permalink":"https://lukelittle.com/posts/2026/02/the-missing-layer-in-your-enterprise-ai-stack-aws-bedrock-guardrails/","summary":"\u003cp\u003eEveryone wants to ship AI into production. Almost no one wants to own what happens when it goes wrong.\nI\u0026rsquo;ve been in enough rooms with financial services clients to know how this plays out. A team builds something impressive on Bedrock — a RAG-powered knowledge assistant, an internal compliance copilot, a customer-facing chatbot. The demo looks great. Then someone in Legal raises their hand. What happens if it leaks a customer\u0026rsquo;s SSN? What if it makes a recommendation that sounds like investment advice? What if a clever user tricks it into ignoring your system prompt?\u003c/p\u003e","title":"The Missing Layer in Your Enterprise AI Stack: AWS Bedrock Guardrails"},{"content":"What does your online presence say about you?\nFor many students, searching their name online brings up little more than a LinkedIn profile and perhaps some social media accounts. This limited visibility can make it challenging to stand out professionally or showcase your actual skills and projects.\nA personal website offers a solution - providing a dedicated space where you can control your professional narrative and present your work on your own terms.\nDespite the benefits, many students don\u0026rsquo;t create personal websites due to perceived barriers:\nThinking web development skills are required Believing they need extensive projects to showcase first Assuming it\u0026rsquo;s expensive or time-consuming The student-branding-starter template addresses these concerns by providing a ready-to-use Hugo website template designed specifically for students. It can be deployed to GitHub Pages in minutes at no cost, comes pre-configured, and makes maintaining your online presence as simple as writing Markdown files.\nThis template isn\u0026rsquo;t just a technical project - it\u0026rsquo;s a foundation for your professional identity online. It creates a platform that can grow with you as you add projects, document your learning, and develop your portfolio.\nIn this article, I\u0026rsquo;ll cover what the student-branding-starter template offers, how to set it up in under 10 minutes, and how you can use it to create an effective online presence.\nWhy This Matters Having a personal website can significantly boost your professional presence as a student. Here\u0026rsquo;s why it\u0026rsquo;s worth considering:\nProfessional Presence Your online presence matters in today\u0026rsquo;s job market. When someone searches your name online, having a personal website gives you control over what they find. It supplements your LinkedIn profile and provides a more comprehensive view of your skills and interests.\nA personal website demonstrates initiative - it shows you\u0026rsquo;re willing to put in effort to present yourself professionally, which is itself a quality employers value.\nPractical Demonstration of Skills Resumes are full of claims about technical skills, but a personal website lets you demonstrate those skills with actual examples. Instead of simply listing \u0026ldquo;Python\u0026rdquo; as a skill, you can show a Python project you built, explain your approach to solving problems, and showcase your coding style.\nThis transforms abstract claims into concrete evidence of your abilities. It\u0026rsquo;s the difference between telling someone you can bake and showing them a cake you made.\nBuilding a Portfolio Over Time One significant advantage of a personal website is that it grows with you. Each project, article, or tutorial you add becomes part of your professional story.\nContent you create can serve multiple purposes:\nShowcase your technical skills Demonstrate your communication abilities Document your progress as a developer Help you remember solutions to problems you\u0026rsquo;ve solved Potentially help others facing similar challenges As you add more content over time, the value of your site increases. By graduation, you\u0026rsquo;ll have built a comprehensive portfolio that tells your story better than a resume ever could.\nAccessibility to Employers A personal site makes your work accessible to anyone interested in your skills - including potential employers. When applying for positions, you can include your website URL on your resume, LinkedIn profile, and in your email signature.\nThis gives reviewers a way to learn more about you beyond your formal application materials and may help you stand out in a competitive field.\nCost-Effective Solution The student-branding-starter template deploys to GitHub Pages, which hosts static websites completely free. There\u0026rsquo;s no monthly subscription or hidden fees.\nMost other portfolio platforms either charge recurring fees or place your content behind their branding. With this approach, you invest only time, not money - and you maintain complete ownership of your content.\nWhat You\u0026rsquo;re Building Let\u0026rsquo;s talk about what you\u0026rsquo;re actually getting with the student-branding-starter template.\nI built this after seeing students struggle with complex web development stacks just to put up a simple portfolio. We don\u0026rsquo;t need React or a database for a personal site – simpler is often better.\nThe template creates a modern, professional website built on proven technology that just works.\nThe Tech Behind It The workflow is intentionally simple:\nYou write content in Markdown files, push your changes to GitHub, and GitHub Actions automatically runs Hugo to build your site. The built site is deployed to GitHub Pages, making your content live on the web with global distribution.\nHugo generates your site – it\u0026rsquo;s blazing fast and converts simple text files to a beautiful website. GitHub handles the version control and hosting, while GitHub Actions automates the deployment process.\nThis means your workflow is dead simple. You write content in Markdown (basically plain text with some formatting). When you save and push your changes, everything rebuilds itself automatically. Your site stays live at yourusername.github.io/repository-name.\nNo databases to configure. No servers to maintain. No complicated deployment pipelines. Just write content, push, and your site updates.\nWhat Comes Pre-Built I\u0026rsquo;ve included everything students need right out of the box:\nA blog section for writing about your projects, what you\u0026rsquo;re learning, or technical concepts you want to explain. A dedicated projects page that showcases your work in a portfolio format. An about page where you tell your story and highlight your skills – often the second most visited page after your homepage.\nThere\u0026rsquo;s also a \u0026ldquo;Now\u0026rdquo; page showing what you\u0026rsquo;re currently focused on (inspired by the /now movement), and a contact page for networking opportunities.\nThe design is responsive and works beautifully on phones, tablets, and desktops. It\u0026rsquo;s SEO-optimized so Google can find and index your content. It includes both dark and light modes for better reading experience. And it even has multilingual support with English and Spanish included by default.\nSocial sharing integration helps your content spread, and the whole thing loads lightning fast, scoring 90+ on Google PageSpeed metrics.\nI\u0026rsquo;ve made it opinionated enough to look professional immediately – no design skills needed – but flexible enough that you can make it your own without getting overwhelmed by options.\nThe 10-Minute Setup The student-branding-starter template is designed for quick setup. Here\u0026rsquo;s how to get your site up and running in about 10 minutes.\nStep 1: Fork the Repository (2 minutes) First, we need to make your own copy of the template.\nHead over to github.com/lukelittle/student-branding-starter and look for the \u0026ldquo;Fork\u0026rdquo; button in the top-right corner. Click it and select your GitHub account as the destination.\nIf you want an especially clean URL, you can rename the repository to yourusername.github.io during this process, but that\u0026rsquo;s optional.\nOnce you click \u0026ldquo;Create Fork,\u0026rdquo; GitHub will take about 30 seconds to copy everything to your account.\nForking is better than starting from scratch because you get a direct copy of the repository that you can modify however you want. All the GitHub Actions workflows come pre-configured, and you maintain a connection to the original repo in case you want updates later.\nChoosing Your Version The student-branding-starter has two versions maintained in separate Git branches:\nmain branch (default): English-only multilingual branch: English + Spanish bilingual support To use the multilingual version:\nWhen forking: Before clicking \u0026ldquo;Create Fork\u0026rdquo;, use the branch dropdown to select multilingual instead of main Switching an existing fork: Clone your repository and run: git checkout multilingual git push origin multilingual --set-upstream Note: Deployment works identically via GitHub Actions regardless of which branch you choose.\nPersonal note: My blog is currently English-only, though I plan to eventually make it bilingual (English/Spanish). Since I learned computer science in English, writing technical content in Spanish is something I\u0026rsquo;m still developing.\nStep 2: Enable GitHub Pages (1 minute) Now we need to turn on the free hosting.\nGo to your newly forked repository and click the Settings tab. In the left sidebar, find and click \u0026ldquo;Pages.\u0026rdquo; Under the \u0026ldquo;Build and deployment\u0026rdquo; section, select \u0026ldquo;GitHub Actions\u0026rdquo; as your source, then save.\nThis small step tells GitHub to use the workflow file that\u0026rsquo;s already included in the template. This file handles building your site with Hugo and deploying it to GitHub Pages automatically whenever you make changes.\nStep 3: Make It Yours (5 minutes) Time to personalize your site. We\u0026rsquo;ll edit the main configuration file directly on GitHub for now.\nFind and click on the hugo.toml file in your repository, then hit the pencil icon to edit it. This file controls how your site looks and functions.\nYou\u0026rsquo;ll want to update a few key sections:\nFirst, change the baseURL to match your GitHub username and repository name:\nbaseURL = \u0026#34;https://yourusername.github.io/repository-name/\u0026#34; Next, replace the title and author information with your name:\ntitle = \u0026#34;Your Name - Personal Site\u0026#34; [params] author = \u0026#34;Your Name\u0026#34; description = \u0026#34;Student, developer, and lifelong learner.\u0026#34; Then update your social links so visitors can connect with you:\n[[params.socialIcons]] name = \u0026#34;github\u0026#34; url = \u0026#34;https://github.com/yourusername\u0026#34; [[params.socialIcons]] name = \u0026#34;linkedin\u0026#34; url = \u0026#34;https://linkedin.com/in/yourprofile\u0026#34; [[params.socialIcons]] name = \u0026#34;email\u0026#34; url = \u0026#34;mailto:your.email@example.com\u0026#34; Finally, personalize your home page information:\n[params.homeInfoParams] Title = \u0026#34;Hi, I\u0026#39;m Your Name 👋\u0026#34; Content = \u0026#34;\u0026#34;\u0026#34; I\u0026#39;m studying computer science at State University with a focus on machine learning. I write about my projects and what I\u0026#39;m learning here. **Currently:** Building a recommendation engine for my senior project. \u0026#34;\u0026#34;\u0026#34; When you\u0026rsquo;re finished, scroll down and click \u0026ldquo;Commit changes,\u0026rdquo; add a message like \u0026ldquo;Customize site with my information,\u0026rdquo; and commit again.\nThis commit triggers your first build and deployment automatically.\nStep 4: Watch the Magic Happen (1 minute) Click over to the \u0026ldquo;Actions\u0026rdquo; tab in your repository. You\u0026rsquo;ll see a workflow running – this is GitHub building your site.\nWait for the green checkmark to appear, indicating that deployment was successful. This usually takes just 1-2 minutes.\nStep 5: See Your Live Site (1 minute) Head back to Settings → Pages. You\u0026rsquo;ll see a message with your site URL: \u0026ldquo;Your site is live at https://yourusername.github.io/repository-name/\u0026quot;\nClick that link, and boom – your professional personal site is live on the internet.\nThat\u0026rsquo;s it. In less than 10 minutes, you\u0026rsquo;ve gone from nothing to having a fully functional personal website. No web development experience required.\nWriting Your First Content Once your site is live, the next step is filling it with content that demonstrates your skills and interests.\nCreating a Blog Post Blog posts are an excellent way to document your learning and showcase your problem-solving approach. Consider writing about:\nProjects you\u0026rsquo;re working on Concepts you\u0026rsquo;ve recently learned Technical challenges you\u0026rsquo;ve overcome Tools or libraries you\u0026rsquo;ve explored To create a new post, you\u0026rsquo;ll use Hugo\u0026rsquo;s command line interface:\n# Clone your repository locally git clone https://github.com/yourusername/repository-name.git cd repository-name # Install Hugo if you haven\u0026#39;t already (macOS example) brew install hugo # Create a new post hugo new posts/my-first-post.md This creates a file at content/posts/my-first-post.md with default front matter:\n--- title: \u0026#34;My First Post\u0026#34; # Post title date: 2026-02-15 # Publication date draft: true # Draft status (true = hidden) tags: [] # Categories for organizing summary: \u0026#34;\u0026#34; # Brief description --- Your content goes here... The front matter (section between --- lines) contains metadata about your post:\ntitle: The main headline date: Publication date draft: Set to false when ready to publish tags: Categories for organization summary: Brief preview shown in listings Below the front matter, you write your content using Markdown, which is a simple formatting syntax. A typical post might include headings, code blocks, lists, and images:\n## Project Overview This post documents my process building a simple data visualization tool. ### The Problem I needed to visualize temperature data collected from multiple sensors. ### The Solution I built a visualization tool using Python with Matplotlib: ```python import matplotlib.pyplot as plt import pandas as pd # Load the data data = pd.read_csv(\u0026#39;temperature_data.csv\u0026#39;) # Create the visualization plt.figure(figsize=(10, 6)) plt.plot(data[\u0026#39;timestamp\u0026#39;], data[\u0026#39;temperature\u0026#39;]) plt.title(\u0026#39;Temperature Readings Over Time\u0026#39;) plt.xlabel(\u0026#39;Time\u0026#39;) plt.ylabel(\u0026#39;Temperature (°C)\u0026#39;) plt.grid(True) plt.show() ``` When your post is ready to publish:\nSet draft: false in the front matter Commit and push to your repository: git add content/posts/my-first-post.md git commit -m \u0026#34;Add my first blog post\u0026#34; git push GitHub Actions will automatically rebuild and deploy your site with the new content.\nCreating Project Pages The Projects section of your site is designed specifically for showcasing your portfolio work. While blog posts can explain concepts or document your learning journey, project pages follow a more structured format to highlight specific work you\u0026rsquo;ve completed.\nTo create a new project page:\nhugo new projects/my-project.md This creates a file with project-specific front matter:\n--- title: \u0026#34;My Project\u0026#34; date: 2026-02-15 draft: true tags: [\u0026#34;python\u0026#34;, \u0026#34;web-dev\u0026#34;] summary: \u0026#34;A brief description of this project\u0026#34; --- ## Overview A short introduction to what this project is and why you built it. **Live Demo:** [Link to demo if available] **GitHub:** [Link to repository] ## Problem Statement What problem were you trying to solve? ## Tech Stack - Technology 1 - Technology 2 - Technology 3 ## Key Features - Feature 1: Brief description - Feature 2: Brief description - Feature 3: Brief description ## What I Learned What skills or knowledge did you gain from this project? ## Challenges What obstacles did you encounter and how did you overcome them? ## Results What was the outcome of the project? Include metrics if possible. ## Future Improvements How would you extend or improve this project given more time? The template prompts you to include important aspects of your project:\nWhat problem it solves Technologies used Key features implemented Learning outcomes Challenges encountered Results achieved Future development plans This structure helps present your work in a comprehensive way that demonstrates both technical abilities and your problem-solving approach.\nWhen documenting projects, remember that even smaller projects can be valuable portfolio pieces when you explain your process and what you learned. Class assignments, hackathon projects, and personal explorations all deserve documentation.\nOnce you\u0026rsquo;ve completed your project page, set draft: false in the front matter, commit and push to make it live.\nUpdating Your About Page The About page tells visitors who you are, what motivates you, and what skills you bring. Edit content/about.md with:\nYour educational background Technical skills and strengths Career interests and goals Personal motivations or philosophy Links to your resume or key projects This page is often the second most-visited after your homepage, so invest time in making it authentic and engaging.\nMaking It Yours Now that you have the basics working, let\u0026rsquo;s customize your site to truly make it your own.\nAdding Your Profile Photo Add your photo to the static/images/ directory Edit hugo.toml to reference it: [params] profileMode = true profilePicture = \u0026#34;/images/your-photo.jpg\u0026#34; Keep your photo professional but authentic. A simple headshot with good lighting works best.\nCustomizing Colors and Style The template uses PaperMod theme, which offers built-in customization:\nCreate a file at assets/css/extended/custom.css Add your custom CSS: :root { --primary-color: #4a89dc; --border-radius: 8px; } .post-title { font-size: 2.5rem; } This overrides the default styles without modifying theme files.\nSetting Up a Custom Domain (Optional) While yourusername.github.io/repository-name works well, a custom domain like yourname.com looks more professional:\nPurchase a domain from Namecheap, Cloudflare, Porkbun, or similar In your repository, go to Settings → Pages Under \u0026ldquo;Custom domain\u0026rdquo;, enter your domain name Set up DNS records as instructed Check \u0026ldquo;Enforce HTTPS\u0026rdquo; Update hugo.toml with your new domain: baseURL = \u0026#34;https://yourname.com/\u0026#34; For detailed instructions, see the Custom Domain section in the repository README.\nAdding Analytics (Optional) To track visitors and popular content:\nCreate a Google Analytics account Get your measurement ID (G-XXXXXXXX) Add to hugo.toml: [params.analytics] google = { SiteVerificationTag = \u0026#34;G-XXXXXXXX\u0026#34; } This gives you insights into which content resonates with visitors.\nContent Strategy for Students A common question is: \u0026ldquo;What should I write about?\u0026rdquo; Here are some content ideas that work well for student portfolio sites:\nContent Ideas 1. Class Project Documentation\nWhen documenting class projects, focus on your unique perspective:\nDesign decisions you made and why What you might do differently in future iterations Any extensions or features you added beyond requirements Specific challenges you encountered and how you solved them 2. Learning Reflections\nConsider writing about concepts you\u0026rsquo;ve recently mastered:\nExplanations of difficult topics in your own words Step-by-step guides through concepts that were challenging Visual aids or diagrams that helped your understanding Resources that were most helpful in your learning process 3. Technology Comparisons\nIf you\u0026rsquo;ve used different approaches to solve similar problems:\nCompare and contrast the technologies or methods Discuss trade-offs between different solutions Explain which scenarios might favor one approach over another Example: \u0026ldquo;Comparing React and Vue for Building a Student Dashboard\u0026rdquo; 4. Problem-Solving Documentation\nWhen you solve technical challenges:\nDocument the error or issue you encountered Explain your debugging process Share the solution you found Reflect on what you learned from the experience 5. Event Summaries\nAfter attending technical events:\nSummarize key takeaways Describe interesting projects or talks Reflect on how the content relates to your interests 6. Learning Resources Collections\nOrganize helpful resources:\nCurate lists of useful learning materials Add your own commentary on why they were valuable Organize by topic or learning objective Posting Frequency Focus on quality over quantity:\nStart with a goal of one thoughtful post per month Maintain consistency rather than posting in bursts Consider shorter, focused posts (500-800 words) that clearly explain one concept SEO Basics (In Plain English) Search Engine Optimization isn\u0026rsquo;t complex for personal sites. Focus on:\nUse descriptive titles - \u0026ldquo;How I Built a Task Management System with React\u0026rdquo; is better than \u0026ldquo;My Project\u0026rdquo; Write clear summaries - These appear in search results and social shares Use relevant tags - These help organize your content Link to relevant sources - Both internal and external Include code snippets - These add value and help with technical searches Most importantly: write naturally about topics you understand. Google rewards authentic expertise.\nBuilding an Audience For students, focus on quality over distribution:\nShare posts in relevant communities (Reddit, Discord, Slack groups) Cross-post to DEV.to or Hashnode (linking back to your site) Include your site URL in your GitHub profile Add it to your LinkedIn profile and resume Be patient—audience growth is typically slow at first and then compounds over time.\nPotential Scenarios Here\u0026rsquo;s how a personal website might benefit your job application process:\nApplication Scenario Example Without a personal website:\nYour application consists of a resume listing skills and projects Reviewer may quickly scan your GitHub repositories but might not explore deeply Limited time to evaluate your experience based on bullet points Your application may look similar to many others With a personal website:\nYour application includes a link to your personal site Reviewer can see your projects with detailed explanations Your problem-solving approach is demonstrated through your posts Your communication skills are evident in your writing Your portfolio provides specific talking points for interviews What Can Improve Your Application While every job application is different, having a personal site can potentially:\nProvide evidence of skills mentioned in your resume Showcase your communication ability Demonstrate initiative and follow-through Give interviewers specific projects to discuss Show how you think about and solve problems Content Suggestions Consider creating these types of content for your site:\nProject documentation with explanations of your approach Tutorials for concepts you\u0026rsquo;ve recently learned Technical challenges you\u0026rsquo;ve faced and how you solved them Learning resources you\u0026rsquo;ve found helpful Code samples that demonstrate your skills The most valuable content tends to be authentic, clearly explained, and shows your problem-solving approach.\nAdvanced Features Once you\u0026rsquo;re comfortable with the basics, you can take advantage of more advanced features:\nLocal Development Workflow Running Hugo locally gives you instant preview of your changes:\n# Start the Hugo development server hugo server -D # View your site at http://localhost:1313 The -D flag shows draft posts, and changes appear immediately as you save files.\nCreating Custom Sections Beyond blog posts and projects, you might want custom sections like \u0026ldquo;Notes\u0026rdquo; or \u0026ldquo;Resources\u0026rdquo;:\nCreate a directory: content/resources/ Add an _index.md file with front matter: --- title: \u0026#34;Resources\u0026#34; description: \u0026#34;Helpful resources I\u0026#39;ve collected\u0026#34; --- Introduction text here... Add to navigation in hugo.toml: [[menu.main]] identifier = \u0026#34;resources\u0026#34; name = \u0026#34;Resources\u0026#34; url = \u0026#34;/resources/\u0026#34; weight = 60 Understanding GitHub Actions Workflow The automation that deploys your site is controlled by .github/workflows/deploy.yml:\nname: Deploy Hugo site to Pages on: push: branches: [\u0026#34;main\u0026#34;] workflow_dispatch: jobs: build: runs-on: ubuntu-latest steps: - name: Checkout uses: actions/checkout@v3 with: submodules: true - name: Setup Hugo uses: peaceiris/actions-hugo@v2 - name: Build run: hugo --minify - name: Upload artifact uses: actions/upload-pages-artifact@v1 deploy: needs: build runs-on: ubuntu-latest permissions: pages: write id-token: write steps: - name: Deploy to GitHub Pages uses: actions/deploy-pages@v1 This workflow:\nChecks out your repository with theme submodules Sets up Hugo in the environment Builds your site with optimization Uploads the built site as an artifact Deploys the artifact to GitHub Pages You can customize this workflow for advanced needs like scheduled posts or multiple environments.\nPerformance Optimization The starter template is already optimized, but for even better performance:\nCompress images before uploading with tools like TinyPNG Use Hugo\u0026rsquo;s built-in asset processing for CSS and JS Lazy-load images for blog posts: ![Alt text](/images/large-image.jpg \u0026#34;Image title\u0026#34; \u0026#34;loading=lazy\u0026#34;) Minimize external dependencies and third-party scripts These optimizations help your site score well on Google\u0026rsquo;s PageSpeed metrics.\nCommon Pitfalls \u0026amp; Solutions \u0026ldquo;I don\u0026rsquo;t have anything interesting to write about\u0026rdquo; Solution: Start with what you\u0026rsquo;re learning right now. Explanation is a form of learning, and writing about a concept helps solidify your understanding. Your unique perspective makes even basic topics valuable.\nPrompt: Write about the last bug that took you more than 30 minutes to solve.\n\u0026ldquo;My writing isn\u0026rsquo;t good enough\u0026rdquo; Solution: Technical writing values clarity over style. Focus on being clear and specific. Use short sentences, include code examples, and explain your thinking. The more you write, the better you\u0026rsquo;ll get.\nPro tip: Start by outlining with bullet points, then expand each point into a paragraph.\n\u0026ldquo;No one will read my content\u0026rdquo; Solution: You\u0026rsquo;re writing for three audiences:\nFuture you (documentation of your learning) Recruiters and potential employers Other students facing similar challenges Even if the wider audience never materializes, the first two provide enough value to justify the effort.\nTechnical Issues During Setup Problem: \u0026ldquo;Theme not loading\u0026rdquo;\nSolution: Make sure you initialized the theme submodule:\ngit submodule update --init --recursive Problem: \u0026ldquo;Site shows 404 after deployment\u0026rdquo;\nSolution: Check your baseURL in hugo.toml. It must match exactly where your site is deployed (including trailing slash).\nProblem: \u0026ldquo;Changes not showing up\u0026rdquo;\nSolution: Hard refresh your browser (Ctrl+Shift+R or Cmd+Shift+R) to clear cache.\nThe Bigger Picture Building a personal site isn\u0026rsquo;t just about having a portfolio—it\u0026rsquo;s about establishing infrastructure for your career.\nBuilding Digital Assets Each post, project write-up, and tutorial you create is a digital asset that:\nDemonstrates your skills and thinking Helps others solve problems Improves your own understanding Builds your professional reputation Creates serendipitous opportunities through discovery Unlike social media posts that disappear in feeds, your personal site content has longevity and compounds in value.\nSkills Beyond the Technical Maintaining a personal site develops crucial meta-skills:\nTechnical writing: Clearly communicating complex ideas Knowledge management: Organizing and structuring information Public learning: Being comfortable showing your growth Digital presence management: Controlling your online narrative Personal brand development: Defining how you want to be known These soft skills often differentiate great engineers from good ones.\nThe Long Game Your personal site is a long-term investment with increasing returns:\nYear 1: Establish baseline presence and early portfolio Year 2: Build content library and refine voice Year 3: Leverage for internships and opportunities Year 4+: Position as industry contributor with viewpoints Students who start early have a significant advantage by graduation.\nNext Steps Ready to build your personal brand? Here\u0026rsquo;s your action plan:\nToday: Fork the repository and deploy your site This week: Customize your About page and add a profile photo Before the weekend: Write your first blog post or project page Within 30 days: Have at least 3 pieces of content published Ongoing: Aim for 1-2 new posts per month Resources to Help You Succeed Technical Writing: Google\u0026rsquo;s Technical Writing Course Markdown Guide: The Markdown Guide Content Ideas: dev.to/t/beginners for inspiration Hugo Documentation: Hugo Docs Commit to Consistency The students who see the biggest impact from their personal sites are those who commit to consistency. It\u0026rsquo;s better to publish one thoughtful post per month than to publish five posts and then abandon your site.\nSchedule time for writing just like you schedule study time. Even 30 minutes twice a week is enough to maintain momentum.\nYour personal brand isn\u0026rsquo;t just what you say about yourself—it\u0026rsquo;s what Google says about you when someone searches your name. With student-branding-starter, you\u0026rsquo;re taking control of that narrative and building a professional presence that will serve you throughout your career.\nThe best part? You can start today, for free, in just 10 minutes.\nYour future self will thank you.\n","permalink":"https://lukelittle.com/posts/2026/02/building-your-personal-brand-a-students-guide-to-online-presence/","summary":"\u003cp\u003eWhat does your online presence say about you?\u003c/p\u003e\n\u003cp\u003eFor many students, searching their name online brings up little more than a LinkedIn profile and perhaps some social media accounts. This limited visibility can make it challenging to stand out professionally or showcase your actual skills and projects.\u003c/p\u003e\n\u003cp\u003eA personal website offers a solution - providing a dedicated space where you can control your professional narrative and present your work on your own terms.\u003c/p\u003e","title":"Building Your Personal Brand: A Student's Guide to Online Presence"},{"content":"Introduction On August 1, 2012, Knight Capital Group—one of the largest market makers on the New York Stock Exchange—lost $440 million in 45 minutes due to a software deployment failure. The incident nearly bankrupted the firm and sent shockwaves through financial markets. While the technical details are fascinating, the real lesson lies in what wasn\u0026rsquo;t there: an effective, centralized mechanism to stop runaway automation before catastrophic losses occurred.\nThis post explores how modern streaming architectures using Apache Kafka and Apache Spark can implement the kind of real-time risk controls that regulations now require—and that Knight Capital desperately needed. We\u0026rsquo;ll connect the dots between a historic trading disaster, regulatory requirements, and a hands-on demo you can deploy yourself.\nThe Knight Capital Incident: What Happened? On that August morning, Knight Capital deployed new trading software to eight servers. Due to an operational error, one server retained old code that had been repurposed. When the market opened, this server began executing a dormant algorithm called \u0026ldquo;Power Peg\u0026rdquo; that was never meant to run in production.\nThe result was catastrophic:\nThe algorithm sent millions of unintended orders to the market Knight accumulated massive, unwanted positions in 154 stocks The firm lost $440 million in approximately 45 minutes Knight Capital required a $400 million emergency bailout to survive The Core Failure Modes Several factors contributed to the disaster:\nPartial Deployment: Not all servers received the correct code update Lack of Centralized Control: No single point could halt all trading activity Insufficient Pre-Trade Controls: Orders weren\u0026rsquo;t validated against risk limits before execution Delayed Detection: The problem wasn\u0026rsquo;t identified and stopped quickly enough The Knight Capital incident wasn\u0026rsquo;t just a software bug—it was a systems design failure. The firm lacked the architectural patterns needed to maintain \u0026ldquo;direct and exclusive control\u0026rdquo; over its market access, a concept that would soon become central to regulatory requirements.\nEnter SEC Rule 15c3-5: The Market Access Rule In response to concerns about the risks posed by direct market access and algorithmic trading, the SEC adopted Rule 15c3-5 in November 2010 (before the Knight incident, though Knight\u0026rsquo;s failure validated the rule\u0026rsquo;s necessity).\nWhat is Market Access? Market access refers to the ability to send orders directly to exchanges or alternative trading systems. Broker-dealers that provide market access—whether for their own trading or for customers—act as gatekeepers to the markets.\nWhat the Rule Requires SEC Rule 15c3-5, formally titled \u0026ldquo;Risk Management Controls for Brokers or Dealers with Market Access,\u0026rdquo; requires broker-dealers to:\nImplement Risk Management Controls: Establish, document, and maintain a system of risk management controls and supervisory procedures reasonably designed to manage the financial, regulatory, and other risks of market access.\nPre-Trade Controls: Implement controls that prevent the entry of orders that exceed appropriate pre-set credit or capital thresholds, or that appear to be erroneous.\nDirect and Exclusive Control: Broker-dealers must have \u0026ldquo;direct and exclusive control\u0026rdquo; over the technology that provides market access. This means they cannot delegate control to customers or third parties—they must retain the ability to stop trading immediately.\nRegular Review: Controls must be reviewed and tested regularly to ensure they\u0026rsquo;re working as intended.\nThe \u0026ldquo;Direct and Exclusive Control\u0026rdquo; Concept This phrase is critical. It means:\nThe broker-dealer must be able to disable or limit market access immediately Control cannot be delegated to customers or outsourced There must be a centralized mechanism to enforce risk limits The firm must maintain supervisory procedures over all market access The rule doesn\u0026rsquo;t prescribe specific technologies (it doesn\u0026rsquo;t say \u0026ldquo;you must use Kafka\u0026rdquo;), but it does mandate capabilities that modern streaming architectures are well-suited to provide.\nRegulatory Text From the SEC\u0026rsquo;s adopting release:\n\u0026ldquo;The rule requires a broker-dealer with market access to establish, document, and maintain a system of risk management controls and supervisory procedures reasonably designed to manage the financial, regulatory, and other risks of this business activity.\u0026rdquo;\nThe rule specifically addresses:\nFinancial risk management (credit and capital thresholds) Regulatory risk management (compliance with regulatory requirements) Erroneous order controls (preventing clearly erroneous orders from reaching the market) Connecting Regulation to Architecture Let\u0026rsquo;s translate regulatory requirements into architectural patterns:\nRegulatory Requirement Architectural Pattern Our Demo Implementation Direct and exclusive control Centralized kill switch with authoritative state Kafka compacted topic for kill state Pre-trade risk controls Real-time order validation before routing Order router checks kill state Prevent erroneous orders Automated detection of anomalous patterns Spark streaming detects threshold breaches Supervisory procedures Audit trail and manual override capability Audit topic + operator console API Regular review and testing Observable, testable system CloudWatch dashboards + demo scripts The key insight: Separation of concerns between detection and enforcement.\nDetection (Spark): Analyzes order patterns, computes risk signals, may suggest kill actions Enforcement (Router): Makes the final decision on every order based on authoritative state Control Plane (Kill Switch): Maintains single source of truth for kill status Audit (Kafka + DynamoDB): Immutable record of all decisions This separation ensures that even if detection fails or is delayed, enforcement remains consistent. The kill switch state is authoritative and replayable.\nWhy Kafka Compaction for Kill State? One of the most interesting architectural choices in our demo is using a Kafka compacted topic for kill switch state. Here\u0026rsquo;s why:\nThe Problem We need a \u0026ldquo;configuration\u0026rdquo; or \u0026ldquo;state\u0026rdquo; that:\nIs the single source of truth Can be updated in real-time Is immediately available to all consumers Has a complete audit trail Can be replayed to bootstrap new services The Solution: Log Compaction Kafka\u0026rsquo;s log compaction retains the latest value for each key while preserving the full history of changes. For kill switch state:\nCompaction Example: Commands Topic (Full History): ACCOUNT:12345 KILL t=100 ACCOUNT:12345 UNKILL t=200 ACCOUNT:12345 KILL t=300 State Topic (Compacted - Latest Only): ACCOUNT:12345 KILL t=300 How it works:\nKey: Scope (e.g., \u0026ldquo;ACCOUNT:12345\u0026rdquo;, \u0026ldquo;SYMBOL:AAPL\u0026rdquo;, \u0026ldquo;GLOBAL\u0026rdquo;) Value: Current status (KILLED or ACTIVE) with metadata Compaction: Kafka automatically retains only the latest state per scope Replayability: New consumers can read the entire topic to bootstrap current state This gives us:\nSingle source of truth: The compacted topic is authoritative Fast bootstrap: New routers can quickly load all current kill states Audit trail: The commands topic retains full history Distributed config: No need for external config store State Compaction Process The following sequence diagram illustrates how Kafka\u0026rsquo;s log compaction maintains the latest state per scope:\nsequenceDiagram participant K1 as Kafka\u0026lt;br/\u0026gt;killswitch.commands.v1\u0026lt;br/\u0026gt;(Full History) participant KSA as Kill Switch\u0026lt;br/\u0026gt;Aggregator participant K2 as Kafka\u0026lt;br/\u0026gt;killswitch.state.v1\u0026lt;br/\u0026gt;(Compacted) participant KC as Kafka\u0026lt;br/\u0026gt;Compaction Process participant OR as Order Router\u0026lt;br/\u0026gt;(New Instance) Note over K1,OR: State Evolution Over Time Note over K1: t=100 K1-\u0026gt;\u0026gt;KSA: KILL command\u0026lt;br/\u0026gt;ACCOUNT:12345 KSA-\u0026gt;\u0026gt;K2: Publish state\u0026lt;br/\u0026gt;Key: ACCOUNT:12345\u0026lt;br/\u0026gt;Value: KILLED (t=100) Note over K1: t=200 K1-\u0026gt;\u0026gt;KSA: UNKILL command\u0026lt;br/\u0026gt;ACCOUNT:12345 KSA-\u0026gt;\u0026gt;K2: Publish state\u0026lt;br/\u0026gt;Key: ACCOUNT:12345\u0026lt;br/\u0026gt;Value: ACTIVE (t=200) Note over K1: t=300 K1-\u0026gt;\u0026gt;KSA: KILL command\u0026lt;br/\u0026gt;ACCOUNT:12345 KSA-\u0026gt;\u0026gt;K2: Publish state\u0026lt;br/\u0026gt;Key: ACCOUNT:12345\u0026lt;br/\u0026gt;Value: KILLED (t=300) Note over K2: Before Compaction:\u0026lt;br/\u0026gt;ACCOUNT:12345 KILLED (t=100)\u0026lt;br/\u0026gt;ACCOUNT:12345 ACTIVE (t=200)\u0026lt;br/\u0026gt;ACCOUNT:12345 KILLED (t=300) K2-\u0026gt;\u0026gt;KC: Compaction triggered\u0026lt;br/\u0026gt;(based on segment.ms\u0026lt;br/\u0026gt;and dirty ratio) KC-\u0026gt;\u0026gt;KC: Retain latest value\u0026lt;br/\u0026gt;per key Note over K2: After Compaction:\u0026lt;br/\u0026gt;ACCOUNT:12345 KILLED (t=300)\u0026lt;br/\u0026gt;(older values removed) Note over OR: New router starts up OR-\u0026gt;\u0026gt;K2: Read from beginning K2--\u0026gt;\u0026gt;OR: ACCOUNT:12345 = KILLED (t=300) Note over OR: Router bootstrapped\u0026lt;br/\u0026gt;with current state\u0026lt;br/\u0026gt;(fast, no history to read) Note over K1: Commands topic still has\u0026lt;br/\u0026gt;full history for audit Compaction Configuration cleanup.policy=compact min.cleanable.dirty.ratio=0.01 # Compact frequently segment.ms=60000 # Small segments for faster compaction These settings ensure kill state updates propagate quickly while maintaining the full history in the commands topic.\nDemo Architecture Walkthrough Our demo implements these patterns using serverless AWS services:\nKey architectural decisions mapped to regulatory requirements:\nRequirement Implementation Direct and exclusive control Operator Console API with manual override capability Pre-trade risk controls Order Router checks kill state before routing every order Prevent erroneous orders Spark detects anomalous patterns in real-time Audit trail Immutable Kafka log + DynamoDB index for queries Supervisory procedures Documented thresholds, operator actions, correlation IDs Normal Order Flow The following sequence diagram illustrates how orders flow through the system when no kill switches are active:\nsequenceDiagram participant OG as Order Generator\u0026lt;br/\u0026gt;(Lambda) participant K1 as Kafka\u0026lt;br/\u0026gt;orders.v1 participant S as Spark\u0026lt;br/\u0026gt;Risk Detector participant K2 as Kafka\u0026lt;br/\u0026gt;risk_signals.v1 participant OR as Order Router\u0026lt;br/\u0026gt;(Lambda) participant KS as Kafka\u0026lt;br/\u0026gt;killswitch.state.v1 participant K3 as Kafka\u0026lt;br/\u0026gt;orders.gated.v1 participant K4 as Kafka\u0026lt;br/\u0026gt;audit.v1 participant DDB as DynamoDB\u0026lt;br/\u0026gt;Audit Index Note over OG,DDB: Normal Operation - No Kill Switches Active OG-\u0026gt;\u0026gt;K1: Publish order\u0026lt;br/\u0026gt;(5 orders/sec) Note right of K1: Key: account_id\u0026lt;br/\u0026gt;Partition by account K1-\u0026gt;\u0026gt;S: Consume orders S-\u0026gt;\u0026gt;S: Compute 60s window\u0026lt;br/\u0026gt;order_count = 50\u0026lt;br/\u0026gt;notional = $500K Note right of S: Below thresholds:\u0026lt;br/\u0026gt;order_rate \u0026lt; 100\u0026lt;br/\u0026gt;notional \u0026lt; $1M S-\u0026gt;\u0026gt;K2: Publish risk signal\u0026lt;br/\u0026gt;(metrics only, no alert) K1-\u0026gt;\u0026gt;OR: Consume order OR-\u0026gt;\u0026gt;KS: Check kill state\u0026lt;br/\u0026gt;for ACCOUNT:12345 KS--\u0026gt;\u0026gt;OR: No kill state found\u0026lt;br/\u0026gt;(ACTIVE by default) Note over OR: Decision: ALLOW OR-\u0026gt;\u0026gt;K3: Forward order\u0026lt;br/\u0026gt;to gated topic OR-\u0026gt;\u0026gt;K4: Publish audit event\u0026lt;br/\u0026gt;decision=ALLOW OR-\u0026gt;\u0026gt;DDB: Write audit record\u0026lt;br/\u0026gt;(async, best effort) Note over OG,DDB: Order successfully routed Topic Flow orders.v1: Raw orders from generator risk_signals.v1: Windowed aggregations from Spark (order rate, notional, concentration) killswitch.commands.v1: Kill/unkill commands (from Spark or operator) killswitch.state.v1: Authoritative kill state (compacted) orders.gated.v1: Orders that passed kill switch check audit.v1: Immutable audit trail of all routing decisions Order Router Enforcement The following diagram shows the detailed logic of how the order router enforces kill switches:\nsequenceDiagram participant K1 as Kafka\u0026lt;br/\u0026gt;orders.v1 participant OR as Order Router\u0026lt;br/\u0026gt;(Lambda) participant Cache as In-Memory\u0026lt;br/\u0026gt;Kill State Cache participant K2 as Kafka\u0026lt;br/\u0026gt;killswitch.state.v1 participant K3 as Kafka\u0026lt;br/\u0026gt;orders.gated.v1 participant K4 as Kafka\u0026lt;br/\u0026gt;audit.v1 participant DDB as DynamoDB\u0026lt;br/\u0026gt;Audit Index Note over K1,DDB: Order Router Processing Logic K1-\u0026gt;\u0026gt;OR: Consume order\u0026lt;br/\u0026gt;account_id: 12345\u0026lt;br/\u0026gt;symbol: AAPL OR-\u0026gt;\u0026gt;Cache: Check kill state\u0026lt;br/\u0026gt;for scopes Note over Cache: Check hierarchy:\u0026lt;br/\u0026gt;1. GLOBAL\u0026lt;br/\u0026gt;2. ACCOUNT:12345\u0026lt;br/\u0026gt;3. SYMBOL:AAPL alt GLOBAL kill active Cache--\u0026gt;\u0026gt;OR: GLOBAL = KILLED Note over OR: Decision: DROP\u0026lt;br/\u0026gt;Reason: Global kill else ACCOUNT kill active Cache--\u0026gt;\u0026gt;OR: ACCOUNT:12345 = KILLED Note over OR: Decision: DROP\u0026lt;br/\u0026gt;Reason: Account kill else SYMBOL kill active Cache--\u0026gt;\u0026gt;OR: SYMBOL:AAPL = KILLED Note over OR: Decision: DROP\u0026lt;br/\u0026gt;Reason: Symbol kill else No kills active Cache--\u0026gt;\u0026gt;OR: All scopes ACTIVE Note over OR: Decision: ALLOW OR-\u0026gt;\u0026gt;K3: Forward order end OR-\u0026gt;\u0026gt;K4: Publish audit event\u0026lt;br/\u0026gt;decision: ALLOW/DROP\u0026lt;br/\u0026gt;scope_matches: [...]\u0026lt;br/\u0026gt;corr_id: uuid-789 OR-\u0026gt;\u0026gt;DDB: Write audit record\u0026lt;br/\u0026gt;(async, best effort) Note over K2,OR: State updates arrive K2-\u0026gt;\u0026gt;OR: New state update OR-\u0026gt;\u0026gt;Cache: Update in-memory cache Note over Cache: Cache always reflects\u0026lt;br/\u0026gt;latest compacted state Latency Considerations This design introduces additional latency relative to in-process risk checks. However, it provides centralized, authoritative enforcement and replayable state — properties essential for satisfying the \u0026ldquo;direct and exclusive control\u0026rdquo; requirement of SEC Rule 15c3-5.\nIn most retail and DMA (Direct Market Access) environments, the added milliseconds (approximately 50ms at most) are an acceptable tradeoff for deterministic control and auditability. This approach is not intended to reflect the architecture of any former employer but rather examines how brokerages can solve these regulatory challenges in a robust, scalable way.\nKill Switch Activation Sequence Here\u0026rsquo;s what happens when Spark detects a threshold breach:\nsequenceDiagram participant OG as Order Generator\u0026lt;br/\u0026gt;(Lambda) participant K1 as Kafka\u0026lt;br/\u0026gt;orders.v1 participant S as Spark\u0026lt;br/\u0026gt;Risk Detector participant K2 as Kafka\u0026lt;br/\u0026gt;risk_signals.v1 participant K3 as Kafka\u0026lt;br/\u0026gt;killswitch.commands.v1 participant KSA as Kill Switch\u0026lt;br/\u0026gt;Aggregator (Lambda) participant K4 as Kafka\u0026lt;br/\u0026gt;killswitch.state.v1 participant DDB as DynamoDB\u0026lt;br/\u0026gt;State Cache participant OR as Order Router\u0026lt;br/\u0026gt;(Lambda) participant K5 as Kafka\u0026lt;br/\u0026gt;audit.v1 Note over OG,K5: Panic Mode Triggered OG-\u0026gt;\u0026gt;K1: Publish orders\u0026lt;br/\u0026gt;(50 orders/sec) Note right of K1: High rate for\u0026lt;br/\u0026gt;ACCOUNT:12345 K1-\u0026gt;\u0026gt;S: Consume orders S-\u0026gt;\u0026gt;S: Compute 60s window\u0026lt;br/\u0026gt;order_count = 150\u0026lt;br/\u0026gt;notional = $2.5M Note over S: BREACH DETECTED!\u0026lt;br/\u0026gt;order_count \u0026gt; 100 S-\u0026gt;\u0026gt;K2: Publish risk signal\u0026lt;br/\u0026gt;with breach flag S-\u0026gt;\u0026gt;K3: Publish KILL command\u0026lt;br/\u0026gt;scope: ACCOUNT:12345\u0026lt;br/\u0026gt;reason: \u0026#34;Order rate breach\u0026#34;\u0026lt;br/\u0026gt;corr_id: uuid-123 Note over K3: Commands topic\u0026lt;br/\u0026gt;(full history retained) K3-\u0026gt;\u0026gt;KSA: Consume KILL command KSA-\u0026gt;\u0026gt;KSA: Process command\u0026lt;br/\u0026gt;Create state record KSA-\u0026gt;\u0026gt;K4: Publish state\u0026lt;br/\u0026gt;Key: ACCOUNT:12345\u0026lt;br/\u0026gt;Value: KILLED\u0026lt;br/\u0026gt;corr_id: uuid-123 Note right of K4: Compacted topic\u0026lt;br/\u0026gt;(latest state per key) KSA-\u0026gt;\u0026gt;DDB: Update state cache\u0026lt;br/\u0026gt;(optional, for fast lookup) Note over K4,OR: State propagates to all routers K4-\u0026gt;\u0026gt;OR: Router reads state update OR-\u0026gt;\u0026gt;OR: Update in-memory cache\u0026lt;br/\u0026gt;ACCOUNT:12345 = KILLED K1-\u0026gt;\u0026gt;OR: New order from 12345 OR-\u0026gt;\u0026gt;OR: Check kill state\u0026lt;br/\u0026gt;ACCOUNT:12345 = KILLED Note over OR: Decision: DROP OR-\u0026gt;\u0026gt;K5: Publish audit event\u0026lt;br/\u0026gt;decision=DROP\u0026lt;br/\u0026gt;reason: \u0026#34;Kill switch active\u0026#34;\u0026lt;br/\u0026gt;corr_id: uuid-123 Note over OG,K5: Order blocked - Kill switch active Key observations:\nDetection (Spark) is decoupled from enforcement (Router) State updates flow through compacted topic (single source of truth) Every decision is audited with correlation IDs Manual override capability (operator can unkill) Spark SQL for Risk Detection The Spark job uses Spark SQL for windowed aggregations:\nSELECT window(event_time, \u0026#39;60 seconds\u0026#39;) as window, account_id, COUNT(*) as order_count, SUM(qty * price) as total_notional, COUNT(DISTINCT symbol) as unique_symbols FROM orders GROUP BY window(event_time, \u0026#39;60 seconds\u0026#39;), account_id When thresholds are breached, Spark emits a kill command:\n{ \u0026#34;cmd_id\u0026#34;: \u0026#34;uuid\u0026#34;, \u0026#34;scope\u0026#34;: \u0026#34;ACCOUNT:12345\u0026#34;, \u0026#34;action\u0026#34;: \u0026#34;KILL\u0026#34;, \u0026#34;reason\u0026#34;: \u0026#34;Order rate breach: 150 orders in 60s\u0026#34;, \u0026#34;triggered_by\u0026#34;: \u0026#34;spark\u0026#34;, \u0026#34;metric\u0026#34;: \u0026#34;order_rate_60s\u0026#34;, \u0026#34;value\u0026#34;: 150 } Enforcement Logic The order router maintains an in-memory cache of kill state (bootstrapped from the compacted topic) and checks every order:\ndef check_kill_status(order): scopes = [ \u0026#39;GLOBAL\u0026#39;, f\u0026#39;ACCOUNT:{order[\u0026#34;account_id\u0026#34;]}\u0026#39;, f\u0026#39;SYMBOL:{order[\u0026#34;symbol\u0026#34;]}\u0026#39; ] for scope in scopes: if scope in kill_state and kill_state[scope][\u0026#39;status\u0026#39;] == \u0026#39;KILLED\u0026#39;: return True, scope, kill_state[scope][\u0026#39;reason\u0026#39;] return False, None, None Every decision is audited with correlation IDs for traceability.\nAlternative Approaches for High-Frequency Trading For high-frequency trading environments, the architecture described above would introduce unacceptable latency. In these ultra-low-latency scenarios, the pre-trade risk gate would be embedded directly in the order handling process—potentially implemented in hardware (e.g., FPGA)—to ensure deterministic, microsecond-level enforcement without introducing network or broker latency.\nKey differences in HFT implementations:\nEmbedded Controls: Risk checks directly in the order path, not as external services Hardware Acceleration: FPGAs or dedicated ASICs for microsecond-level checks Local State: State maintained in local memory with minimal or no network calls Minimal Serialization: Custom binary protocols instead of JSON Deterministic Performance: Bounded, predictable latency for all operations In such environments, you wouldn\u0026rsquo;t add Kafka to the hot path, as even the most optimized message broker would introduce unacceptable latency. Instead, while still maintaining the regulatory requirements for \u0026ldquo;direct and exclusive control,\u0026rdquo; risk configurations would be loaded at startup and updated via side channels, with enforcement happening directly within the order processing pipeline.\nWhy This Matters for Students This demo teaches several critical concepts:\nEvent-Driven Architecture: Using Kafka as the backbone for real-time systems Stream Processing: Spark Structured Streaming for windowed aggregations Separation of Concerns: Detection vs. enforcement vs. control Operational Patterns: Compaction, idempotency, correlation IDs Regulatory Thinking: How compliance requirements shape architecture Serverless at Scale: Building production-grade systems without managing servers Most importantly, it connects abstract concepts (regulations, risk management) to concrete implementations you can deploy and experiment with.\nTry It Yourself The complete demo is available in the repository. You can:\nDeploy to AWS: Full serverless stack with Terraform Run locally: Docker Compose for quick iteration Experiment: Change thresholds, add new scopes, implement throttling Learn: Detailed workshop docs with exercises See the repository for step-by-step instructions.\nSuggested Exercises Add a SYMBOL-level kill switch that triggers on concentration Implement throttling (rate limiting) instead of binary kill/allow Add deduplication to prevent duplicate order IDs Build a dashboard to visualize risk signals in real-time Implement automatic unkill after a cooldown period Conclusion The Knight Capital incident taught the industry a painful lesson about the importance of centralized control and pre-trade risk management. SEC Rule 15c3-5 codified these lessons into regulatory requirements that all broker-dealers must follow.\nModern streaming architectures using Kafka and Spark provide elegant solutions to these requirements:\nKafka\u0026rsquo;s compacted topics give us authoritative, replayable state Spark\u0026rsquo;s streaming SQL enables real-time risk detection Separation of detection and enforcement ensures consistent control Immutable audit trails provide full traceability While this demo uses synthetic data and simplified logic, the architectural patterns are production-grade. Real broker-dealers use similar approaches to maintain the \u0026ldquo;direct and exclusive control\u0026rdquo; that regulations require and that Knight Capital lacked.\nThe next time you hear about a trading glitch or market disruption, ask: \u0026ldquo;Where was the kill switch?\u0026rdquo;\nSources and Further Reading Primary Regulatory Sources SEC Rule 15c3-5 Final Adopting Release\nSecurities and Exchange Commission, Release No. 34-63241 (November 3, 2010)\nhttps://www.sec.gov/files/rules/final/2010/34-63241.pdf\nSEC Small Entity Compliance Guide for Rule 15c3-5\nhttps://www.sec.gov/files/rules/final/2010/34-63241-secg.htm\nCode of Federal Regulations: 17 CFR § 240.15c3-5\nhttps://www.law.cornell.edu/cfr/text/17/240.15c3-5\nKnight Capital Incident SEC Administrative Proceeding Against Knight Capital\nFile No. 3-15570 (October 16, 2013)\nDetails the regulatory findings and penalties related to the incident.\nNanex Research: Knight Capital\u0026rsquo;s Trading Glitch\nTechnical analysis of the order flow during the incident (secondary source).\nTechnical Resources Apache Kafka Documentation: Log Compaction\nhttps://kafka.apache.org/documentation/#compaction\nApache Spark Structured Streaming Guide\nhttps://spark.apache.org/docs/latest/structured-streaming-programming-guide.html\nDisclaimer: This blog post and associated demo are for educational purposes only. They do not constitute trading advice, legal advice, or compliance guidance. The architecture described does not represent any former employer\u0026rsquo;s actual systems or implementations. The demo uses synthetic data and simplified logic to illustrate concepts rather than real production implementations. Actual production trading systems require extensive additional controls, testing, and regulatory review. Always consult with legal and compliance professionals when implementing market access systems.\n","permalink":"https://lukelittle.com/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eOn August 1, 2012, Knight Capital Group—one of the largest market makers on the New York Stock Exchange—lost $440 million in 45 minutes due to a software deployment failure. The incident nearly bankrupted the firm and sent shockwaves through financial markets. While the technical details are fascinating, the real lesson lies in what wasn\u0026rsquo;t there: an effective, centralized mechanism to stop runaway automation before catastrophic losses occurred.\u003c/p\u003e\n\u003cp\u003eThis post explores how modern streaming architectures using Apache Kafka and Apache Spark can implement the kind of real-time risk controls that regulations now require—and that Knight Capital desperately needed. We\u0026rsquo;ll connect the dots between a historic trading disaster, regulatory requirements, and a hands-on demo you can deploy yourself.\u003c/p\u003e","title":"Designing Pre-Trade Risk Controls on AWS (SEC Rule 15c3-5)"},{"content":"Managing AWS costs becomes increasingly complex as infrastructure grows. Organizations often struggle with cloud cost management, spending valuable engineering time manually analyzing Cost Explorer data, identifying optimization opportunities, and implementing changes. Even with dedicated cost management tools, the analysis and remediation process remains largely manual, requiring specialized expertise to interpret cost data and translate it into actionable steps.\nThis post demonstrates how to build an automated agent that analyzes AWS costs and generates actionable recommendations to reduce cloud spend. By combining AWS Bedrock\u0026rsquo;s analytical capabilities with Cost Explorer data, the system identifies cost outliers and provides specific optimization steps that go beyond basic visualizations to deliver meaningful insights.\nWhat we\u0026rsquo;re building A cost optimization agent that:\nRuns on a schedule (weekly or monthly) Fetches detailed cost data from AWS Cost Explorer Analyzes spending patterns and identifies optimization opportunities Generates a markdown report with specific recommendations Sends a summary to Slack or email Tracks recommendations and their potential savings This solution goes beyond basic cost visualization by providing specific, actionable steps to optimize AWS spending—essentially turning data into decisions.\nReal-World Applications This solution addresses cost management challenges across different contexts:\nEnterprise FinOps Teams: In larger organizations, the agent provides consistent, ongoing cost analysis that augments the FinOps team\u0026rsquo;s capabilities, ensuring no optimization opportunity goes unnoticed even as the infrastructure grows in complexity.\nStartups and Small Teams: For organizations without dedicated cloud financial analysts, the agent provides expert-level cost optimization recommendations that would otherwise require specialized knowledge or expensive consultants.\nManaged Service Providers: MSPs can deploy the agent across client environments, standardizing cost optimization practices while customizing thresholds and priorities for each client\u0026rsquo;s specific needs.\nDevelopment Environments: The agent can enforce stricter cost controls in non-production environments, identifying development and testing resources that can be safely downsized, scheduled, or terminated to reduce costs without affecting production workloads.\nMulti-Cloud Strategies: While this implementation focuses on AWS, the architecture pattern can be extended to analyze costs across multiple cloud providers, giving organizations a unified view of optimization opportunities.\nArchitecture overview Here\u0026rsquo;s the high-level architecture:\nEventBridge (scheduled) → Lambda → Bedrock Agent with Action Groups → S3 (report) → SNS (notifications) The key components:\nEventBridge: Triggers the agent on a schedule Lambda: Initializes and coordinates the analysis process Bedrock Agent: Orchestrates the data gathering and analysis Action Groups: Custom Lambda functions for specific tasks S3: Stores the generated reports SNS: Sends notifications with the summary DynamoDB: Tracks recommendations and their implementation status This event-driven architecture ensures the cost optimization process runs automatically on schedule, eliminating the need for manual intervention and ensuring consistent analysis.\nWhat you\u0026rsquo;ll need AWS Account with Bedrock and Cost Explorer access IAM role with appropriate permissions S3 bucket for storing reports SNS topic or email for notifications Basic understanding of AWS services Step 1: Create the Action Group Lambda functions First, let\u0026rsquo;s create the Lambda functions our Bedrock Agent will use as action groups:\n1. Cost Data Retrieval Lambda This function fetches comprehensive cost data from AWS Cost Explorer:\nimport json import os import boto3 import datetime from dateutil.relativedelta import relativedelta def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) time_period = payload.get(\u0026#39;time_period\u0026#39;, \u0026#39;MONTH\u0026#39;) services = payload.get(\u0026#39;services\u0026#39;, []) # Configure time period based on request end_date = datetime.datetime.now() if time_period == \u0026#39;MONTH\u0026#39;: start_date = end_date - relativedelta(months=1) granularity = \u0026#39;DAILY\u0026#39; elif time_period == \u0026#39;QUARTER\u0026#39;: start_date = end_date - relativedelta(months=3) granularity = \u0026#39;MONTHLY\u0026#39; elif time_period == \u0026#39;WEEK\u0026#39;: start_date = end_date - relativedelta(weeks=1) granularity = \u0026#39;DAILY\u0026#39; else: # Default to monthly start_date = end_date - relativedelta(months=1) granularity = \u0026#39;DAILY\u0026#39; # Format dates for Cost Explorer start_str = start_date.strftime(\u0026#39;%Y-%m-%d\u0026#39;) end_str = end_date.strftime(\u0026#39;%Y-%m-%d\u0026#39;) # Initialize Cost Explorer client ce_client = boto3.client(\u0026#39;ce\u0026#39;) # Get cost data, anomalies, and recommendations response = get_cost_data(ce_client, start_str, end_str, granularity) anomalies = get_anomalies(ce_client, start_str, end_str) reservation_recs = get_reservation_recommendations(ce_client) savings_plans_recs = get_savings_plans_recommendations(ce_client) # Return the compiled data return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({ \u0026#39;cost_data\u0026#39;: response, \u0026#39;anomalies\u0026#39;: anomalies, \u0026#39;reservation_recommendations\u0026#39;: reservation_recs, \u0026#39;savings_plans_recommendations\u0026#39;: savings_plans_recs }) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} # Helper functions (implementation details omitted for brevity) def get_cost_data(ce_client, start_str, end_str, granularity): # Implementation details omitted return {} def get_anomalies(ce_client, start_str, end_str): # Implementation details omitted return [] def get_reservation_recommendations(ce_client): # Implementation details omitted return [] def get_savings_plans_recommendations(ce_client): # Implementation details omitted return [] 2. Resource Analysis Lambda This function analyzes AWS resources for optimization opportunities:\nimport json import boto3 import datetime def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) resource_types = payload.get(\u0026#39;resource_types\u0026#39;, [\u0026#39;ec2\u0026#39;, \u0026#39;rds\u0026#39;, \u0026#39;ebs\u0026#39;, \u0026#39;lambda\u0026#39;]) results = {} # Analyze each resource type if \u0026#39;ec2\u0026#39; in resource_types: results[\u0026#39;ec2\u0026#39;] = analyze_ec2_instances() if \u0026#39;rds\u0026#39; in resource_types: results[\u0026#39;rds\u0026#39;] = analyze_rds_instances() if \u0026#39;ebs\u0026#39; in resource_types: results[\u0026#39;ebs\u0026#39;] = analyze_ebs_volumes() if \u0026#39;lambda\u0026#39; in resource_types: results[\u0026#39;lambda\u0026#39;] = analyze_lambda_functions() return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps(results) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} def analyze_ec2_instances(): ec2_client = boto3.client(\u0026#39;ec2\u0026#39;) cloudwatch = boto3.client(\u0026#39;cloudwatch\u0026#39;) # Get instances and identify optimization opportunities instances = get_all_instances(ec2_client) low_utilization = find_low_utilization_instances(instances, cloudwatch) potential_downsizing = find_downsizing_opportunities(instances, cloudwatch) return { \u0026#39;total_instances\u0026#39;: len(instances), \u0026#39;running_instances\u0026#39;: count_running_instances(instances), \u0026#39;stopped_instances\u0026#39;: count_stopped_instances(instances), \u0026#39;low_utilization_instances\u0026#39;: low_utilization, \u0026#39;potential_downsizing\u0026#39;: potential_downsizing } # Helper functions (implementation details omitted for brevity) def get_all_instances(ec2_client): # Implementation details omitted return [] def find_low_utilization_instances(instances, cloudwatch): # Implementation details omitted return [] def find_downsizing_opportunities(instances, cloudwatch): # Implementation details omitted return [] def count_running_instances(instances): # Implementation details omitted return 0 def count_stopped_instances(instances): # Implementation details omitted return 0 def analyze_rds_instances(): # Implementation details omitted return {} def analyze_ebs_volumes(): # Implementation details omitted return {} def analyze_lambda_functions(): # Implementation details omitted return {} 3. Report Generation Lambda This function generates cost optimization reports and sends notifications:\nimport json import boto3 import os import time from datetime import datetime def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) cost_data = payload.get(\u0026#39;cost_data\u0026#39;, {}) resource_analysis = payload.get(\u0026#39;resource_analysis\u0026#39;, {}) destination = payload.get(\u0026#39;destination\u0026#39;, \u0026#39;S3\u0026#39;) # Generate detailed markdown report report_content = generate_markdown_report(cost_data, resource_analysis) # Create a unique filename with timestamp timestamp = datetime.now().strftime(\u0026#39;%Y-%m-%d-%H-%M-%S\u0026#39;) filename = f\u0026#34;cost-optimization-report-{timestamp}.md\u0026#34; results = {} # Save to S3 if requested if destination in [\u0026#39;S3\u0026#39;, \u0026#39;BOTH\u0026#39;]: s3_url = save_to_s3(report_content, filename) results[\u0026#39;s3_url\u0026#39;] = s3_url # Send notification if requested if destination in [\u0026#39;SNS\u0026#39;, \u0026#39;BOTH\u0026#39;]: summary = generate_summary(cost_data, resource_analysis) if \u0026#39;s3_url\u0026#39; in results: summary += f\u0026#34;\\n\\nDetailed report: {results[\u0026#39;s3_url\u0026#39;]}\u0026#34; send_notification(summary, timestamp) results[\u0026#39;sns_notification\u0026#39;] = \u0026#39;Sent\u0026#39; # Store recommendations in DynamoDB for tracking store_recommendations(cost_data, resource_analysis) return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps(results) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} # Helper functions (implementation details omitted for brevity) def generate_markdown_report(cost_data, resource_analysis): # Implementation details omitted return \u0026#34;\u0026#34; def save_to_s3(report_content, filename): # Implementation details omitted return \u0026#34;\u0026#34; def generate_summary(cost_data, resource_analysis): # Implementation details omitted return \u0026#34;\u0026#34; def send_notification(summary, timestamp): # Implementation details omitted pass def store_recommendations(cost_data, resource_analysis): # Implementation details omitted pass Step 2: Set up DynamoDB for tracking recommendations You\u0026rsquo;ll need DynamoDB tables to track reports and recommendations. In production, you\u0026rsquo;d define these in your infrastructure-as-code using Terraform or CloudFormation. The tables need:\nCostOptimizationRecommendations table:\nPartition key: recommendation_id (String) PAY_PER_REQUEST billing mode for cost efficiency CostOptimizationReports table:\nPartition key: report_id (String) PAY_PER_REQUEST billing mode for cost efficiency The first table stores individual recommendations with their implementation status, while the second table tracks metadata about generated reports.\nStep 3: Create the Bedrock Agent Now let\u0026rsquo;s create the agent that will orchestrate the entire analysis process:\nIn the Bedrock console, go to \u0026ldquo;Agents\u0026rdquo; → \u0026ldquo;Create agent\u0026rdquo;\nName it \u0026ldquo;CostOptimizationAgent\u0026rdquo;\nSelect Claude 3.5 Sonnet for the foundation model\nCreate three action groups:\na. GetCostData\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;GetCostData\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Retrieves cost and usage data from AWS Cost Explorer\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: Cost Explorer API\\n version: 1.0.0\\npaths:\\n /getCostData:\\n post:\\n summary: Get cost and usage data from AWS Cost Explorer\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n properties:\\n time_period:\\n type: string\\n description: The time period to analyze (WEEK, MONTH, QUARTER)\\n services:\\n type: array\\n items:\\n type: string\\n description: Optional filter for specific AWS services\\n responses:\\n 200:\\n description: Successful response with cost data\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-COST-DATA-LAMBDA-ARN]\u0026#34; } } b. AnalyzeResources\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;AnalyzeResources\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Analyzes AWS resources for optimization opportunities\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: Resource Analysis API\\n version: 1.0.0\\npaths:\\n /analyzeResources:\\n post:\\n summary: Analyze AWS resources for optimization opportunities\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n properties:\\n resource_types:\\n type: array\\n items:\\n type: string\\n description: Resource types to analyze, e.g., \u0026#39;ec2\u0026#39;, \u0026#39;rds\u0026#39;, \u0026#39;ebs\u0026#39;, \u0026#39;lambda\u0026#39;\\n responses:\\n 200:\\n description: Successful analysis of resources\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-RESOURCE-ANALYSIS-LAMBDA-ARN]\u0026#34; } } c. GenerateReport\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;GenerateReport\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Generates a cost optimization report and sends notifications\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: Report Generation API\\n version: 1.0.0\\npaths:\\n /generateReport:\\n post:\\n summary: Generate a cost optimization report and send notifications\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n required:\\n - cost_data\\n - resource_analysis\\n properties:\\n cost_data:\\n type: object\\n description: Cost and usage data from Cost Explorer\\n resource_analysis:\\n type: object\\n description: Results of resource analysis\\n destination:\\n type: string\\n description: Where to send the report (S3, SNS, or BOTH)\\n responses:\\n 200:\\n description: Successful report generation\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-REPORT-GENERATION-LAMBDA-ARN]\u0026#34; } } Configure the agent\u0026rsquo;s instructions:\nYou are a helpful cost optimization agent for AWS. Your purpose is to analyze AWS costs and resource usage to identify potential savings opportunities. When asked to analyze costs: 1. Get cost and usage data using GetCostData action 2. Analyze resources for optimization opportunities using AnalyzeResources action 3. Generate a report of findings and recommendations using GenerateReport action Your recommendations should be practical and actionable, focusing on: - EC2 instance optimization (rightsizing, stopping idle instances) - Unattached or underutilized EBS volumes - Rarely used Lambda functions - Reserved Instance or Savings Plans opportunities - Multi-AZ configurations that might not be needed for non-production Present findings clearly in order of potential savings, with the highest-impact opportunities first. These instructions are crucial as they define how the agent will behave when analyzing costs. The careful wording ensures it focuses on the most impactful optimization opportunities.\nStep 4: Create the main Lambda function Create the main Lambda function that will be triggered by the schedule:\nimport json import os import boto3 import logging import time # Initialize Bedrock Runtime client bedrock_agent_runtime = boto3.client(\u0026#39;bedrock-agent-runtime\u0026#39;) logger = logging.getLogger() logger.setLevel(logging.INFO) def lambda_handler(event, context): try: # Get parameters from the event time_period = event.get(\u0026#39;time_period\u0026#39;, \u0026#39;MONTH\u0026#39;) # Invoke the Bedrock agent response = bedrock_agent_runtime.invoke_agent( agentId=os.environ[\u0026#39;BEDROCK_AGENT_ID\u0026#39;], agentAliasId=os.environ[\u0026#39;BEDROCK_AGENT_ALIAS_ID\u0026#39;], sessionId=f\u0026#34;cost-analysis-{int(time.time())}\u0026#34;, inputText=f\u0026#34;Analyze AWS costs for the past {time_period.lower()}, look for optimization opportunities, and generate a comprehensive report with specific actionable recommendations.\u0026#34; ) # Process agent response completion = \u0026#39;\u0026#39; for event in response.get(\u0026#39;completion\u0026#39;, []): chunk = json.loads(event[\u0026#39;chunk\u0026#39;][\u0026#39;bytes\u0026#39;].decode()) if chunk[\u0026#39;type\u0026#39;] == \u0026#39;message\u0026#39;: completion += chunk[\u0026#39;message\u0026#39;][\u0026#39;content\u0026#39;][0][\u0026#39;text\u0026#39;] return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;Cost analysis completed\u0026#39;}) } except Exception as e: logger.error(f\u0026#34;Error: {str(e)}\u0026#34;) return { \u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)}) } Step 5: Set up the EventBridge rule Create an EventBridge rule to run the analysis on a schedule. In a production environment, you\u0026rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation, setting parameters like:\nRule name: \u0026ldquo;WeeklyCostOptimizationAnalysis\u0026rdquo; Schedule expression: \u0026ldquo;cron(0 8 ? * MON *)\u0026rdquo; (runs every Monday at 8 AM) Target: Your main Lambda function EventBridge ensures the cost analysis runs automatically at your chosen interval without manual intervention.\nStep 6: Set up SNS for notifications Set up an SNS topic for notifications by creating a topic and adding subscribers. In production, you\u0026rsquo;d define this in your infrastructure-as-code, configuring:\nTopic name: \u0026ldquo;CostOptimizationAlerts\u0026rdquo; Protocol: Email, SMS, or webhook (based on your preferred notification channel) Subscribers: Finance team, cloud administrators, or a Slack webhook The notification system ensures key stakeholders are informed of optimization opportunities as they\u0026rsquo;re identified.\nAnalysis Capabilities The agent can identify several types of cost optimization opportunities:\n1. EC2 Instance Optimization Idle Instances: Identifies running instances with CPU utilization consistently below 5% Rightsizing Opportunities: Finds instances that could be downsized based on utilization patterns Stopped Instances: Locates instances that have been stopped for extended periods Instance Family Upgrades: Suggests moving to newer generation instance families Reserved Instance Coverage: Identifies on-demand instances that should be covered by RIs 2. Storage Optimization Unattached EBS Volumes: Finds volumes not attached to instances Old Snapshots: Identifies EBS snapshots older than 6 months Underutilized Volumes: Locates volumes with consistently low I/O patterns Storage Class Transitions: Recommends moving infrequently accessed data to lower-cost storage tiers 3. Database Optimization Overprovisioned RDS Instances: Identifies oversized database instances Multi-AZ in Development: Flags multi-AZ deployments in non-production environments Idle Databases: Finds database instances with minimal connection counts Reserved Instance Opportunities: Suggests RIs for stable database workloads 4. Serverless Optimization Overallocated Memory: Identifies Lambda functions with excessive memory allocation Rarely Used Functions: Finds functions that are rarely invoked but consume resources Long-Running Functions: Suggests optimizations for functions that consistently run long Cost considerations This solution is cost-efficient:\nLambda costs: Most usage will fall under the free tier EventBridge: No additional cost for scheduled rules Bedrock API: ~$0.015 per 1,000 tokens with Claude Sonnet S3: Minimal storage costs for reports DynamoDB: Pay-per-request pricing keeps costs very low SNS: Practically free for email notifications For a weekly execution schedule, the infrastructure costs typically remain under $5 per month for most organizations due to the minimal compute resources required.\nExtending the solution Here are some ways to enhance your cost optimization agent:\nMulti-account analysis: Extend to analyze costs across an AWS Organization. This offers a comprehensive view of spending across your entire cloud estate.\nImplementation tracking: Track which recommendations were implemented and their actual savings. This helps quantify the ROI of the optimization agent.\nAutomated remediation: Add capability to automatically implement low-risk optimizations like removing unattached EBS volumes. The agent could implement changes automatically during off-hours.\nSlack integration: Send reports directly to Slack channels, enabling team discussions around cost optimization opportunities and tagging responsible teams.\nTagging compliance: Check for resources without proper cost allocation tags, ensuring your organization maintains visibility into spend by department, team, or project.\nBudget alerts integration: Combine cost optimization with proactive budget alerts, automatically triggering more aggressive analysis when a budget threshold is approaching.\nCustom thresholds: Allow different teams or environments to set custom thresholds for what constitutes underutilization based on their specific workload patterns.\nConclusion This solution leverages several AWS services to create an intelligent cost optimization system that analyzes cloud spending and provides specific recommendations for reducing costs.\nKey advantages of this approach include:\nAutomation: Regular, scheduled analysis without manual intervention Actionable insights: Specific recommendations rather than just data visualization Comprehensive coverage: Analysis across multiple resource types (EC2, RDS, Lambda, etc.) Prioritization: Recommendations sorted by potential impact Integration: Works with existing AWS services and notification systems The architecture combines the data collection capabilities of AWS Cost Explorer with the analytical power of Amazon Bedrock to generate insights similar to those from a cloud financial analyst. By implementing this solution, organizations can transform cost management from a periodic, manual exercise into an ongoing, automated process.\nThe system is particularly effective at identifying unused resources, rightsizing opportunities, and reservation/Savings Plans recommendations - areas that often yield significant savings when properly optimized.\n","permalink":"https://lukelittle.com/posts/2026/02/building-a-cost-optimization-agent-with-aws-bedrock-and-cost-explorer/","summary":"\u003cp\u003eManaging AWS costs becomes increasingly complex as infrastructure grows. Organizations often struggle with cloud cost management, spending valuable engineering time manually analyzing Cost Explorer data, identifying optimization opportunities, and implementing changes. Even with dedicated cost management tools, the analysis and remediation process remains largely manual, requiring specialized expertise to interpret cost data and translate it into actionable steps.\u003c/p\u003e\n\u003cp\u003eThis post demonstrates how to build an automated agent that analyzes AWS costs and generates actionable recommendations to reduce cloud spend. By combining AWS Bedrock\u0026rsquo;s analytical capabilities with Cost Explorer data, the system identifies cost outliers and provides specific optimization steps that go beyond basic visualizations to deliver meaningful insights.\u003c/p\u003e","title":"Building a Cost Optimization Agent with AWS Bedrock and Cost Explorer"},{"content":"Code reviews are essential for maintaining code quality, but they can be time-consuming and often repetitive. Developers find themselves commenting on the same issues across multiple pull requests: missing tests, inconsistent naming, inadequate error handling, and numerous other routine concerns. This creates a bottleneck in the development process, as team members wait for their code to be reviewed while reviewers struggle to balance thorough reviews with their own development work.\nAn AI assistant can address this challenge by analyzing pull requests before human reviewers, catching common issues and allowing the team to focus on more complex aspects of the code review. This approach doesn\u0026rsquo;t replace human judgment but enhances it, ensuring that routine issues are caught consistently while freeing up developer time for deeper analysis.\nThis post demonstrates how to build an automated PR reviewer using AWS Bedrock Agents that analyzes code changes and provides feedback directly in GitHub.\nWhat we\u0026rsquo;re building A PR review agent that:\nGets triggered automatically when a new PR is opened or updated Fetches the PR diff from GitHub Analyzes the changes using AWS Bedrock Posts a detailed review comment with suggestions Tracks review history in DynamoDB The goal isn\u0026rsquo;t to replace human reviewers, but to complement them by identifying common issues, style violations, and potential bugs before human review begins.\nReal-World Applications This solution addresses common development challenges across different contexts:\nEnterprise Development Teams: In large organizations with strict coding standards, the agent ensures consistency across hundreds of developers, reducing the burden on senior engineers who often shoulder the bulk of review responsibilities.\nOpen Source Projects: Maintainers can use the PR reviewer to handle the initial assessment of community contributions, ensuring they meet project guidelines before dedicating their limited time to review.\nEducational Settings: Computer science programs can deploy the agent to provide students with immediate feedback on their code submissions, helping them learn best practices without requiring instructor intervention for every issue.\nContinuous Integration Pipelines: The PR reviewer can become part of a broader CI/CD strategy, providing code quality assessments alongside traditional test runs and builds.\nOnboarding New Team Members: New developers on a project receive immediate feedback on their work that helps them understand team coding standards more quickly, accelerating their integration into the team.\nArchitecture overview Here\u0026rsquo;s how the system works:\nGitHub Webhook → EventBridge → Lambda → Bedrock Agent with Action Groups → GitHub API → PR Comments The key components:\nGitHub Webhook: Triggers on PR events EventBridge: Routes events to Lambda Lambda: Processes GitHub events and invokes Bedrock Agent Bedrock Agent: Coordinates the review process Action Groups: Custom Lambda functions that the agent can call DynamoDB: Tracks review history and status This event-driven architecture ensures the review process begins automatically whenever a pull request is opened or updated, without requiring any manual intervention.\nWhat you\u0026rsquo;ll need AWS Account with Bedrock access GitHub repository where you want to enable reviews GitHub Personal Access Token with repo permissions Basic familiarity with AWS and GitHub Step 1: Create the Action Group Lambda functions First, let\u0026rsquo;s create the Lambda functions our Bedrock Agent will use as action groups. We need three functions:\n1. Get PR Diff Lambda This function fetches the pull request details and diff from GitHub:\nimport json import os import boto3 import requests from github import Github def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) repo_name = payload[\u0026#39;repo_name\u0026#39;] pr_number = payload[\u0026#39;pr_number\u0026#39;] # Initialize GitHub g = Github(os.environ[\u0026#39;GITHUB_TOKEN\u0026#39;]) repo = g.get_repo(repo_name) pr = repo.get_pull(int(pr_number)) # Get PR details and diff pr_title = pr.title pr_description = pr.body or \u0026#34;\u0026#34; pr_author = pr.user.login pr_files = [f.filename for f in pr.get_files()] diff_url = f\u0026#34;https://github.com/{repo_name}/pull/{pr_number}.diff\u0026#34; headers = {\u0026#39;Authorization\u0026#39;: f\u0026#34;token {os.environ[\u0026#39;GITHUB_TOKEN\u0026#39;]}\u0026#34;} diff_response = requests.get(diff_url, headers=headers) diff = diff_response.text # Return everything the agent needs return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({ \u0026#39;pr_title\u0026#39;: pr_title, \u0026#39;pr_description\u0026#39;: pr_description, \u0026#39;pr_author\u0026#39;: pr_author, \u0026#39;pr_files\u0026#39;: pr_files, \u0026#39;pr_diff\u0026#39;: diff }) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} 2. Analyze Code Lambda This function sends the code diff to Bedrock for analysis:\nimport json import os import boto3 def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) diff = payload[\u0026#39;diff\u0026#39;] languages = payload.get(\u0026#39;languages\u0026#39;, []) # Create prompt for the LLM prompt = f\u0026#34;\u0026#34;\u0026#34; Please review the following code diff and provide actionable feedback: {diff} In your analysis, please look for: 1. Potential bugs or logic errors 2. Security vulnerabilities 3. Performance issues 4. Code style and best practices 5. Missing tests or documentation Format your response as: ## Summary (Brief overview of the changes and their purpose) ## Critical Issues (List any serious problems that must be fixed) ## Suggestions (List minor issues and improvements) ## Positive Notes (Highlight good practices in the code) \u0026#34;\u0026#34;\u0026#34; # Call Bedrock for analysis bedrock = boto3.client(\u0026#39;bedrock-runtime\u0026#39;) response = bedrock.invoke_model( modelId=\u0026#39;anthropic.claude-3-sonnet-20240229-v1:0\u0026#39;, contentType=\u0026#39;application/json\u0026#39;, accept=\u0026#39;application/json\u0026#39;, body=json.dumps({ \u0026#34;anthropic_version\u0026#34;: \u0026#34;bedrock-2023-05-31\u0026#34;, \u0026#34;max_tokens\u0026#34;: 2000, \u0026#34;messages\u0026#34;: [{\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: prompt}] }) ) # Parse and return response response_body = json.loads(response[\u0026#39;body\u0026#39;].read()) analysis = response_body[\u0026#39;content\u0026#39;][0][\u0026#39;text\u0026#39;] return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;analysis\u0026#39;: analysis}) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} 3. Post Comment Lambda This function posts the review feedback to the GitHub PR and records the review in DynamoDB:\nimport json import os import boto3 import time from github import Github def lambda_handler(event, context): try: # Extract parameters from the event payload = json.loads(event[\u0026#39;body\u0026#39;]) repo_name = payload[\u0026#39;repo_name\u0026#39;] pr_number = payload[\u0026#39;pr_number\u0026#39;] comment = payload[\u0026#39;comment\u0026#39;] # Record review in DynamoDB dynamodb = boto3.resource(\u0026#39;dynamodb\u0026#39;) table = dynamodb.Table(os.environ[\u0026#39;REVIEW_TABLE\u0026#39;]) review_id = f\u0026#34;{repo_name}-{pr_number}-{int(time.time())}\u0026#34; table.put_item( Item={ \u0026#39;review_id\u0026#39;: review_id, \u0026#39;repo_name\u0026#39;: repo_name, \u0026#39;pr_number\u0026#39;: pr_number, \u0026#39;timestamp\u0026#39;: int(time.time()), \u0026#39;comment\u0026#39;: comment } ) # Post comment to GitHub g = Github(os.environ[\u0026#39;GITHUB_TOKEN\u0026#39;]) repo = g.get_repo(repo_name) pr = repo.get_pull(int(pr_number)) # Format the comment for GitHub formatted_comment = f\u0026#34;\u0026#34;\u0026#34; ## AI Code Review 🤖 {comment} --- *This review was automatically generated by the PR Review Agent. [Learn more](https://example.com/pr-agent)* \u0026#34;\u0026#34;\u0026#34; pr.create_issue_comment(formatted_comment) return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({ \u0026#39;review_id\u0026#39;: review_id, \u0026#39;status\u0026#39;: \u0026#39;Comment posted successfully\u0026#39; }) } except Exception as e: return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} Step 2: Set up DynamoDB for tracking reviews You\u0026rsquo;ll need a DynamoDB table to track review history. In production, you\u0026rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation. The table needs:\nPartition key: review_id (String) PAY_PER_REQUEST billing mode for cost efficiency This table will store metadata about each review, including repository, PR number, timestamp, and the content of the review comment.\nStep 3: Create the Bedrock Agent Now let\u0026rsquo;s create the agent that will orchestrate the entire review process:\nIn the Bedrock console, go to \u0026ldquo;Agents\u0026rdquo; → \u0026ldquo;Create agent\u0026rdquo;\nName it \u0026ldquo;PRReviewAgent\u0026rdquo;\nSelect Claude 3.5 Sonnet for the foundation model\nCreate three action groups:\na. GetPRDiff\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;GetPRDiff\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Retrieves the diff for a GitHub pull request\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: GitHub PR Diff API\\n version: 1.0.0\\npaths:\\n /getPRDiff:\\n post:\\n summary: Get the diff for a GitHub pull request\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n required:\\n - repo_name\\n - pr_number\\n properties:\\n repo_name:\\n type: string\\n description: The repository in format \u0026#39;owner/repo\u0026#39;\\n pr_number:\\n type: integer\\n description: The pull request number\\n responses:\\n 200:\\n description: Successful response\\n content:\\n application/json:\\n schema:\\n type: object\\n properties:\\n pr_title:\\n type: string\\n pr_description:\\n type: string\\n pr_author:\\n type: string\\n pr_files:\\n type: array\\n items:\\n type: string\\n pr_diff:\\n type: string\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-GET-PR-DIFF-LAMBDA-ARN]\u0026#34; } } b. AnalyzeCode\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;AnalyzeCode\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Analyzes code for potential issues and improvements\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: Code Analysis API\\n version: 1.0.0\\npaths:\\n /analyzeCode:\\n post:\\n summary: Analyze code diff for issues and suggestions\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n required:\\n - diff\\n properties:\\n diff:\\n type: string\\n description: The code diff to analyze\\n languages:\\n type: array\\n items:\\n type: string\\n description: Programming languages in the diff (optional)\\n responses:\\n 200:\\n description: Successful analysis\\n content:\\n application/json:\\n schema:\\n type: object\\n properties:\\n analysis:\\n type: string\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-ANALYZE-CODE-LAMBDA-ARN]\u0026#34; } } c. PostComment\n{ \u0026#34;actionGroupName\u0026#34;: \u0026#34;PostComment\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;Posts a review comment to a GitHub pull request\u0026#34;, \u0026#34;apiSchema\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;openapi\u0026#34;, \u0026#34;payload\u0026#34;: \u0026#34;openapi: 3.0.0\\ninfo:\\n title: GitHub Comment API\\n version: 1.0.0\\npaths:\\n /postComment:\\n post:\\n summary: Post a comment to a GitHub pull request\\n requestBody:\\n required: true\\n content:\\n application/json:\\n schema:\\n type: object\\n required:\\n - repo_name\\n - pr_number\\n - comment\\n properties:\\n repo_name:\\n type: string\\n description: The repository in format \u0026#39;owner/repo\u0026#39;\\n pr_number:\\n type: integer\\n description: The pull request number\\n comment:\\n type: string\\n description: The comment text to post\\n responses:\\n 200:\\n description: Successful response\\n content:\\n application/json:\\n schema:\\n type: object\\n properties:\\n review_id:\\n type: string\\n status:\\n type: string\u0026#34; }, \u0026#34;actionGroupExecutor\u0026#34;: { \u0026#34;type\u0026#34;: \u0026#34;lambda\u0026#34;, \u0026#34;lambdaArn\u0026#34;: \u0026#34;[YOUR-POST-COMMENT-LAMBDA-ARN]\u0026#34; } } Configure the agent\u0026rsquo;s instructions:\nYou are a helpful GitHub Pull Request reviewer designed to analyze code changes and provide constructive feedback. Your goal is to help developers improve their code by identifying issues and suggesting improvements. When a PR is submitted: 1. Get the PR diff using the GetPRDiff action 2. Analyze the code for issues using the AnalyzeCode action 3. Post a helpful review comment using the PostComment action Your reviews should be constructive and educational. Focus on: - Potential bugs or logic errors - Security vulnerabilities - Performance issues - Code style and best practices - Missing tests or documentation Use a professional tone and be concise but thorough in your feedback. These instructions are crucial as they define how the agent will behave when reviewing code. The careful wording encourages constructive feedback while maintaining a professional tone.\nStep 4: Create the main Lambda function Now create the main Lambda function that will be triggered by GitHub events and coordinate the review process:\nimport json import os import boto3 import logging # Initialize Bedrock Runtime client bedrock_agent_runtime = boto3.client(\u0026#39;bedrock-agent-runtime\u0026#39;) logger = logging.getLogger() logger.setLevel(logging.INFO) def lambda_handler(event, context): try: # Parse the GitHub webhook event github_event = json.loads(event[\u0026#39;body\u0026#39;]) headers = event.get(\u0026#39;headers\u0026#39;, {}) # Check if it\u0026#39;s a pull request event we care about if headers.get(\u0026#39;X-GitHub-Event\u0026#39;) == \u0026#39;pull_request\u0026#39;: action = github_event.get(\u0026#39;action\u0026#39;) # Process only opened or synchronized (updated) PRs if action in (\u0026#39;opened\u0026#39;, \u0026#39;synchronize\u0026#39;): pr = github_event[\u0026#39;pull_request\u0026#39;] repo_name = github_event[\u0026#39;repository\u0026#39;][\u0026#39;full_name\u0026#39;] pr_number = pr[\u0026#39;number\u0026#39;] # Invoke the Bedrock agent to perform the review response = bedrock_agent_runtime.invoke_agent( agentId=os.environ[\u0026#39;BEDROCK_AGENT_ID\u0026#39;], agentAliasId=os.environ[\u0026#39;BEDROCK_AGENT_ALIAS_ID\u0026#39;], sessionId=f\u0026#34;{repo_name}-{pr_number}\u0026#34;, inputText=f\u0026#34;Review pull request #{pr_number} in repository {repo_name}\u0026#34; ) # Process agent response completion = \u0026#39;\u0026#39; for event in response.get(\u0026#39;completion\u0026#39;, []): chunk = json.loads(event[\u0026#39;chunk\u0026#39;][\u0026#39;bytes\u0026#39;].decode()) if chunk[\u0026#39;type\u0026#39;] == \u0026#39;message\u0026#39;: completion += chunk[\u0026#39;message\u0026#39;][\u0026#39;content\u0026#39;][0][\u0026#39;text\u0026#39;] return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;Agent invoked successfully\u0026#39;}) } return {\u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;Ignored PR action\u0026#39;})} return {\u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;Ignored webhook event\u0026#39;})} except Exception as e: logger.error(f\u0026#34;Error: {str(e)}\u0026#34;) return {\u0026#39;statusCode\u0026#39;: 500, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;error\u0026#39;: str(e)})} Step 5: Create the API Gateway Create a new REST API in API Gateway Add a POST method for the root resource Set the integration type to Lambda Function and select your main function Deploy the API to a stage and note the URL The API Gateway serves as the entry point for GitHub webhooks, receiving events when pull requests are opened or updated.\nStep 6: Set up the GitHub webhook Go to your GitHub repository Navigate to Settings → Webhooks → Add webhook Set the Payload URL to your API Gateway URL Set Content type to application/json Select \u0026ldquo;Let me select individual events\u0026rdquo; and check \u0026ldquo;Pull requests\u0026rdquo; Click \u0026ldquo;Add webhook\u0026rdquo; With the webhook in place, GitHub will now notify your system whenever pull requests are created or updated, triggering the automated review process.\nTesting the system To test your PR review agent:\nCreate a simple PR with some code changes Check that the webhook is triggered (visible in GitHub webhook settings) Verify your Lambda is invoked (check CloudWatch logs) Check that a comment appears on the PR with the review You should see a comprehensive review comment that identifies potential issues, makes suggestions for improvement, and highlights positive aspects of the code changes.\nEnhancing the agent Here are some ways to improve your PR reviewer:\nLanguage-specific rules: Extend the AnalyzeCode lambda to apply language-specific linters or static analysis tools for more precise feedback based on the programming language.\nContextual awareness: Include repository history, architecture documentation, or team standards in the analysis to provide more relevant and contextual suggestions.\nReview customization: Allow teams to set review focus areas through configuration. For example, some teams might prioritize security checks while others emphasize performance.\nLearning from feedback: Track which suggestions developers implement versus ignore to improve future recommendations and reduce false positives.\nPR metadata analysis: Consider PR size, files changed, and complexity when determining the review strategy. Large PRs might receive different handling than small, focused changes.\nInline comments: Enhance the agent to post specific comments on individual lines in the diff rather than just a summary comment.\nPre-commit integration: Offer developers the ability to run the same analysis locally before submitting their PR, using pre-commit hooks.\nCost considerations This architecture is cost-efficient for most development teams:\nLambda costs: Typically covered by the free tier for small to medium teams API Gateway: Approximately $1 per million requests Bedrock API: Around $0.015 per 1,000 tokens with Claude Sonnet DynamoDB: Pay-per-request pricing with minimal storage needs The total cost will vary based on your team\u0026rsquo;s PR volume and the size of code changes being reviewed, but for most teams, it remains quite affordable compared to the developer time saved.\nSecurity considerations When implementing this:\nStore API tokens in AWS Secrets Manager Set IAM permissions using least privilege Consider the sensitivity of code being reviewed Remember that agents aren\u0026rsquo;t perfect at security analysis While the PR reviewer can identify many common security issues, it should not be your only security control. Critical security reviews should still be performed by security experts, especially for sensitive components.\nConclusion By combining AWS Bedrock Agents with Action Groups, this architecture creates an intelligent PR review system that helps maintain code quality while reducing the time developers spend on repetitive aspects of code reviews.\nThe solution provides several key benefits:\nConsistency: Every PR receives the same baseline level of review Early detection: Catches common issues before human reviewers see the code Educational feedback: Provides context and explanations for suggested improvements Focus: Allows human reviewers to concentrate on architecture and business logic Historical tracking: Stores reviews for future reference and improvement While this implementation uses specific AWS services, the architectural patterns can be adapted to other cloud providers or integrated with different version control systems beyond GitHub.\nThis approach to automated code review represents a practical application of generative AI that delivers immediate value to development teams: better code quality, faster reviews, and more time for developers to focus on creative and complex challenges rather than routine feedback.\n","permalink":"https://lukelittle.com/posts/2026/02/building-a-github-pr-reviewer-with-bedrock-agents-and-action-groups/","summary":"\u003cp\u003eCode reviews are essential for maintaining code quality, but they can be time-consuming and often repetitive. Developers find themselves commenting on the same issues across multiple pull requests: missing tests, inconsistent naming, inadequate error handling, and numerous other routine concerns. This creates a bottleneck in the development process, as team members wait for their code to be reviewed while reviewers struggle to balance thorough reviews with their own development work.\u003c/p\u003e","title":"Building a GitHub PR Reviewer with Bedrock Agents and Action Groups"},{"content":"\u0026ldquo;Where can I find our vacation policy?\u0026rdquo; \u0026ldquo;What\u0026rsquo;s the process for requesting new hardware?\u0026rdquo; \u0026ldquo;Can you explain our security guidelines?\u0026rdquo; These questions echo through company Slack channels daily, interrupting workflows and creating redundant work for team leads and HR staff. The same questions get asked repeatedly, and answers are buried in documentation that\u0026rsquo;s difficult to navigate.\nIn this post, I\u0026rsquo;ll show you how to build a simple yet powerful Q\u0026amp;A bot for Slack that leverages your company\u0026rsquo;s documentation to provide accurate, contextual answers. The best part? It runs entirely on AWS managed services, minimizing operational overhead while delivering immediate value to your organization.\nReal-World Applications This solution addresses documentation challenges across different departments:\nHR and People Teams: Employees constantly ask about benefits, PTO policies, and workplace guidelines. An AI bot can instantly answer \u0026ldquo;How many vacation days do I have?\u0026rdquo; or \u0026ldquo;What\u0026rsquo;s our parental leave policy?\u0026rdquo; by citing the exact paragraph from your handbook.\nEngineering Teams: Technical documentation grows exponentially with your codebase. When an engineer asks \u0026ldquo;How do I set up the development environment?\u0026rdquo; or \u0026ldquo;What\u0026rsquo;s our database migration process?\u0026rdquo;, the bot can provide step-by-step instructions from your wiki.\nProduct and Sales Teams: Sales representatives need quick access to product specifications, pricing details, and competitive positioning. A knowledge bot can answer \u0026ldquo;What are the enterprise tier limits?\u0026rdquo; during a client call without disrupting other team members.\nCustomer Support: Support teams juggle hundreds of internal processes. When an agent needs to know \u0026ldquo;What\u0026rsquo;s our escalation policy?\u0026rdquo; or \u0026ldquo;How do I process a refund?\u0026rdquo;, immediate answers improve customer response times.\nNew Employee Onboarding: The first weeks at a new job involve absorbing massive amounts of information. A knowledge bot gives new hires an accessible way to ask questions without feeling like they\u0026rsquo;re bothering colleagues.\nWhat We\u0026rsquo;re Building A Slack bot that:\nReceives questions from users in a channel or DM Uses AWS Bedrock Knowledge Base to search through your company documentation Generates accurate answers with citations to source documents Handles follow-up questions with conversation history The system uses Retrieval-Augmented Generation (RAG), combining the reasoning capabilities of large language models with retrieval from your own data sources—giving you the benefits of generative AI while keeping your data within your AWS account.\nUnderstanding Bedrock Knowledge Bases AWS Bedrock Knowledge Bases represents a significant advancement in enterprise knowledge management. Let\u0026rsquo;s explore how it works and why it\u0026rsquo;s superior to traditional search or direct LLM prompting.\nThe RAG Architecture Retrieval-Augmented Generation (RAG) addresses a fundamental limitation of LLMs: they have no knowledge of your internal documents. RAG works by:\nDocument Processing: Your documents are divided into chunks of an appropriate size for retrieval Vector Embedding: Each chunk is converted into a numerical vector representation using an embedding model Vector Storage: These embeddings are stored in a vector database (OpenSearch Serverless in Bedrock\u0026rsquo;s case) Semantic Search: When a question arrives, it\u0026rsquo;s converted to the same vector space and semantically similar chunks are retrieved Context Augmentation: Retrieved chunks are injected as context into the prompt sent to the LLM Answer Generation: The LLM generates an answer based on this context, citing the relevant sources This approach dramatically improves accuracy by giving the model direct access to your internal knowledge, while maintaining the reasoning capabilities of foundation models.\nSupported Document Types Bedrock Knowledge Bases supports a wide range of document formats:\nPDF files (text and scanned documents with OCR) Microsoft Office (Word, PowerPoint, Excel) Text and Markdown files HTML and web pages CSV and JSON data This versatility means you can ingest existing documentation without reformatting.\nSynchronization and Updates A key feature is automatic synchronization. When documents in your S3 bucket are updated, Bedrock Knowledge Bases can automatically detect these changes and update the vector store, ensuring your bot always has the latest information.\nSemantic vs. Keyword Search Traditional search systems match keywords, but Bedrock Knowledge Bases understands concepts. If someone asks about \u0026ldquo;time off,\u0026rdquo; it can retrieve documents about \u0026ldquo;vacation,\u0026rdquo; \u0026ldquo;PTO,\u0026rdquo; and \u0026ldquo;leave of absence\u0026rdquo; because it understands these concepts are related—even if they don\u0026rsquo;t share exact keywords.\nArchitecture Overview Here\u0026rsquo;s how the solution components work together:\nWhen a user asks a question in Slack, the message triggers a webhook to API Gateway. This request is processed by a Lambda function that maintains conversation context in DynamoDB and communicates with Bedrock. The Bedrock Agent uses the Knowledge Base to search your documentation, retrieves relevant information, and formulates a response that\u0026rsquo;s sent back to the user through Slack.\nEach component serves a specific purpose in this flow:\nSlack API handles the user interface, making the experience seamless within your existing communication platform. API Gateway provides a secure endpoint for Slack to send events to. Lambda orchestrates the process, maintaining conversation history and managing the interaction between Slack and Bedrock. Bedrock Agent uses Claude to interpret questions, retrieve information, and generate natural-sounding responses. Bedrock Knowledge Base indexes and searches your company documentation, finding the most relevant information for each question. DynamoDB stores conversation history so the bot can understand follow-up questions in context. This serverless architecture scales automatically with usage and requires minimal maintenance once deployed.\nPrerequisites Before starting, make sure you have:\nAWS Account with Bedrock access (you\u0026rsquo;ll need quota for Claude models) Slack workspace with permissions to create apps Company documentation organized in a folder structure Basic familiarity with AWS services and Terraform (or CloudFormation) Step 1: Create the Knowledge Base First, we\u0026rsquo;ll create a Knowledge Base to store and index your company documentation:\nUpload documents to an S3 bucket:\naws s3 mb s3://your-company-docs aws s3 sync ./docs s3://your-company-docs/ Organize your documents logically—folders like HR, Engineering, and Sales help the system understand document context.\nCreate the Knowledge Base in Bedrock:\nNavigate to AWS Bedrock in the console, select \u0026ldquo;Knowledge bases\u0026rdquo; → \u0026ldquo;Create knowledge base,\u0026rdquo; and follow the wizard:\nName: \u0026ldquo;CompanyDocs\u0026rdquo; Data source: Select your S3 bucket Vector store: \u0026ldquo;Create new Amazon OpenSearch Serverless vector store\u0026rdquo; Embedding model: \u0026ldquo;Titan Embeddings G1\u0026rdquo; (offers excellent performance for most use cases) Enable automatic synchronization to keep your knowledge base updated The initial data synchronization process will take several minutes depending on the volume of your documents. During this time, Bedrock is analyzing your documents, chunking them appropriately, and converting them into vector embeddings for semantic search.\nStep 2: Create the Bedrock Agent Now let\u0026rsquo;s create an agent that will use our Knowledge Base:\nIn the Bedrock console, go to \u0026ldquo;Agents\u0026rdquo; → \u0026ldquo;Create agent\u0026rdquo;\nName it \u0026ldquo;CompanyDocsAssistant\u0026rdquo; and select Claude 3.5 Sonnet for the foundation model\nIn the \u0026ldquo;Action groups\u0026rdquo; section, add a Knowledge Base action group and select the \u0026ldquo;CompanyDocs\u0026rdquo; knowledge base we created\nConfigure the agent\u0026rsquo;s instructions with detailed guidance:\nYou are a helpful assistant that answers questions about company documentation, policies, and procedures. When answering: 1. Be concise but thorough 2. Always cite sources by document name when you provide information 3. If you don\u0026#39;t know or can\u0026#39;t find relevant information, say so clearly 4. For follow-up questions, maintain context from previous exchanges 5. Format responses with appropriate Slack formatting (bullets, bold, etc.) where helpful 6. Present step-by-step procedures in numbered lists when applicable For the IAM role, create a new service role with the necessary permissions to access your Knowledge Base\nThe detailed instructions are crucial—they set the tone and behavior of your assistant, determining how it will respond to various types of questions.\nStep 3: Create the Lambda Function Next, create a Lambda function to handle Slack events and communicate with our Bedrock agent:\nimport json import os import boto3 import logging import urllib.request import time from boto3.dynamodb.conditions import Key # Initialize clients bedrock_agent_runtime = boto3.client(\u0026#39;bedrock-agent-runtime\u0026#39;) dynamodb = boto3.resource(\u0026#39;dynamodb\u0026#39;) conversation_table = dynamodb.Table(os.environ[\u0026#39;CONVERSATION_TABLE\u0026#39;]) logger = logging.getLogger() logger.setLevel(logging.INFO) def lambda_handler(event, context): # Parse the incoming event from Slack body = json.loads(event[\u0026#39;body\u0026#39;]) # Handle URL verification challenge if body.get(\u0026#39;type\u0026#39;) == \u0026#39;url_verification\u0026#39;: return {\u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;challenge\u0026#39;: body[\u0026#39;challenge\u0026#39;]})} # Process message events (app_mention or direct message) if body.get(\u0026#39;event\u0026#39;, {}).get(\u0026#39;type\u0026#39;) == \u0026#39;app_mention\u0026#39; or \\ (body.get(\u0026#39;event\u0026#39;, {}).get(\u0026#39;type\u0026#39;) == \u0026#39;message\u0026#39; and body.get(\u0026#39;event\u0026#39;, {}).get(\u0026#39;channel_type\u0026#39;) == \u0026#39;im\u0026#39;): event_data = body[\u0026#39;event\u0026#39;] user_id = event_data[\u0026#39;user\u0026#39;] channel_id = event_data[\u0026#39;channel\u0026#39;] text = event_data.get(\u0026#39;text\u0026#39;, \u0026#39;\u0026#39;).replace(f\u0026#34;\u0026lt;@{os.environ[\u0026#39;BOT_USER_ID\u0026#39;]}\u0026gt;\u0026#34;, \u0026#39;\u0026#39;).strip() # Get conversation history and invoke Bedrock agent conversation_id = f\u0026#34;{user_id}:{channel_id}\u0026#34; history = get_conversation_history(conversation_id) response = invoke_bedrock_agent(text, history, conversation_id) # Send response back to Slack send_slack_message(channel_id, response) return {\u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;ok\u0026#39;})} return {\u0026#39;statusCode\u0026#39;: 200, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;status\u0026#39;: \u0026#39;ignored\u0026#39;})} def invoke_bedrock_agent(question, history, conversation_id): try: # Format history for Bedrock and add the current question messages = format_conversation_history(history) messages.append({ \u0026#39;role\u0026#39;: \u0026#39;user\u0026#39;, \u0026#39;content\u0026#39;: [{\u0026#39;text\u0026#39;: question}] }) # Invoke the Bedrock agent response = bedrock_agent_runtime.invoke_agent( agentId=os.environ[\u0026#39;BEDROCK_AGENT_ID\u0026#39;], agentAliasId=os.environ[\u0026#39;BEDROCK_AGENT_ALIAS_ID\u0026#39;], sessionId=conversation_id, inputText=question, enableTrace=True ) # Extract and process the response completion = process_agent_response(response) # Store conversation in DynamoDB for history store_conversation_entry(conversation_id, \u0026#39;user\u0026#39;, question) store_conversation_entry(conversation_id, \u0026#39;assistant\u0026#39;, completion) return completion except Exception as e: logger.error(f\u0026#34;Error invoking Bedrock agent: {str(e)}\u0026#34;) return f\u0026#34;I\u0026#39;m having trouble answering that right now. Technical details: {str(e)}\u0026#34; # Additional helper functions for conversation history, messaging, etc. # (implementation details omitted for brevity) This Lambda function handles:\nReceiving events from Slack Maintaining conversation context Communicating with the Bedrock agent Sending responses back to the user The actual implementation includes additional helpers for conversation history management, message formatting, and error handling that we\u0026rsquo;ve omitted here for brevity.\nStep 4: Set up the DynamoDB Table You\u0026rsquo;ll need a DynamoDB table to track conversation history. In production, you\u0026rsquo;d define this in your infrastructure-as-code using Terraform or CloudFormation. The table needs:\nPartition key: conversation_id (String) Sort key: timestamp (Number) PAY_PER_REQUEST billing mode for cost efficiency This table enables the bot to understand follow-up questions by maintaining context from previous exchanges.\nStep 5: Create the Slack App Go to api.slack.com/apps and create a new app \u0026ldquo;From scratch\u0026rdquo; Under \u0026ldquo;OAuth \u0026amp; Permissions,\u0026rdquo; add these scopes: app_mentions:read chat:write im:history im:read Under \u0026ldquo;Event Subscriptions\u0026rdquo;: Enable Events and set the Request URL to your API Gateway endpoint Subscribe to bot events: app_mention and message.im Install the app to your workspace and copy the Bot User OAuth Token The Slack app configuration establishes the permissions and event subscriptions needed for the bot to receive messages and respond to users.\nStep 6: Deploy the API Gateway Create an API Gateway to receive events from Slack:\nCreate a new HTTP API with a POST route that integrates with your Lambda function Deploy the API and note the URL Update your Slack app\u0026rsquo;s Event Subscriptions URL with this endpoint Add environment variables to your Lambda function: SLACK_BOT_TOKEN: The OAuth token from your Slack app BOT_USER_ID: The user ID of your Slack bot BEDROCK_AGENT_ID: The ID of your Bedrock agent BEDROCK_AGENT_ALIAS_ID: The alias ID of your agent CONVERSATION_TABLE: Your DynamoDB table name Cost Optimization This solution is cost-effective, but there are a few considerations:\nBedrock API calls: ~$0.015 per 1,000 tokens with Claude Sonnet Knowledge Base storage: ~$0.023/GB for S3 + OpenSearch Serverless vector storage Lambda: Free tier likely covers most usage patterns DynamoDB: Pay-per-request pricing keeps costs low API Gateway: ~$1 per million requests For a team of 20 people asking 10 questions per day, expect costs around $30-50 per month. You can implement usage tracking to monitor and control costs as adoption grows.\nExtending the Solution Here are some ways to enhance this basic implementation:\nMulti-channel support: Monitor multiple Slack channels with channel-specific knowledge bases Document syncing: Set up automatic synchronization with your documentation systems Permissions: Implement access controls based on Slack user groups Analytics: Track common questions to identify gaps in your documentation Multi-model support: Use a simpler model for basic questions and Claude for complex ones Conversation summarization: Periodically summarize long conversations for better context management Conclusion With just a few AWS services, you\u0026rsquo;ve built an intelligent assistant that makes your company\u0026rsquo;s documentation accessible via Slack. No more hunting through SharePoint or Confluence—just ask the bot and get instant answers with citations to the source material.\nThe real power here is that your data remains within your AWS account, the system only has access to approved documents, and it continuously improves as you add more documentation. As AWS enhances Bedrock\u0026rsquo;s capabilities, your bot will automatically benefit from these improvements without any changes to your architecture.\nThis solution demonstrates how easily companies can now deploy practical AI applications using managed services. What used to require a specialized ML team and months of development can now be built in days using serverless components.\nWhat documentation would you connect to your knowledge bot first? Let me know on LinkedIn!\n","permalink":"https://lukelittle.com/posts/2026/02/building-a-company-knowledge-bot-slack--bedrock-knowledge-bases/","summary":"\u003cp\u003e\u0026ldquo;Where can I find our vacation policy?\u0026rdquo; \u0026ldquo;What\u0026rsquo;s the process for requesting new hardware?\u0026rdquo; \u0026ldquo;Can you explain our security guidelines?\u0026rdquo; These questions echo through company Slack channels daily, interrupting workflows and creating redundant work for team leads and HR staff. The same questions get asked repeatedly, and answers are buried in documentation that\u0026rsquo;s difficult to navigate.\u003c/p\u003e\n\u003cp\u003eIn this post, I\u0026rsquo;ll show you how to build a simple yet powerful Q\u0026amp;A bot for Slack that leverages your company\u0026rsquo;s documentation to provide accurate, contextual answers. The best part? It runs entirely on AWS managed services, minimizing operational overhead while delivering immediate value to your organization.\u003c/p\u003e","title":"Building a Company Knowledge Bot: Slack + Bedrock Knowledge Bases"},{"content":"I wrote this article for Ippon on February 10, 2026. The enthusiasm around projects like OpenClaw—an open-source framework enabling autonomous task execution across messaging platforms, file systems, and enterprise APIs—reveals a critical blind spot in enterprise technology governance.\nFor an entire week, I heard nothing but discussions of moltbots, OpenClaw, Clawd, and various implementations being created with these tools. But when I realized the permissions being exposed and how enterprises were approaching agent governance, I felt compelled to document this moment in AI history.\nThe Mental Model That\u0026rsquo;s Cracking The core of the issue is that technology executives are making permission and architecture decisions based on an outdated mental model of where AI lives in the enterprise. That understanding is molting. The old shell has cracked. The new one hasn\u0026rsquo;t hardened.\nFive years ago, AI lived in models. You sent a query, you got a prediction. The AI was a service you called when you needed analysis. It was passive, responsive, and contained within clear boundaries.\nThat mental model shaped how enterprises provisioned access. But this model is now too small. It has cracked. And most technology executives still provision access as if it were intact.\nThe Reality of Agent Operation AI agents no longer live in models. They live inside your operational environment. They run continuously, not episodically. They identify problems and pursue solutions autonomously. They compose capabilities across systems in ways you didn\u0026rsquo;t anticipate. They optimize for goals, not for following prescribed paths.\nAgent Operation Cycle:\nThe cycle begins when a user grants initial permissions to the agent. From there, the agent perceives the environment, including available systems, data, and APIs. It evaluates the current state against its objectives, then plans specific actions to achieve its goals. Once planned, it executes tools against enterprise systems and returns to perception with new information.\nThis creates a continuous loop of operation. The agent constantly explores what\u0026rsquo;s possible within its granted permissions, discovers novel ways to combine capabilities, optimizes relentlessly toward objectives, and persists until goals are achieved. Unlike human operators who might respect implied boundaries, agents operate to the full extent of their technical permissions.\nWhen you grant an agent permission to access systems, you\u0026rsquo;re not just giving it the ability to perform specific documented actions. You\u0026rsquo;re giving it the capability to explore what\u0026rsquo;s possible, discover paths through your systems, compose capabilities in unexpected ways, and persist until it achieves its objectives.\nThe Permission Gap The gap between intended permissions and actual capabilities is where enterprise failures occur. Not through compromise or malicious behavior, but through agents doing exactly what they were designed to do using permissions you willingly granted.\nWhen you grant email access for calendar integration, you\u0026rsquo;ve actually given permission to read all messages, not just calendar invites. The agent can extract any information from email content, map communication patterns and relationships, and leverage that data for any goal, not just scheduling.\nWith file system access supposedly for document generation, the agent can traverse all directories, not just document folders. It can discover credentials in config files, find and extract sensitive data, and move information between previously segregated systems.\nIf you provide API access for specific workflows, the agent will enumerate all endpoints, not just documented ones. It can compose novel workflows that bypass governance controls, create unauthorized integrations between systems, and persist until objectives are met, regardless of constraints or organizational boundaries.\nWhen you grant an agent \u0026ldquo;API access to the financial system,\u0026rdquo; here\u0026rsquo;s what you actually provisioned: continuous exploration of what that API can do; composition across boundaries that were previously respected; optimization for stated goals without recognizing unstated constraints; and persistence until objectives are met regardless of unexpected consequences.\nThis isn\u0026rsquo;t theoretical. These are the natural outcomes when you provision access for agents using mental models designed for passive services.\nA Moment in History We\u0026rsquo;re watching the world transform around us in ways we can\u0026rsquo;t really predict. Agentic AI is fundamentally different from what came before—not just better, but qualitatively different in how it operates.\nWhat\u0026rsquo;s happening now with OpenClaw and moltbots will be remembered as a pivotal transition point. One day we\u0026rsquo;ll recall this era the same way we look back on AltaVista or those curated lists of URLs people used to maintain to find websites in the early web.\nWe\u0026rsquo;re in that transitional period where the old shell has cracked but the new one hasn\u0026rsquo;t hardened. The decisions enterprises make during this molting period—about how they provision access, how they architect controls, and how they govern autonomous systems—will establish patterns that persist for years.\nMy full article on Ippon\u0026rsquo;s blog explores this governance gap in much greater detail, examining specific examples of permission mismatches, architectural controls that can help, and governance frameworks designed for the new reality of agent-based operations.\nFebruary 2026. The week of the moltbots. We were here. We saw it happen.\n","permalink":"https://lukelittle.com/posts/2026/02/the-week-of-the-moltbots/","summary":"\u003cp\u003eI wrote \u003ca href=\"https://blog.ippon.tech/openclaw-and-the-molting-of-enterprise-ai-governance\"\u003ethis article for Ippon\u003c/a\u003e on February 10, 2026. The enthusiasm around projects like \u003ca href=\"https://openclaw.ai/\"\u003eOpenClaw\u003c/a\u003e—an open-source framework enabling autonomous task execution across messaging platforms, file systems, and enterprise APIs—reveals a critical blind spot in enterprise technology governance.\u003c/p\u003e\n\u003cp\u003eFor an entire week, I heard nothing but discussions of moltbots, OpenClaw, Clawd, and various implementations being created with these tools. But when I realized the permissions being exposed and how enterprises were approaching agent governance, I felt compelled to document this moment in AI history.\u003c/p\u003e","title":"The Week of the Moltbots"},{"content":"Live Demonstration of FastMCP on AWS The February meetup of the Richmond AWS User Group featured a hands-on demonstration of FastMCP on AWS, exploring how modern agent frameworks can be deployed and operated in real cloud environments. Rather than focusing on theoretical concepts, the session provided attendees with practical insights into what it actually takes to run AI agent frameworks in production.\nBehind the Scenes: Unscripted AI Engineering What made this demonstration particularly authentic was that I hadn\u0026rsquo;t tested the solution beforehand. Armed with an impressively detailed prompt I\u0026rsquo;d crafted (available on GitHub), I wanted to make this a genuinely live experience—including all the potential hiccups and surprises that come with real AI development.\nAs part of my introduction, I surveyed attendees about their programming language preferences and AI tooling adoption. Interestingly, only about half were currently using agentic coding tools like Kiro or Claude Code in their workflows, highlighting the adoption curve many teams are still navigating.\nGoing Beyond the Theory Unlike high-level overviews that often gloss over implementation details, this talk walked through a real working example. I connected an AI agent to my personal record album collection hosted on a website and worked toward building a page where the model could answer specific questions about that collection. This was an evolution of the project I described in my January article on FastMCP and the Vinyl Collection Chatbot.\nThe demo highlighted some important realities of working with AI agents: even with carefully prepared prompts, much of the work involves orchestration, managing component spin-up times, and making adjustments throughout the development process.\nIn one of the evening\u0026rsquo;s more memorable moments, I deliberately let Kiro loop for about 45 minutes while attendees engaged in a spirited debate over whether it was stuck or still processing. This unplanned \u0026ldquo;teachable moment\u0026rdquo; perfectly illustrated the sometimes opaque nature of these systems and the patience required when working with them. Eventually, we were rewarded with a working solution, but the journey there—with all its uncertainty—was perhaps more valuable than a polished, pre-prepared demo would have been.\nThis practical perspective gave attendees a clearer understanding of the day-to-day challenges and solutions when working with these technologies.\nThe Nuances of Prompt Design One of the most valuable discussions during the session centered around prompt design. We explored how changing topics mid-prompt can quietly degrade accuracy—not because the system is fundamentally flawed, but because the model is constantly making statistical predictions about what comes next.\nWatching these effects play out in real-time reinforced an important lesson: small decisions in prompt design can have compounding effects as system complexity increases. This practical demonstration helped attendees understand the nuanced relationship between prompt structure and system performance.\nIncremental Progress and Real Results By the end of the session, the agent was successfully answering targeted questions about the album collection. While this represented a relatively simple use case, it effectively demonstrated that effective AI systems are typically built incrementally, with careful attention to:\nContext management Structural considerations Patience during the iteration process For professionals working in cloud, security, and related fields, the session underscored that understanding how these systems behave under real-world conditions is just as critical as familiarity with the underlying theory.\nAccess to Demo Resources The code demonstrated during the session is available on GitHub for those who want to explore the implementation details further: https://github.com/lukelittle/rawsug-fastmcp-demo\nCommunity Engagement Thanks to everyone who attended, asked thoughtful questions, and stayed after for discussions about architecture, AI agents, and practical AWS implementations. Special appreciation goes to Lucas Ward and Devin Veasna for their contributions, and to the entire RAWSUG community for maintaining an environment that\u0026rsquo;s practical, curious, and enjoyable.\nAs we continue exploring the intersection of AI and cloud technology, these community gatherings provide invaluable opportunities to learn from each other\u0026rsquo;s experiences and collectively advance our understanding of these powerful tools.\nLooking Forward The Richmond AWS User Group remains committed to delivering practical, hands-on sessions that go beyond marketing slides to show real implementation details. Stay tuned for upcoming meetups that will continue to explore cutting-edge technologies with a focus on practical application.\n","permalink":"https://lukelittle.com/posts/2026/02/richmond-aws-user-group-fastmcp-demo-on-aws/","summary":"\u003ch2 id=\"live-demonstration-of-fastmcp-on-aws\"\u003eLive Demonstration of FastMCP on AWS\u003c/h2\u003e\n\u003cp\u003eThe February meetup of the Richmond AWS User Group featured a hands-on demonstration of FastMCP on AWS, exploring how modern agent frameworks can be deployed and operated in real cloud environments. Rather than focusing on theoretical concepts, the session provided attendees with practical insights into what it actually takes to run AI agent frameworks in production.\u003c/p\u003e\n\u003ch2 id=\"behind-the-scenes-unscripted-ai-engineering\"\u003eBehind the Scenes: Unscripted AI Engineering\u003c/h2\u003e\n\u003cp\u003eWhat made this demonstration particularly authentic was that I hadn\u0026rsquo;t tested the solution beforehand. Armed with an impressively detailed prompt I\u0026rsquo;d crafted (\u003ca href=\"https://github.com/lukelittle/rawsug-fastmcp-demo/blob/main/prompt.txt\"\u003eavailable on GitHub\u003c/a\u003e), I wanted to make this a genuinely live experience—including all the potential hiccups and surprises that come with real AI development.\u003c/p\u003e","title":"Richmond AWS User Group: FastMCP Demo on AWS"},{"content":"The Real Problem: Production, Not Prototypes Everyone can demo generative AI. Almost no one can run it safely in production.\nEnterprises in finance, healthcare, and the public sector aren\u0026rsquo;t blocked by technology capabilities—they\u0026rsquo;re blocked by governance requirements that today\u0026rsquo;s AI implementations rarely satisfy.\nThese organizations face three critical blockers:\nData leakage risk: Sensitive information, from PII to trade secrets, flowing through public model APIs Lack of auditability: No reliable record of prompts, responses, or who accessed what information Unclear ownership: Ambiguous rights over prompt engineering IP, training data, and generated outputs AWS customers don\u0026rsquo;t want AI that behaves like a chatbot toy. They need AI that behaves like enterprise infrastructure: secured, monitored, audited, governed, and compliant with their existing security posture.\nDesign Goals for Enterprise-Ready GenAI When designing generative AI systems for regulated environments, your architecture must satisfy these non-negotiable requirements:\nNo public internet exposure for sensitive data No training on customer data without explicit permission Full audit trail of all prompts and responses IAM-first access control integrated with enterprise identity Serverless and scalable by default This checklist maps directly to AWS Well-Architected Framework principles, particularly in security and operational excellence.\nReference Architecture Overview Here\u0026rsquo;s a reference architecture that meets these requirements using AWS services:\nWalking the Architecture: Building for Security and Scale 1️⃣ Edge \u0026amp; Entry: CloudFront + API Gateway The edge layer serves as your first line of defense:\nGlobal edge protection through CloudFront Request validation and throttling via API Gateway Clear API contract for AI access WAF rules to block suspicious patterns This approach frames AI as just another AWS workload, not an exception to your security rules. Your existing infrastructure and compliance controls extend naturally to your AI services.\n2️⃣ Prompt Handling: Lambda The Prompt Handler Lambda is where policy meets AI:\nSanitizes inputs to prevent prompt injection Injects system prompts to enforce guardrails Enforces token limits (cost control) Attaches request metadata (user ID, application, purpose) This layer ensures all model interactions are appropriately structured and traced. Every prompt includes context about who sent it, why, and what constraints apply.\n3️⃣ Private Model Access: Bedrock via VPC Endpoint The model interaction layer guarantees data privacy:\nNo public internet egress No customer-managed model hosting No fine-tuning on customer prompts VPC integration with existing security controls The model is consumed like a managed AWS service—not an external API. This distinction is critical for security teams evaluating AI adoption.\n4️⃣ Response Filtering: Post-processing Lambda The response handler implements safety guardrails:\nContent moderation (PII, offensive content) Output validation against schema Optional redaction of sensitive information Confidence scoring and hallucination detection This layer acknowledges and mitigates hallucination risk without fear-mongering, providing mechanisms to validate and filter model outputs.\n5️⃣ Audit \u0026amp; Evidence: DynamoDB / S3 The audit layer addresses compliance requirements:\nPersistent storage of prompt hashes Model ID and version tracking Timestamped responses Immutable audit logs This creates a defensible evidentiary trail that satisfies governance requirements for regulated industries.\nWhy This Works for Regulated Industries This architecture succeeds where most AI implementations fail because it addresses the key requirements that matter to enterprise stakeholders:\nSecurity: IAM-based access control, VPC endpoints, and private networking eliminate public exposure risks. The system operates entirely within your security perimeter, following the principle of \u0026ldquo;default deny\u0026rdquo; with explicit allow policies.\nCompliance: Complete prompt/response traceability enables regulatory reporting and satisfies audit requirements. You can demonstrate who used the system, when, how, and what results they received—critical for SOC2, HIPAA, and FedRAMP.\nCost control: Serverless scaling plus token limits provide predictable, manageable costs. Unlike self-hosted options, you\u0026rsquo;re not paying for idle infrastructure, and unlike public APIs, you have fine-grained control over usage patterns.\nOperational clarity: The system is observable, debuggable, and auditable using the same tools you already use for the rest of your AWS infrastructure. There\u0026rsquo;s no AI-specific monitoring to implement.\nThis approach works particularly well for financial services (handling sensitive financial data), healthcare (maintaining PHI compliance), and public sector (satisfying FedRAMP requirements)—precisely the industries with the most to gain from AI and the most stringent security requirements.\nImplementation Considerations When implementing this pattern in production, several practical considerations emerge:\nIAM roles and boundaries: Create specific IAM roles for each component with least-privilege access. The prompt handler needs Bedrock access but not S3 write access; the response filter needs DynamoDB write access but not Bedrock APIs. Use service control policies (SCPs) to enforce guardrails.\nVPC design: Depending on your existing network topology, you may need to adjust the VPC design. For large enterprises with transit gateways, consider routing AI traffic through dedicated VPCs with specific security monitoring.\nCost management: Monitor token usage carefully. Implement token quotas at the API Gateway layer and consider using smaller context window models for initial responses, reserving larger context models for specific use cases.\nScaling characteristics: Lambda\u0026rsquo;s concurrency model handles traffic spikes well, but Bedrock has model-specific quotas and SLAs. Request quota increases proactively if you anticipate high volume. Consider implementing queue-based architectures for asynchronous workloads.\nCross-account patterns: For large organizations, implement a hub-and-spoke model where a central AI governance account hosts the Bedrock endpoint, with workload accounts accessing it through cross-account roles. This centralizes auditing while enabling distributed usage.\nConclusion The gap between AI demos and AI in production isn\u0026rsquo;t primarily a technical gap—it\u0026rsquo;s a governance gap.\nThis reference architecture bridges that gap by treating generative AI as enterprise infrastructure rather than a standalone tool. It integrates with existing security controls, creates auditability, and provides the governance hooks necessary for regulated environments.\nThe result? AI that can safely navigate the journey from prompt to production, enabling organizations to capture AI\u0026rsquo;s business value without compromising on security and compliance requirements.\nIn regulated environments, the future of AI isn\u0026rsquo;t about building fancy demos—it\u0026rsquo;s about building trust. By architecting generative AI systems that behave like proper enterprise infrastructure—secured, monitored, audited, and governed—we allow organizations to focus on business value rather than security firefighting.\nRemember that this is a reference architecture, not a one-size-fits-all solution. Your specific implementation should be tailored to your compliance requirements, existing infrastructure, and risk profile. But the principles outlined here—isolation, auditability, IAM-first access, and metadata enrichment—remain universal best practices for any enterprise AI deployment.\n","permalink":"https://lukelittle.com/posts/2026/02/from-prompt-to-production-designing-safe-generative-ai-on-aws-for-regulated-environments/","summary":"\u003ch2 id=\"the-real-problem-production-not-prototypes\"\u003eThe Real Problem: Production, Not Prototypes\u003c/h2\u003e\n\u003cp\u003eEveryone can demo generative AI. Almost no one can run it safely in production.\u003c/p\u003e\n\u003cp\u003eEnterprises in finance, healthcare, and the public sector aren\u0026rsquo;t blocked by technology capabilities—they\u0026rsquo;re blocked by governance requirements that today\u0026rsquo;s AI implementations rarely satisfy.\u003c/p\u003e\n\u003cp\u003eThese organizations face three critical blockers:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eData leakage risk\u003c/strong\u003e: Sensitive information, from PII to trade secrets, flowing through public model APIs\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eLack of auditability\u003c/strong\u003e: No reliable record of prompts, responses, or who accessed what information\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eUnclear ownership\u003c/strong\u003e: Ambiguous rights over prompt engineering IP, training data, and generated outputs\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eAWS customers don\u0026rsquo;t want AI that behaves like a chatbot toy. They need AI that behaves like enterprise infrastructure: secured, monitored, audited, governed, and compliant with their existing security posture.\u003c/p\u003e","title":"From Prompt to Production: Designing Safe Generative AI on AWS for Regulated Environments"},{"content":"What is the Model Context Protocol? The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Think of it as a universal adapter that lets any AI agent talk to any tool or data source without custom integration code.\nAnthropic announced MCP in November 2024 and donated it to the Linux Foundation\u0026rsquo;s Agentic AI Foundation about a year later, in December 2025. The adoption has been swift: OpenAI integrated it into ChatGPT, Google DeepMind uses it for Gemini agents, AWS built AgentCore around it, and development tools like Zed, Sourcegraph, Replit, and Codeium all support it. In just a few months, the community has built thousands of MCP servers. The protocol has become the de-facto standard for agent-to-tool communication.\nIf you\u0026rsquo;ve read my previous articles on AWS DevOps Agent, Security Agent, and Kiro, you\u0026rsquo;ve seen what these frontier agents do. This article explains how they actually work—the protocol layer that makes integration possible. More importantly, it shows how you can extend these agents to work with your proprietary systems.\nThe N×M integration problem Here\u0026rsquo;s why MCP matters: You\u0026rsquo;ve built an amazing AI agent. It reasons brilliantly, writes elegant code, debugs complex issues. But it can\u0026rsquo;t access your company\u0026rsquo;s data. Your customer records are in Salesforce. Your code is in GitHub. Your metrics are in Datadog. Your tickets are in Jira.\nThe traditional approach is building custom integrations. One for each pairing. Want your agent to read Salesforce? Build a Salesforce connector. Want it to access GitHub? Build a GitHub connector. Want it to work with both? Build both connectors. Want to switch LLM providers? Rebuild everything.\nThis is the N×M integration problem: N agents times M data sources equals N×M custom integrations. As your ecosystem grows, the complexity becomes unmanageable. Three agents talking to four services means twelve custom connectors. Ten agents and fifty systems means five hundred integrations to build and maintain.\nEach connector requires understanding the target system\u0026rsquo;s API, building authentication flows, handling rate limits and retries, writing serialization and deserialization logic, maintaining the connector as APIs change, and rebuilding for each new agent or LLM provider. This doesn\u0026rsquo;t scale.\nWith MCP, you build the Salesforce MCP server once. After that, DevOps Agent, Security Agent, Kiro, Claude Desktop—any MCP client—can access Salesforce. No custom integration needed. This is why every major AI company adopted MCP within months of its announcement. It solves an existential scaling problem.\nHow MCP works: The architecture MCP follows a client-server architecture using JSON-RPC 2.0 for communication. The protocol exposes three core primitives that servers can implement.\nResources are like RESTful GET endpoints. They load information into LLM context—files, database records, API responses. When you ask an agent to retrieve a customer\u0026rsquo;s support ticket history, it\u0026rsquo;s accessing a resource.\nTools are like RESTful POST endpoints. They execute code and produce side effects—creating Jira tickets, deploying code, sending Slack messages. When you tell an agent to create a high-priority ticket for a customer, it\u0026rsquo;s using a tool.\nPrompts are reusable templates that guide interactions with structured patterns. Think \u0026ldquo;Analyze this code for security vulnerabilities\u0026rdquo; or \u0026ldquo;Generate release notes from commits.\u0026rdquo; They standardize how agents approach common tasks.\nThe communication flow Here\u0026rsquo;s how an agent actually communicates with an MCP server:\nsequenceDiagram participant Agent as MCP Client (Agent) participant Server as MCP Server (Tool) Agent-\u0026gt;\u0026gt;Server: 1. Initialize Connection Agent-\u0026gt;\u0026gt;Server: 2. List Available Tools Server-\u0026gt;\u0026gt;Agent: 3. Tool Definitions (JSON Schema) Note over Agent: Agent decides\u0026lt;br/\u0026gt;which tool to call Agent-\u0026gt;\u0026gt;Server: 4. Call Tool (with parameters) Note over Server: Validates params\u0026lt;br/\u0026gt;Executes logic\u0026lt;br/\u0026gt;Queries external API Server-\u0026gt;\u0026gt;Agent: 5. Tool Result Note over Agent: Agent uses result\u0026lt;br/\u0026gt;in response The agent initializes a connection, requests the list of available tools, receives their JSON Schema definitions, decides which tool to call based on the user\u0026rsquo;s request, sends the tool invocation with parameters, waits for the server to validate inputs and execute the logic, receives the result, and uses it to generate the final response. It\u0026rsquo;s a clean request-response pattern.\nTransport mechanisms MCP supports two primary transport methods. Standard Input/Output (stdio) is designed for local integration—the server runs as a subprocess and communicates via stdin/stdout. This is what Claude Desktop uses to run local MCP servers. It\u0026rsquo;s simple, fast, and synchronous, but only works locally.\nServer-Sent Events (SSE) is designed for remote integration. It uses HTTP-based streaming where the server pushes updates to the client. This is what you\u0026rsquo;d use for cloud-hosted MCP servers and enterprise integrations. It\u0026rsquo;s network-capable and supports real-time updates, but requires HTTP server infrastructure.\nThe protocol layer MCP messages are structured as JSON-RPC 2.0 calls. Here\u0026rsquo;s what a tool call looks like: Request from client to server:\n{ \u0026#34;jsonrpc\u0026#34;: \u0026#34;2.0\u0026#34;, \u0026#34;id\u0026#34;: 1, \u0026#34;method\u0026#34;: \u0026#34;tools/call\u0026#34;, \u0026#34;params\u0026#34;: { \u0026#34;name\u0026#34;: \u0026#34;get_customer_info\u0026#34;, \u0026#34;arguments\u0026#34;: { \u0026#34;customer_id\u0026#34;: \u0026#34;cust_12345\u0026#34; } } } Response from server to client:\n{ \u0026#34;jsonrpc\u0026#34;: \u0026#34;2.0\u0026#34;, \u0026#34;id\u0026#34;: 1, \u0026#34;result\u0026#34;: { \u0026#34;content\u0026#34;: [ { \u0026#34;type\u0026#34;: \u0026#34;text\u0026#34;, \u0026#34;text\u0026#34;: \u0026#34;{\\\u0026#34;name\\\u0026#34;: \\\u0026#34;Acme Corp\\\u0026#34;, \\\u0026#34;tier\\\u0026#34;: \\\u0026#34;Enterprise\\\u0026#34;, \\\u0026#34;health\\\u0026#34;: \\\u0026#34;green\\\u0026#34;}\u0026#34; } ] } } The beauty of this: you rarely write this JSON by hand. MCP SDKs for Python, TypeScript, Java, C#, and Kotlin abstract it away completely. Which brings us to FastMCP.\nBuilding with FastMCP Building MCP servers from scratch means handling JSON-RPC protocol details, managing connections, writing serialization boilerplate. It\u0026rsquo;s tedious work that distracts from the actual business logic you want to implement.\nFastMCP is the solution—a decorator-based Python framework that makes building MCP servers feel like writing FastAPI applications. Think of it as FastAPI for AI agents with the same elegant decorator patterns, Flask for tool integration with minimal boilerplate, or Express.js for MCP with simple and intuitive design.\nHere\u0026rsquo;s a complete, working MCP server:\nfrom fastmcp import FastMCP # Initialize server mcp = FastMCP(\u0026#34;My First Server\u0026#34;) # Define a tool @mcp.tool() def add_numbers(a: int, b: int) -\u0026gt; int: \u0026#34;\u0026#34;\u0026#34;Add two numbers together\u0026#34;\u0026#34;\u0026#34; return a + b # Run it if __name__ == \u0026#34;__main__\u0026#34;: mcp.run() That\u0026rsquo;s it. You now have an MCP server that exposes a tool, automatically generates JSON Schema from type hints, handles serialization and deserialization, provides error handling, supports both stdio and SSE transports, and works with any MCP client. No boilerplate. No protocol details. Just your logic.\nA real example: GitHub integration Let\u0026rsquo;s build something practical—an MCP server that lets agents interact with GitHub. This demonstrates how FastMCP handles real-world integrations with external APIs, authentication, and error handling.\nfrom fastmcp import FastMCP from github import Github import os # Initialize server mcp = FastMCP( name=\u0026#34;GitHub Integration\u0026#34;, instructions=\u0026#34;Use this server to interact with GitHub repositories\u0026#34; ) # Initialize GitHub client github_token = os.getenv(\u0026#34;GITHUB_TOKEN\u0026#34;) gh = Github(github_token) @mcp.tool() def get_repository_info(owner: str, repo: str) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34; Get information about a GitHub repository. Args: owner: Repository owner (username or organization) repo: Repository name Returns: Repository details including stars, forks, issues \u0026#34;\u0026#34;\u0026#34; repository = gh.get_repo(f\u0026#34;{owner}/{repo}\u0026#34;) return { \u0026#34;name\u0026#34;: repository.name, \u0026#34;description\u0026#34;: repository.description, \u0026#34;stars\u0026#34;: repository.stargazers_count, \u0026#34;forks\u0026#34;: repository.forks_count, \u0026#34;open_issues\u0026#34;: repository.open_issues_count, \u0026#34;language\u0026#34;: repository.language, \u0026#34;created_at\u0026#34;: repository.created_at.isoformat(), \u0026#34;updated_at\u0026#34;: repository.updated_at.isoformat() } @mcp.tool() def create_issue( owner: str, repo: str, title: str, body: str, labels: list[str] = None ) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34; Create a new issue in a GitHub repository. Args: owner: Repository owner repo: Repository name title: Issue title body: Issue description labels: Optional list of labels to apply Returns: Created issue details \u0026#34;\u0026#34;\u0026#34; repository = gh.get_repo(f\u0026#34;{owner}/{repo}\u0026#34;) issue = repository.create_issue( title=title, body=body, labels=labels or [] ) return { \u0026#34;number\u0026#34;: issue.number, \u0026#34;url\u0026#34;: issue.html_url, \u0026#34;state\u0026#34;: issue.state, \u0026#34;created_at\u0026#34;: issue.created_at.isoformat() } @mcp.tool() def list_pull_requests( owner: str, repo: str, state: str = \u0026#34;open\u0026#34; ) -\u0026gt; list[dict]: \u0026#34;\u0026#34;\u0026#34; List pull requests in a repository. Args: owner: Repository owner repo: Repository name state: PR state (\u0026#39;open\u0026#39;, \u0026#39;closed\u0026#39;, \u0026#39;all\u0026#39;) Returns: List of pull requests \u0026#34;\u0026#34;\u0026#34; repository = gh.get_repo(f\u0026#34;{owner}/{repo}\u0026#34;) pulls = repository.get_pulls(state=state) return [ { \u0026#34;number\u0026#34;: pr.number, \u0026#34;title\u0026#34;: pr.title, \u0026#34;author\u0026#34;: pr.user.login, \u0026#34;state\u0026#34;: pr.state, \u0026#34;created_at\u0026#34;: pr.created_at.isoformat(), \u0026#34;url\u0026#34;: pr.html_url } for pr in pulls[:10] # Limit to 10 for brevity ] @mcp.resource(\u0026#34;repo://{owner}/{repo}/README\u0026#34;) def get_readme(owner: str, repo: str) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34; Get the README content for a repository. This is exposed as a resource (not a tool) because it\u0026#39;s primarily for loading context into the LLM. \u0026#34;\u0026#34;\u0026#34; repository = gh.get_repo(f\u0026#34;{owner}/{repo}\u0026#34;) readme = repository.get_readme() return readme.decoded_content.decode(\u0026#39;utf-8\u0026#39;) if __name__ == \u0026#34;__main__\u0026#34;: # Run with stdio (for local Claude Desktop) mcp.run(transport=\u0026#34;stdio\u0026#34;) # Or run with SSE (for remote access) # mcp.run(transport=\u0026#34;sse\u0026#34;, port=8000) FastMCP handled type validation—ensuring owner and repo are strings—error handling so GitHub API failures return proper MCP errors, schema generation from docstrings and type hints, authentication flow using the GitHub token from environment variables, and transport abstraction so the same code works with stdio or SSE.\nYou focused on your business logic—what the tool actually does—and clear documentation where docstrings become tool descriptions. The framework handles everything else.\nAdvanced FastMCP features FastMCP isn\u0026rsquo;t just decorators—it\u0026rsquo;s a full-featured framework with capabilities that become important as your integrations grow more sophisticated.\nContext injection lets you access MCP context and capabilities within your tools. This is useful for long-running operations where you want to send progress updates to the client:\nfrom fastmcp import FastMCP, Context from mcp.server.session import ServerSession mcp = FastMCP(\u0026#34;Progress Example\u0026#34;) @mcp.tool() async def long_running_task( task_name: str, ctx: Context[ServerSession, None], steps: int = 5 ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Execute a task with progress updates.\u0026#34;\u0026#34;\u0026#34; # Send progress notifications to client await ctx.info(f\u0026#34;Starting: {task_name}\u0026#34;) for i in range(steps): progress = (i + 1) / steps * 100 await ctx.progress(progress, f\u0026#34;Step {i+1}/{steps}\u0026#34;) # Do actual work here await asyncio.sleep(1) await ctx.info(f\u0026#34;Completed: {task_name}\u0026#34;) return f\u0026#34;Task {task_name} completed successfully\u0026#34; Sampling lets your server request LLM completions. This is powerful for agentic workflows where the server orchestrates LLM calls without needing API keys:\n@mcp.tool() async def generate_commit_message( diff: str, ctx: Context[ServerSession, None] ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Generate a commit message from a git diff.\u0026#34;\u0026#34;\u0026#34; # Ask the client\u0026#39;s LLM to generate the message result = await ctx.session.create_message( messages=[{ \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: f\u0026#34;Generate a concise commit message for this diff:\\n\\n{diff}\u0026#34; }], max_tokens=100 ) return result.content[0].text Elicitation lets servers request additional information mid-operation. This is useful when the tool needs clarification from the user:\nfrom mcp.types import ElicitRequestTextParams @mcp.tool() async def deploy_to_environment( service: str, ctx: Context[ServerSession, None] ) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Deploy a service to an environment.\u0026#34;\u0026#34;\u0026#34; # Ask user which environment result = await ctx.elicit( ElicitRequestTextParams( mode=\u0026#34;text\u0026#34;, message=\u0026#34;Which environment? (staging/production)\u0026#34;, placeholder=\u0026#34;staging\u0026#34; ) ) environment = result.value # Validate input if environment not in [\u0026#34;staging\u0026#34;, \u0026#34;production\u0026#34;]: raise ValueError(\u0026#34;Environment must be \u0026#39;staging\u0026#39; or \u0026#39;production\u0026#39;\u0026#34;) # Proceed with deployment return f\u0026#34;Deploying {service} to {environment}...\u0026#34; Filesystem roots define security boundaries for which directories agents can access. FastMCP enforces that paths must be within configured roots that the client specifies when connecting:\n@mcp.tool() def read_file(path: str, ctx: Context[ServerSession, None]) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Read a file from the allowed workspace.\u0026#34;\u0026#34;\u0026#34; # FastMCP enforces that path must be within configured roots # Client specifies roots when connecting: # roots=[\u0026#34;file:///safe/workspace\u0026#34;] with open(path, \u0026#39;r\u0026#39;) as f: return f.read() Case study: Vinyl collection chatbot To see how easy FastMCP makes building real-world integrations, I built a serverless chatbot that answers questions about my vinyl collection. The data comes from a Discogs export—a CSV file containing my complete record collection that I store in S3.\nAsk it \u0026ldquo;What Grimes records do I have?\u0026rdquo; and it queries the CSV, parses the Discogs data, and returns results. Ask it \u0026ldquo;What is vinyl?\u0026rdquo; and it just answers from general knowledge. The bot intelligently decides when to use tools and when to rely on its training data.\nHere\u0026rsquo;s the complete implementation of the MCP server that queries the Discogs collection data:\nfrom fastmcp import FastMCP import boto3 import csv from io import StringIO # Initialize FastMCP server mcp = FastMCP(\u0026#34;vinyl-collection-server\u0026#34;) # S3 client for accessing Discogs export s3_client = boto3.client(\u0026#39;s3\u0026#39;) DATA_BUCKET = \u0026#34;my-vinyl-collection\u0026#34; @mcp.tool() def query_vinyl_collection(query_type: str, search_term: str, limit: int = 10) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34; Query Luke\u0026#39;s vinyl record collection from Discogs export data. Args: query_type: One of: artist, label, year, title, all search_term: What to search for limit: Max results (default 10) \u0026#34;\u0026#34;\u0026#34; # Download Discogs CSV export from S3 response = s3_client.get_object(Bucket=DATA_BUCKET, Key=\u0026#39;discogs.csv\u0026#39;) csv_content = response[\u0026#39;Body\u0026#39;].read().decode(\u0026#39;utf-8\u0026#39;) # Parse CSV records = [] csv_reader = csv.DictReader(StringIO(csv_content)) for row in csv_reader: records.append(row) # Filter based on query type matches = [] for record in records: if query_type == \u0026#34;artist\u0026#34; and search_term.lower() in record[\u0026#39;Artist\u0026#39;].lower(): matches.append(record) elif query_type == \u0026#34;label\u0026#34; and search_term.lower() in record[\u0026#39;Label\u0026#39;].lower(): matches.append(record) elif query_type == \u0026#34;year\u0026#34; and search_term in record[\u0026#39;Released\u0026#39;]: matches.append(record) elif query_type == \u0026#34;title\u0026#34; and search_term.lower() in record[\u0026#39;Title\u0026#39;].lower(): matches.append(record) elif query_type == \u0026#34;all\u0026#34;: if (search_term.lower() in record[\u0026#39;Artist\u0026#39;].lower() or search_term.lower() in record[\u0026#39;Title\u0026#39;].lower()): matches.append(record) # Format results results = [] for record in matches[:limit]: results.append( f\u0026#34;{record[\u0026#39;Artist\u0026#39;]} - {record[\u0026#39;Title\u0026#39;]} \u0026#34; f\u0026#34;({record[\u0026#39;Label\u0026#39;]}, {record[\u0026#39;Released\u0026#39;]})\u0026#34; ) return \u0026#34;\\n\u0026#34;.join(results) if results else \u0026#34;No matches found\u0026#34; if __name__ == \u0026#34;__main__\u0026#34;: mcp.run(transport=\u0026#34;sse\u0026#34;, port=8080) That\u0026rsquo;s it. No JSON schema writing. No protocol handling. No tool registration boilerplate. FastMCP generates the schema from type hints and docstrings, handles the Model Context Protocol communication, converts everything to Bedrock\u0026rsquo;s format, and manages tool execution. You write the business logic for parsing Discogs data. FastMCP handles everything else.\nThe whole system—frontend, Lambda, Bedrock integration, S3 data access—deployed in under two minutes. This isn\u0026rsquo;t a toy example. It\u0026rsquo;s running in production, costs about fifteen dollars a month, and demonstrates why FastMCP is becoming the standard way to build tools for AI agents.\nThis is the full loop: user asks a question, Bedrock decides whether to use a tool, FastMCP executes it against the Discogs export data, and the result is returned—all serverless.\nHow AWS frontier agents use MCP Now let\u0026rsquo;s connect this to the agents I\u0026rsquo;ve covered in previous articles. These patterns show how MCP enables coordination between specialized agents.\nDevOps Agent scenario: A Lambda function starts timing out in production. DevOps Agent detects the anomaly through CloudWatch MCP Server where it reads metrics, logs, and traces. It queries the GitHub MCP Server to analyze recent code changes that might have introduced the issue. It creates an incident through the PagerDuty MCP Server and notifies the on-call engineer. Finally, it posts a summary to the incidents channel via the Slack MCP Server.\nWithout MCP, AWS would need custom connectors for every observability tool, ticketing system, and chat platform—an integration nightmare that doesn\u0026rsquo;t scale. With MCP, DevOps Agent uses a standard MCP client and any MCP-compatible tool works instantly.\nSecurity Agent scenario: During a penetration test, Security Agent finds a SQL injection vulnerability. It reads the vulnerable code through the GitHub MCP Server, creates a security ticket via the Jira MCP Server, requests a code fix from Kiro through MCP, and once Kiro generates the fix, it creates a pull request back through GitHub\u0026rsquo;s MCP Server and notifies the security team via Slack\u0026rsquo;s MCP Server.\nThe agents coordinate through MCP without custom integration code. Security Agent discovers, Kiro fixes, DevOps Agent validates deployment—all using the same protocol.\nKiro scenario: A developer asks Kiro to \u0026ldquo;Add authentication to the user API endpoint.\u0026rdquo; Kiro reads the existing code through the Filesystem MCP Server, checks authentication patterns in other services via the GitHub MCP Server, queries your company\u0026rsquo;s authentication library documentation through an Internal Auth MCP Server, writes the updated code back through the Filesystem MCP Server, and creates a pull request via GitHub\u0026rsquo;s MCP Server.\nKiro doesn\u0026rsquo;t just generate code—it actively researches your codebase and standards through MCP servers, understanding context before making changes.\nBuilding enterprise MCP servers: Real patterns Here are three production-grade patterns that demonstrate how FastMCP handles enterprise integrations.\nSalesforce MCP Server shows integration with a major CRM system:\nfrom fastmcp import FastMCP from simple_salesforce import Salesforce import os mcp = FastMCP(\u0026#34;Salesforce CRM\u0026#34;) sf = Salesforce( username=os.getenv(\u0026#34;SF_USERNAME\u0026#34;), password=os.getenv(\u0026#34;SF_PASSWORD\u0026#34;), security_token=os.getenv(\u0026#34;SF_SECURITY_TOKEN\u0026#34;) ) @mcp.tool() def get_account(account_id: str) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34;Retrieve Salesforce account details.\u0026#34;\u0026#34;\u0026#34; return sf.Account.get(account_id) @mcp.tool() def create_case( account_id: str, subject: str, description: str, priority: str = \u0026#34;Medium\u0026#34; ) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34;Create a support case in Salesforce.\u0026#34;\u0026#34;\u0026#34; return sf.Case.create({ \u0026#39;AccountId\u0026#39;: account_id, \u0026#39;Subject\u0026#39;: subject, \u0026#39;Description\u0026#39;: description, \u0026#39;Priority\u0026#39;: priority }) @mcp.tool() def search_accounts(query: str, limit: int = 10) -\u0026gt; list[dict]: \u0026#34;\u0026#34;\u0026#34;Search Salesforce accounts by name.\u0026#34;\u0026#34;\u0026#34; soql = f\u0026#34;SELECT Id, Name, Industry, AnnualRevenue FROM Account WHERE Name LIKE \u0026#39;%{query}%\u0026#39; LIMIT {limit}\u0026#34; return sf.query(soql)[\u0026#39;records\u0026#39;] With this server deployed, any agent can now access Salesforce—DevOps Agent, Security Agent, Kiro, or any future agent you build. No custom integration needed per agent. Build once, use everywhere.\nInternal Wiki MCP Server demonstrates knowledge base integration:\nfrom fastmcp import FastMCP from elasticsearch import Elasticsearch import os mcp = FastMCP(\u0026#34;Company Wiki\u0026#34;) es = Elasticsearch( os.getenv(\u0026#34;ELASTICSEARCH_URL\u0026#34;), api_key=os.getenv(\u0026#34;ELASTICSEARCH_API_KEY\u0026#34;) ) @mcp.resource(\u0026#34;wiki://{article_id}\u0026#34;) def get_wiki_article(article_id: str) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Get wiki article by ID.\u0026#34;\u0026#34;\u0026#34; result = es.get(index=\u0026#34;wiki\u0026#34;, id=article_id) return result[\u0026#39;_source\u0026#39;][\u0026#39;content\u0026#39;] @mcp.tool() def search_wiki( query: str, limit: int = 5 ) -\u0026gt; list[dict]: \u0026#34;\u0026#34;\u0026#34; Search company wiki for relevant articles. Returns articles ranked by relevance. \u0026#34;\u0026#34;\u0026#34; response = es.search( index=\u0026#34;wiki\u0026#34;, body={ \u0026#34;query\u0026#34;: { \u0026#34;multi_match\u0026#34;: { \u0026#34;query\u0026#34;: query, \u0026#34;fields\u0026#34;: [\u0026#34;title^2\u0026#34;, \u0026#34;content\u0026#34;, \u0026#34;tags\u0026#34;] } }, \u0026#34;size\u0026#34;: limit } ) return [ { \u0026#34;id\u0026#34;: hit[\u0026#39;_id\u0026#39;], \u0026#34;title\u0026#34;: hit[\u0026#39;_source\u0026#39;][\u0026#39;title\u0026#39;], \u0026#34;excerpt\u0026#34;: hit[\u0026#39;_source\u0026#39;][\u0026#39;content\u0026#39;][:200] + \u0026#34;...\u0026#34;, \u0026#34;score\u0026#34;: hit[\u0026#39;_score\u0026#39;], \u0026#34;url\u0026#34;: f\u0026#34;https://wiki.company.com/articles/{hit[\u0026#39;_id\u0026#39;]}\u0026#34; } for hit in response[\u0026#39;hits\u0026#39;][\u0026#39;hits\u0026#39;] ] This lets agents access institutional knowledge—Security Agent can read security policies, DevOps Agent can consult runbooks, and Kiro can reference coding standards, all from your internal wiki.\nMulti-Tool Orchestration MCP Server aggregates data from multiple sources:\nfrom fastmcp import FastMCP from datadog_api_client import ApiClient, Configuration from datadog_api_client.v1.api.metrics_api import MetricsApi mcp = FastMCP(\u0026#34;Operations Dashboard\u0026#34;) # Initialize multiple clients datadog_config = Configuration() datadog_config.api_key[\u0026#39;apiKeyAuth\u0026#39;] = os.getenv(\u0026#34;DD_API_KEY\u0026#34;) datadog_config.api_key[\u0026#39;appKeyAuth\u0026#39;] = os.getenv(\u0026#34;DD_APP_KEY\u0026#34;) @mcp.tool() async def get_service_health(service_name: str) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34; Get comprehensive health status for a service. Aggregates data from multiple monitoring systems. \u0026#34;\u0026#34;\u0026#34; # Get Datadog metrics with ApiClient(datadog_config) as api_client: api_instance = MetricsApi(api_client) metrics = api_instance.query_metrics( _from=int(time.time()) - 3600, to=int(time.time()), query=f\u0026#34;avg:system.cpu.user{{service:{service_name}}}\u0026#34; ) # Get PagerDuty incidents (if integrated) # Get deployment status from CI/CD # Aggregate everything return { \u0026#34;service\u0026#34;: service_name, \u0026#34;status\u0026#34;: \u0026#34;healthy\u0026#34;, # or \u0026#34;degraded\u0026#34;, \u0026#34;down\u0026#34; \u0026#34;cpu_usage\u0026#34;: metrics[\u0026#39;series\u0026#39;][0][\u0026#39;pointlist\u0026#39;][-1][1], \u0026#34;open_incidents\u0026#34;: 0, \u0026#34;last_deployment\u0026#34;: \u0026#34;2025-01-25T10:30:00Z\u0026#34;, \u0026#34;error_rate\u0026#34;: 0.02 # 2% } This gives agents a unified view across disparate monitoring tools—one call returns comprehensive health status aggregated from Datadog, PagerDuty, and your CI/CD system.\nDeploying MCP servers: Production patterns Claude Desktop is the easiest way to test MCP servers during development. Create your MCP server using the patterns shown above, then configure Claude Desktop to use it. Edit the configuration file at ~/Library/Application Support/Claude/claude_desktop_config.json on Mac or %APPDATA%\\Claude\\claude_desktop_config.json on Windows:\n{ \u0026#34;mcpServers\u0026#34;: { \u0026#34;github\u0026#34;: { \u0026#34;command\u0026#34;: \u0026#34;python\u0026#34;, \u0026#34;args\u0026#34;: [\u0026#34;/path/to/github_server.py\u0026#34;], \u0026#34;env\u0026#34;: { \u0026#34;GITHUB_TOKEN\u0026#34;: \u0026#34;your_token_here\u0026#34; } }, \u0026#34;salesforce\u0026#34;: { \u0026#34;command\u0026#34;: \u0026#34;python\u0026#34;, \u0026#34;args\u0026#34;: [\u0026#34;/path/to/salesforce_server.py\u0026#34;], \u0026#34;env\u0026#34;: { \u0026#34;SF_USERNAME\u0026#34;: \u0026#34;your_username\u0026#34;, \u0026#34;SF_PASSWORD\u0026#34;: \u0026#34;your_password\u0026#34; } } } } Restart Claude Desktop and your agents will have access to these tools.\nFor production deployments, use remote MCP servers with SSE transport:\n# server.py from fastmcp import FastMCP mcp = FastMCP(\u0026#34;Production GitHub Server\u0026#34;) # ... (your tools here) if __name__ == \u0026#34;__main__\u0026#34;: # Run with SSE on custom port mcp.run(transport=\u0026#34;sse\u0026#34;, port=8080) You can deploy this on AWS Lambda with a Lambda function URL, ECS or Fargate for containerized auto-scaling workloads, traditional EC2 instances, or AWS App Runner for fully managed hosting.\nFor AWS Bedrock AgentCore integration, configure the client like this:\nfrom mcp import ClientSession, SseServerParameters from mcp.client.sse import sse_client server_params = SseServerParameters( url=\u0026#34;https://mcp.company.com/github\u0026#34;, headers={\u0026#34;Authorization\u0026#34;: \u0026#34;Bearer YOUR_TOKEN\u0026#34;} ) async with sse_client(server_params) as (read, write): async with ClientSession(read, write) as session: await session.initialize() tools = await session.list_tools() # Use tools... AWS Bedrock AgentCore and MCP integration This is where everything comes together. AgentCore natively supports MCP through its Gateway service.\ngraph TB subgraph AgentCore[\u0026#34;AWS Bedrock AgentCore\u0026#34;] Runtime[AgentCore Runtime\u0026lt;br/\u0026gt;LangChain/LangGraph/Custom] Gateway[AgentCore Gateway\u0026lt;br/\u0026gt;Tool Discovery\u0026lt;br/\u0026gt;MCP Client\u0026lt;br/\u0026gt;Policy Enforcement\u0026lt;br/\u0026gt;Authentication] Runtime --\u0026gt; Gateway end Gateway --\u0026gt; GH[GitHub\u0026lt;br/\u0026gt;MCP Server] Gateway --\u0026gt; SF[Salesforce\u0026lt;br/\u0026gt;MCP Server] Gateway --\u0026gt; Custom[Your Custom\u0026lt;br/\u0026gt;MCP Server] Integrating your FastMCP server with AgentCore follows three steps:\n# Step 1: Build your MCP server with FastMCP from fastmcp import FastMCP mcp = FastMCP(\u0026#34;Custom Enterprise Tool\u0026#34;) @mcp.tool() def query_internal_database(sql: str) -\u0026gt; list[dict]: \u0026#34;\u0026#34;\u0026#34;Query internal PostgreSQL database.\u0026#34;\u0026#34;\u0026#34; # Your implementation pass # Deploy to AWS (Lambda, ECS, etc.) mcp.run(transport=\u0026#34;sse\u0026#34;, port=8080) # Step 2: Register with AgentCore Gateway # In AgentCore Console or via SDK: import boto3 agentcore = boto3.client(\u0026#39;bedrock-agentcore\u0026#39;) agentcore.register_mcp_server( name=\u0026#34;internal-database\u0026#34;, url=\u0026#34;https://internal-mcp.company.com\u0026#34;, authentication={ \u0026#39;type\u0026#39;: \u0026#39;bearer\u0026#39;, \u0026#39;tokenSecret\u0026#39;: \u0026#39;arn:aws:secretsmanager:...\u0026#39; } ) # Step 3: Use in your agent from langchain.agents import create_react_agent from bedrock_agentcore import BedrockAgentCoreApp app = BedrockAgentCoreApp() @app.entrypoint def my_agent(request): # Your agent automatically has access to all # MCP servers registered in AgentCore Gateway agent = create_react_agent(model, tools) return agent.invoke(request) AgentCore Gateway gives you automatic tool discovery where it lists all registered MCP servers, policy enforcement so you can define which agents can use which tools, authentication handling through AgentCore Identity for OAuth and API keys, observability with every MCP call logged to CloudWatch, and scaling managed entirely by AgentCore infrastructure.\nSecurity and best practices Input validation is critical. Always validate and sanitize inputs before execution:\n@mcp.tool() def query_database(sql: str) -\u0026gt; list[dict]: \u0026#34;\u0026#34;\u0026#34;Execute a SQL query (SELECT only).\u0026#34;\u0026#34;\u0026#34; # ✅ GOOD: Validate before execution dangerous_keywords = [\u0026#39;DROP\u0026#39;, \u0026#39;DELETE\u0026#39;, \u0026#39;UPDATE\u0026#39;, \u0026#39;INSERT\u0026#39;, \u0026#39;ALTER\u0026#39;] if any(keyword in sql.upper() for keyword in dangerous_keywords): raise ValueError(\u0026#34;Only SELECT queries allowed\u0026#34;) # Additional validation if \u0026#39;;\u0026#39; in sql: raise ValueError(\u0026#34;Multiple statements not allowed\u0026#34;) return execute_query(sql) Apply the principle of least privilege by granting minimal permissions. Instead of using full admin access tokens, use read-only or scoped tokens:\n# BAD: Full admin access github_token = os.getenv(\u0026#34;GITHUB_ADMIN_TOKEN\u0026#34;) # GOOD: Read-only or scoped token github_token = os.getenv(\u0026#34;GITHUB_READONLY_TOKEN\u0026#34;) Rate limiting protects your backend systems from abuse:\nfrom functools import lru_cache from time import time @lru_cache(maxsize=128) def rate_limit_key(user_id: str) -\u0026gt; int: return int(time() / 60) # 1-minute windows @mcp.tool() def expensive_operation(user_id: str, query: str) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34;Rate-limited operation.\u0026#34;\u0026#34;\u0026#34; # Simple rate limiting (use Redis in production) key = f\u0026#34;{user_id}:{rate_limit_key(user_id)}\u0026#34; if request_count(key) \u0026gt; 10: raise Exception(\u0026#34;Rate limit exceeded: 10 requests per minute\u0026#34;) increment_count(key) return perform_operation(query) Error handling should provide helpful errors without leaking internal details:\nfrom mcp.types import ToolError @mcp.tool() def get_sensitive_data(resource_id: str) -\u0026gt; dict: \u0026#34;\u0026#34;\u0026#34;Access protected resource.\u0026#34;\u0026#34;\u0026#34; try: return fetch_resource(resource_id) except PermissionError: # GOOD: Generic error message raise ToolError(\u0026#34;Access denied: insufficient permissions\u0026#34;) except Exception as e: # BAD: Would leak internal details # raise ToolError(f\u0026#34;Database error: {str(e)}\u0026#34;) # GOOD: Log internally, return generic error logger.error(f\u0026#34;Internal error: {e}\u0026#34;, exc_info=True) raise ToolError(\u0026#34;An internal error occurred\u0026#34;) The future of MCP Based on current trajectory and industry patterns, MCP is evolving rapidly. In the near term—the next six to twelve months—expect broader ecosystem adoption with every major AI platform supporting MCP (OpenAI, Google, Microsoft, AWS), IDE integration in VS Code, JetBrains, and Zed, and browser extensions for Arc and Chrome. Enhanced security will standardize OAuth 2.0 flows, introduce fine-grained permission models, and establish audit logging specifications. Performance optimizations will add caching strategies, batch operations, and streaming support for large results.\nThe medium term—one to two years out—will bring enterprise features like an MCP server marketplace similar to AWS Marketplace, a certified and verified server registry, and SLA guarantees with monitoring. Advanced capabilities will enable multi-step workflows that chain tools across servers, state management for sessions and transactions, and push notifications where servers can alert clients. Developer experience improvements will include visual server builders for low-code MCP development, testing frameworks like pytest for MCP, and observability SDKs with OpenTelemetry integration.\nLong term—two to five years—MCP becomes the HTTP of agentic AI. Every enterprise system will provide an MCP interface, legacy systems will get wrapped in MCP adapters, and universal agent communication becomes standard. Economic models will emerge with pay-per-call MCP servers, SaaS subscriptions for premium MCP servers, and even agent-to-agent commerce. The platform shift means developers stop building custom integrations entirely, \u0026ldquo;MCP-first\u0026rdquo; becomes the default architecture, and agents compose tools like developers compose libraries today.\nNext steps If you\u0026rsquo;re building agents, start by learning FastMCP basics in about thirty minutes—install it with pip, build the hello world server, and test with Claude Desktop. Then build an MCP server for your most-used tool in a couple hours—GitHub, Jira, Salesforce, or your internal API—following the patterns in this article and deploying locally to test thoroughly. Finally, integrate with your agents in about an hour by adding it to your Claude Desktop config or integrating with AgentCore Gateway, then verify agents can discover and use the tools.\nIf you\u0026rsquo;re at an enterprise, inventory your tool landscape over a day—list all systems agents need to access, identify which have existing MCP servers, and plan which you need to build custom. Build two to three pilot MCP servers in a week, starting with high-value, low-risk systems, using FastMCP for rapid development, and deploying to staging to test with real agents. Establish governance over two weeks by defining which agents can use which tools, implementing policy enforcement through AgentCore Policy, and setting up observability and monitoring. Then scale gradually over time, adding MCP servers incrementally, monitoring usage and costs, and iterating based on agent behavior.\nWhy this matters Here\u0026rsquo;s the reality: The value of AI agents isn\u0026rsquo;t in the agents themselves—it\u0026rsquo;s in what they can access. An agent that can reason brilliantly but can\u0026rsquo;t touch your systems is a toy. An agent that can access your systems through MCP is a tool. A fleet of specialized agents, coordinating through MCP, is a transformation.\nDevOps Agent is powerful. Security Agent is powerful. Kiro is powerful. But they\u0026rsquo;re only as powerful as the tools you give them through MCP.\nThat\u0026rsquo;s why understanding MCP isn\u0026rsquo;t optional—it\u0026rsquo;s foundational. It\u0026rsquo;s the difference between agents that demo well versus agents that actually work, prototypes in staging versus production deployments, vendor lock-in versus true interoperability.\nFastMCP makes building these integrations trivial. The protocol is stable. The ecosystem is growing rapidly. The question isn\u0026rsquo;t \u0026ldquo;Should we adopt MCP?\u0026rdquo; It\u0026rsquo;s \u0026ldquo;Which systems do we connect first?\u0026rdquo;\nResources For further exploration, start with the MCP Official Specification at modelcontextprotocol.io, the FastMCP GitHub repository at github.com/jlowin/fastmcp, and FastMCP Documentation at gofastmcp.com. The MCP Python SDK is available at github.com/modelcontextprotocol/python-sdk. AWS AgentCore Gateway documentation can be found in the AWS Bedrock AgentCore Documentation, and the MCP Server Registry is at github.com/modelcontextprotocol/servers.\nIf you\u0026rsquo;re building agents or integrating AI into enterprise systems, MCP is infrastructure you need to understand. Not as a nice-to-have, but as a prerequisite. Start with FastMCP. Build a server for your most-used tool. See how quickly you can give agents superpowers. The agents are here. The protocol is stable. The ecosystem is exploding. Time to plug in.\n","permalink":"https://lukelittle.com/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/","summary":"\u003ch2 id=\"what-is-the-model-context-protocol\"\u003eWhat is the Model Context Protocol?\u003c/h2\u003e\n\u003cp\u003eThe Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Think of it as a universal adapter that lets any AI agent talk to any tool or data source without custom integration code.\u003c/p\u003e\n\u003cp\u003eAnthropic announced MCP in November 2024 and donated it to the Linux Foundation\u0026rsquo;s Agentic AI Foundation about a year later, in December 2025. The adoption has been swift: OpenAI integrated it into ChatGPT, Google DeepMind uses it for Gemini agents, AWS built AgentCore around it, and development tools like Zed, Sourcegraph, Replit, and Codeium all support it. In just a few months, the community has built thousands of MCP servers. The protocol has become the de-facto standard for agent-to-tool communication.\u003c/p\u003e","title":"FastMCP and the Vinyl Collection Chatbot: Serverless Agentic AI in Action"},{"content":"Here\u0026rsquo;s what should make every security leader uncomfortable: organizations routinely deploy vulnerable code to production to meet delivery deadlines.\nNot because they don\u0026rsquo;t care about security. Because security can\u0026rsquo;t keep up.\nOver 60% of organizations update their web applications weekly or more frequently. Nearly 75% test those applications for security monthly or less. The math doesn\u0026rsquo;t work. The gap between development velocity and security validation grows wider every sprint.\nAt re:Invent 2025, AWS CEO Matt Garman announced AWS Security Agent—not as another security scanning tool to add to the pile, but as a fundamentally different approach to the problem.\nSecurity Agent is a frontier agent that operates autonomously throughout the development lifecycle, conducting design reviews, analyzing code, and executing penetration tests on-demand, matching the pace of modern development instead of being its bottleneck.\nYou can watch the AWS Security Agent announcement here: https://www.youtube.com/watch?v=oMY0tUDEhtY\nThis is AWS\u0026rsquo;s bet that security doesn\u0026rsquo;t scale through more manual reviews—it scales through intelligent automation that understands your applications, your standards, and your threats.\nWhat makes an agent \u0026ldquo;frontier-class\u0026rdquo; AWS uses the term \u0026ldquo;frontier agent\u0026rdquo; to mean something specific. It\u0026rsquo;s not just GPT-4 with security tools.\n1. Autonomous goal-directed behavior\nTraditional security: \u0026ldquo;Run this SAST scan and give me findings.\u0026rdquo;\nFrontier agent: \u0026ldquo;Validate this application meets our security requirements\u0026rdquo; → agent figures out how\nYou give it an objective, it decomposes the problem, forms hypotheses, collects evidence, and executes—without asking you for step-by-step guidance.\n2. Multi-agent coordination\nSecurity Agent doesn\u0026rsquo;t work alone. When conducting a penetration test, it spawns specialized sub-agents—one for authentication bypass, another for authorization flaws, a third for injection attacks. These agents work concurrently, investigating multiple attack vectors simultaneously and coordinating across findings.\n3. Long-running and context-aware\nHere\u0026rsquo;s the paradigm shift: Security Agent maintains persistent understanding of your applications.\nTraditional security tools forget everything between scans. Security Agent learns:\nYour organizational security requirements Your application architecture and data flows Your code patterns and common vulnerabilities Your historical findings and remediation approaches When testing your API, it doesn\u0026rsquo;t just throw generic payloads. It understands your authentication mechanism, maps your business logic, and targets application-specific vulnerabilities.\nHow it actually works Security Agent operates across three phases of the development lifecycle, each with different capabilities:\nPhase 1: Design Security Review\nBefore code exists, upload design documents, architecture diagrams, and threat models. Security Agent analyzes against:\nAWS security best practices Your organization\u0026rsquo;s security requirements Common architectural vulnerabilities Threat modeling patterns Output: Security risk analysis with specific remediation guidance, in minutes instead of days.\nPhase 2: Code Security Review\nDuring development, GitHub integration provides automated security feedback on every pull request. Security Agent validates:\nOrganizational security requirements (approved libraries, logging standards, data policies) OWASP Top 10 vulnerabilities Code-level security anti-patterns Compliance with your defined security standards Developers get immediate feedback in their workflow—no context switching required.\nPhase 3: On-Demand Penetration Testing\nWhenever needed—pre-deployment, post-change, or on a schedule—Security Agent conducts comprehensive penetration testing.\nUnlike traditional scanners, it:\nBuilds understanding from your source code and architecture Creates customized attack plans based on your specific stack Executes multi-step attack chains (not just single-payload scans) Tests business logic vulnerabilities Validates findings to eliminate false positives Generates pull requests with remediation code The penetration testing loop When you trigger a pentest, here\u0026rsquo;s what happens:\ngraph TB Input[Target URLs + Code + Docs] --\u0026gt; Context[Build Application Understanding] Context --\u0026gt; Discovery[Discover Attack Surface\u0026lt;br/\u0026gt;Map endpoints, APIs, flows] Discovery --\u0026gt; Planning[Planning Agent\u0026lt;br/\u0026gt;Create customized attack plan] Planning --\u0026gt; Testing[Specialized Testing Agents] subgraph Testing[\u0026#34; \u0026#34;] Auth[Auth Bypass] Authz[Authorization] Inject[Injection] Logic[Business Logic] end Testing --\u0026gt; Validate[Validator Agent\u0026lt;br/\u0026gt;Eliminate false positives] Validate --\u0026gt; Remediate[Remediation Agent\u0026lt;br/\u0026gt;Generate code fixes] Remediate --\u0026gt; PR[Pull Request with Fix] The clever part is the context-aware testing. Security Agent analyzes your source code to understand:\nWhich endpoints are public vs. internal What authentication patterns you use How data flows through your system What your actual threat model looks like When it sees JWT tokens, it doesn\u0026rsquo;t just test for SQL injection—it focuses on JWT-specific attacks like algorithm confusion, token tampering, and replay attacks.\nAgent Spaces and organizational requirements Everything starts with an Agent Space—the workspace where Security Agent operates and the security boundary for what it can access.\nYou might structure Agent Spaces as:\nPer-application: One space for your customer portal, another for your admin dashboard Per-team: One space per development team managing their portfolio Per-environment: Separate spaces for staging vs. production testing The powerful part: you define your organization\u0026rsquo;s security standards once, centrally:\nAuthentication Requirements: - OAuth 2.0 for all API endpoints - JWT tokens with 15-minute expiration - MFA required for admin functions Data Protection Requirements: - PII encrypted at rest (KMS) - TLS 1.3 for data in transit - No credit card data in logs Logging Requirements: - Correlation IDs on all requests - Auth failures logged with IP - No PII in application logs These requirements automatically apply to all Agent Spaces, enforced during both design reviews and code reviews. Consistent enforcement across the organization—no more \u0026ldquo;this team follows standards, that team doesn\u0026rsquo;t.\u0026rdquo;\nDeploying it (practical walkthrough) AWS provides console-based setup. Here\u0026rsquo;s the flow:\nStep 1: Create Agent Space\nAWS Console → Security Agent → Create Agent Space - Name: \u0026#34;production-security\u0026#34; - Agent role: Auto-created Step 2: Define Security Requirements\nSecurity Requirements → Add requirements: - Authentication standards - Authorization patterns - Data protection policies - Logging requirements - Compliance frameworks Step 3: Configure Penetration Testing\nAgent Space → Enable penetration testing - Add target domains (verify ownership) - Configure CloudWatch logging - Set up VPC access (for private apps) - Add credentials to Secrets Manager Step 4: Integrate with GitHub\nInstall AWS Security Agent GitHub App - Authorize repository access - Enable code review for Agent Space - Auto-review triggered on PRs Step 5: Execute Penetration Test\nSecurity Agent Web App → Create pentest - Target: https://staging.app.example.com - Authentication: From Secrets Manager - Attach: Source code + design docs - Enable automatic remediation PRs Watch as Security Agent:\nDiscovers attack surface Executes targeted scenarios Validates findings Creates PRs with fixes Testing it AWS provides real test scenarios. Run these before connecting production:\nTest 1: API authentication bypass\nDeploy an API with intentionally weak JWT validation:\n# Weak JWT verification def verify_token(token): # Missing signature validation payload = jwt.decode(token, verify=False) return payload[\u0026#39;user_id\u0026#39;] Trigger Security Agent pentest, watch it:\nIdentify JWT usage Test algorithm confusion Attempt signature bypass Generate exploit proof Create PR with proper validation Test 2: SQL injection in query\nDeploy an endpoint with SQL injection:\n# Vulnerable query def get_user(user_id): query = f\u0026#34;SELECT * FROM users WHERE id = {user_id}\u0026#34; return db.execute(query) Security Agent should:\nDetect SQL construction Test injection vectors Confirm exploitability Recommend parameterized queries What\u0026rsquo;s not ready yet (the honest limitations) 1. us-east-1 only Security Agent is currently only available in us-east-1.\nIf you have data residency requirements (GDPR, finance, healthcare), this is a blocker. Applications in other regions must be accessible from us-east-1.\nMitigation: Test staging/dev environments. AWS will expand regions post-GA.\n2. GitHub-only code review Currently only GitHub is supported for automated code review.\nWhat\u0026rsquo;s missing:\nGitLab Bitbucket AWS CodeCommit Azure DevOps Workaround: You can still use penetration testing by providing code via S3.\n3. Not a replacement for professional pentesting Security Agent is powerful but not guaranteed to discover all vulnerabilities. It\u0026rsquo;s best used for continuous testing at development velocity.\nAWS\u0026rsquo;s position: \u0026ldquo;AWS Security Agent is not a professional penetration testing service, and we encourage users to integrate AWS Security Agent into their security review workflow.\u0026rdquo;\nThe right mental model: Security Agent is your continuous validation layer. Professional pentesters are your comprehensive pre-launch audit.\n4. False positive management While Validator Agents significantly reduce false positives, AI-powered testing will never be 100% perfect.\nWhat AWS does:\nOnly reports high/medium confidence findings Hides unverified findings by default Provides reproducible exploit paths What you should do:\nReview findings with security expertise Validate critical findings independently Use CloudWatch logs to understand agent reasoning 5. Learning curve Security Agent builds topology understanding over time. Early investigations might be less accurate.\nMitigation:\nRun test assessments to let it learn Tag resources consistently Document dependencies explicitly Should you actually use this? Use it if:\nYour development velocity is outpacing security capacity You deploy weekly but test security monthly You knowingly ship vulnerable code to meet deadlines You want to scale security across your entire portfolio You\u0026rsquo;re AWS-native (tight integration benefits) You\u0026rsquo;re comfortable with preview-phase tech Wait if:\nYou need multi-region support now You use source control other than GitHub (for code review) Your security process already keeps pace with development You need production SLAs (preview = no SLAs) You require deterministic pricing Key insight: Security Agent amplifies good practices and exposes bad ones. If your infrastructure is poorly tagged, deployments aren\u0026rsquo;t tracked, and requirements are scattered, it\u0026rsquo;ll struggle. But if you have solid foundations, it can be transformative.\nMy honest take This is the future of application security. Not because AI replaces security engineers, but because it handles undifferentiated heavy lifting.\nThe traditional model: Security is a gate. Development builds features, security reviews them, findings go back to development, repeat until deadlines force compromise.\nThe agentic model: Security is embedded. AI agents continuously validate security throughout development, provide real-time guidance, and scale security expertise to match development velocity.\nSecurity Agent doesn\u0026rsquo;t eliminate the need for security expertise—it amplifies what your security team can accomplish. One senior security engineer using Security Agent can cover more applications than a team of five without it.\nThe question isn\u0026rsquo;t whether agentic security is coming—it\u0026rsquo;s here. The question is whether you\u0026rsquo;re ready when GA drops.\nIf you\u0026rsquo;re experimenting with this or have questions, reach out on LinkedIn. The technology is moving fast, and we\u0026rsquo;re all figuring it out together.\nResources:\nAWS Security Agent User Guide AWS Security Agent Product Page AWS Frontier Agents Overview ","permalink":"https://lukelittle.com/posts/2026/01/your-ai-security-engineer-inside-aws-security-agent/","summary":"\u003cp\u003eHere\u0026rsquo;s what should make every security leader uncomfortable: organizations routinely deploy vulnerable code to production to meet delivery deadlines.\u003c/p\u003e\n\u003cp\u003eNot because they don\u0026rsquo;t care about security. Because security can\u0026rsquo;t keep up.\u003c/p\u003e\n\u003cp\u003eOver 60% of organizations update their web applications weekly or more frequently. Nearly 75% test those applications for security monthly or less. The math doesn\u0026rsquo;t work. The gap between development velocity and security validation grows wider every sprint.\u003c/p\u003e\n\u003cp\u003eAt re:Invent 2025, AWS CEO Matt Garman announced \u003cstrong\u003eAWS Security Agent\u003c/strong\u003e—not as another security scanning tool to add to the pile, but as a fundamentally different approach to the problem.\u003c/p\u003e","title":"Your AI Security Engineer: Inside AWS Security Agent"},{"content":"Our serverless survey application is a great example of a modern cloud native application. It\u0026rsquo;s fast, scalable, and cost-effective. But it\u0026rsquo;s missing one critical feature: user authentication. In this post, we\u0026rsquo;ll walk through how to add robust, secure authentication using AWS Cognito.\nWhy Add Authentication? Right now, anyone can vote, and anyone can reset the entire survey. In a real-world application, we need to control access. Authentication allows us to:\nPrevent abuse: Stop users from voting multiple times. Secure administrative functions: Only allow authorized users to reset the survey. Personalize the user experience: (Future enhancement) Show users their past votes. Exploring the Options When adding authentication to a serverless app on AWS, there are a few common patterns:\nAPI Keys: The simplest approach. We could generate an API key and require it for certain API endpoints.\nPros: Easy to implement. Cons: Not true user authentication. Keys can be shared or leaked. Doesn\u0026rsquo;t scale for managing individual users. Lambda Authorizers (Custom Authorizers): We can write a custom Lambda function that is triggered by API Gateway before the main handler. This function would be responsible for validating a token (e.g., a session token you manage yourself).\nPros: Full control over the authentication logic. Cons: You have to build and manage the entire user lifecycle yourself (sign-up, sign-in, password reset, etc.). This is a lot of work and easy to get wrong. AWS Cognito: A fully managed user identity and authentication service. Cognito handles all the heavy lifting of user management, including sign-up, sign-in, password recovery, and multi-factor authentication (MFA). It integrates seamlessly with API Gateway.\nPros: Secure, scalable, and feature-rich. Offloads the undifferentiated heavy lifting of authentication. Cons: Can seem complex at first due to the number of features. For our application, AWS Cognito is the clear winner. It provides the best balance of security, features, and ease of integration.\nUnderstanding the Current Architecture Before we add authentication, let\u0026rsquo;s visualize how our serverless application currently works:\nRight now, anyone can call our API endpoints. There\u0026rsquo;s no way to verify who\u0026rsquo;s making the request or prevent abuse.\nThe Plan: Integrating Cognito We\u0026rsquo;ll use a Cognito User Pool to manage our users. Think of it as a user database with built-in authentication logic. Here\u0026rsquo;s the high-level plan:\nInfrastructure (Terraform):\nCreate a Cognito User Pool to store and manage users. Create a Cognito User Pool Client, which allows our frontend application to interact with the User Pool. Add an authorizer to our API Gateway to protect our endpoints. Frontend (HTML/JavaScript):\nCreate a login page (login.html). Add sign-up and sign-in forms. Use the Amazon Cognito Identity SDK for JavaScript to communicate with Cognito. On successful sign-in, store the JWT (JSON Web Token) from Cognito in local storage. Include the JWT in the Authorization header of all subsequent API requests. Add a \u0026ldquo;Logout\u0026rdquo; button. Backend (Lambda):\nUpdate our Lambda functions to expect and validate the JWT from Cognito. API Gateway will do most of the validation for us. We can use the claims inside the JWT to identify the user. For instance, we can use the user\u0026rsquo;s sub (subject) claim instead of the sessionId to track votes. Here\u0026rsquo;s what the authenticated architecture will look like:\nLet\u0026rsquo;s Get Building! Step 1: Beefing up our Terraform First, we need to add the Cognito resources to our terraform/main.tf file.\n# --- Cognito User Pool --- resource \u0026#34;aws_cognito_user_pool\u0026#34; \u0026#34;survey_user_pool\u0026#34; { name = \u0026#34;SurveyUserPool\u0026#34; auto_verified_attributes = [\u0026#34;email\u0026#34;] } resource \u0026#34;aws_cognito_user_pool_client\u0026#34; \u0026#34;survey_user_pool_client\u0026#34; { name = \u0026#34;SurveyUserPoolClient\u0026#34; user_pool_id = aws_cognito_user_pool.survey_user_pool.id generate_secret = false # This is a public client explicit_auth_flows = [\u0026#34;ALLOW_USER_PASSWORD_AUTH\u0026#34;, \u0026#34;ALLOW_REFRESH_TOKEN_AUTH\u0026#34;] } We also need to update our API Gateway to use a Cognito authorizer.\n# --- API Gateway --- resource \u0026#34;aws_apigatewayv2_api\u0026#34; \u0026#34;survey_api\u0026#34; { name = \u0026#34;CrackingTheCloudSurveyAPI\u0026#34; protocol_type = \u0026#34;HTTP\u0026#34; cors_configuration { allow_origins = [\u0026#34;*\u0026#34;] allow_methods = [\u0026#34;POST\u0026#34;, \u0026#34;GET\u0026#34;, \u0026#34;OPTIONS\u0026#34;] allow_headers = [\u0026#34;*\u0026#34;] } } resource \u0026#34;aws_apigatewayv2_authorizer\u0026#34; \u0026#34;cognito_authorizer\u0026#34; { api_id = aws_apigatewayv2_api.survey_api.id authorizer_type = \u0026#34;JWT\u0026#34; identity_sources = [\u0026#34;$request.header.Authorization\u0026#34;] name = \u0026#34;CognitoAuthorizer\u0026#34; jwt_configuration { audience = [aws_cognito_user_pool_client.survey_user_pool_client.id] issuer = \u0026#34;https://${aws_cognito_user_pool.survey_user_pool.endpoint}\u0026#34; } } # Update the routes to use the authorizer resource \u0026#34;aws_apigatewayv2_route\u0026#34; \u0026#34;vote_route\u0026#34; { api_id = aws_apigatewayv2_api.survey_api.id route_key = \u0026#34;POST /vote\u0026#34; target = \u0026#34;integrations/${aws_apigatewayv2_integration.vote_integration.id}\u0026#34; authorization_type = \u0026#34;JWT\u0026#34; authorizer_id = aws_apigatewayv2_authorizer.cognito_authorizer.id } resource \u0026#34;aws_apigatewayv2_route\u0026#34; \u0026#34;reset_route\u0026#34; { api_id = aws_apigatewayv2_api.survey_api.id route_key = \u0026#34;POST /reset\u0026#34; target = \u0026#34;integrations/${aws_apigatewayv2_integration.reset_integration.id}\u0026#34; # Note: we might want more fine-grained control here in a real app authorization_type = \u0026#34;JWT\u0026#34; authorizer_id = aws_apigatewayv2_authorizer.cognito_authorizer.id } Step 2: The Frontend - Where the Magic Happens This is where we\u0026rsquo;ll see the biggest changes. We need a way for users to sign up and sign in.\nFirst, let\u0026rsquo;s create a new login.html page. This will be a simple page with forms for sign-up and sign-in.\nWe\u0026rsquo;ll also need to include the AWS Cognito Identity SDK in our project. We can either download it and host it ourselves, or use a CDN. For simplicity, we\u0026rsquo;ll use a CDN in our HTML files.\n\u0026lt;!-- In login.html, index.html, etc. --\u0026gt; \u0026lt;script src=\u0026#34;https://cdn.jsdelivr.net/npm/amazon-cognito-identity-js@5.2.10/dist/amazon-cognito-identity.min.js\u0026#34;\u0026gt;\u0026lt;/script\u0026gt; Our main.js will need a significant update. Here are the key parts:\n// Add these variables at the top of main.js const API_URL = \u0026#39;${API_URL}\u0026#39;; const COGNITO_USER_POOL_ID = \u0026#39;${COGNITO_USER_POOL_ID}\u0026#39;; const COGNITO_CLIENT_ID = \u0026#39;${COGNITO_CLIENT_ID}\u0026#39;; const poolData = { UserPoolId: COGNITO_USER_POOL_ID, ClientId: COGNITO_CLIENT_ID }; const userPool = new AmazonCognitoIdentity.CognitoUserPool(poolData); // On page load, check if the user is logged in window.onload = function() { const idToken = localStorage.getItem(\u0026#39;idToken\u0026#39;); if (!idToken \u0026amp;\u0026amp; !window.location.href.endsWith(\u0026#39;login.html\u0026#39;)) { window.location.href = \u0026#39;login.html\u0026#39;; } else if (idToken \u0026amp;\u0026amp; window.location.href.endsWith(\u0026#39;index.html\u0026#39;)) { addSignOutButton(); } }; function addSignOutButton() { const container = document.getElementById(\u0026#39;signout-container\u0026#39;); if (container) { const signOutButton = document.createElement(\u0026#39;button\u0026#39;); signOutButton.textContent = \u0026#39;Sign Out\u0026#39;; signOutButton.onclick = signOutUser; container.appendChild(signOutButton); } } function signUpUser() { const email = document.getElementById(\u0026#39;signUpEmail\u0026#39;).value; const password = document.getElementById(\u0026#39;signUpPassword\u0026#39;).value; const messageDiv = document.getElementById(\u0026#39;message\u0026#39;); const attributeList = []; const dataEmail = { Name: \u0026#39;email\u0026#39;, Value: email, }; const attributeEmail = new AmazonCognitoIdentity.CognitoUserAttribute(dataEmail); attributeList.push(attributeEmail); userPool.signUp(email, password, attributeList, null, function(err, result){ if (err) { messageDiv.textContent = err.message || JSON.stringify(err); return; } messageDiv.textContent = \u0026#39;Sign up successful! Please check your email for a verification code.\u0026#39;; document.getElementById(\u0026#39;confirm-container\u0026#39;).style.display = \u0026#39;block\u0026#39;; }); } function confirmSignUpUser() { const email = document.getElementById(\u0026#39;signUpEmail\u0026#39;).value; const code = document.getElementById(\u0026#39;confirmationCode\u0026#39;).value; const messageDiv = document.getElementById(\u0026#39;message\u0026#39;); const userData = { Username: email, Pool: userPool }; const cognitoUser = new AmazonCognitoIdentity.CognitoUser(userData); cognitoUser.confirmRegistration(code, true, function(err, result) { if (err) { messageDiv.textContent = err.message || JSON.stringify(err); return; } messageDiv.textContent = \u0026#39;Confirmation successful! You can now sign in.\u0026#39;; document.getElementById(\u0026#39;confirm-container\u0026#39;).style.display = \u0026#39;none\u0026#39;; }); } function signInUser() { const email = document.getElementById(\u0026#39;signInEmail\u0026#39;).value; const password = document.getElementById(\u0026#39;signInPassword\u0026#39;).value; const messageDiv = document.getElementById(\u0026#39;message\u0026#39;); const authenticationData = { Username: email, Password: password, }; const authenticationDetails = new AmazonCognitoIdentity.AuthenticationDetails(authenticationData); const userData = { Username: email, Pool: userPool }; const cognitoUser = new AmazonCognitoIdentity.CognitoUser(userData); cognitoUser.authenticateUser(authenticationDetails, { onSuccess: function (result) { const idToken = result.getIdToken().getJwtToken(); localStorage.setItem(\u0026#39;idToken\u0026#39;, idToken); window.location.href = \u0026#39;index.html\u0026#39;; }, onFailure: function(err) { messageDiv.textContent = err.message || JSON.stringify(err); }, }); } function signOutUser() { localStorage.removeItem(\u0026#39;idToken\u0026#39;); const cognitoUser = userPool.getCurrentUser(); if (cognitoUser) { cognitoUser.signOut(); } window.location.href = \u0026#39;login.html\u0026#39;; } async function vote(option) { const idToken = localStorage.getItem(\u0026#39;idToken\u0026#39;); if (!idToken) { window.location.href = \u0026#39;login.html\u0026#39;; return; } const messageDiv = document.getElementById(\u0026#39;message\u0026#39;); messageDiv.textContent = \u0026#39;Submitting your vote...\u0026#39;; try { const response = await fetch(`${API_URL}vote`, { method: \u0026#39;POST\u0026#39;, headers: { \u0026#39;Content-Type\u0026#39;: \u0026#39;application/json\u0026#39;, \u0026#39;Authorization\u0026#39;: `Bearer ${idToken}` }, body: JSON.stringify({ vote: option }), }); if (!response.ok) { throw new Error(`HTTP error! status: $${response.status}`); } messageDiv.textContent = \u0026#39;Thank you for voting!\u0026#39;; document.querySelectorAll(\u0026#39;.vote-btn\u0026#39;).forEach(button =\u0026gt; { button.disabled = true; }); } catch (error) { console.error(\u0026#39;Error submitting vote:\u0026#39;, error); messageDiv.textContent = \u0026#39;Sorry, there was an error submitting your vote.\u0026#39;; } } Step 3: Backend Adjustments Our backend Lambda functions need a small change. Since we are now identifying users by their JWT, we can remove the sessionId logic. The vote.py function will now get the user\u0026rsquo;s unique ID from the JWT claims that API Gateway passes along.\nHere\u0026rsquo;s how the Lambda function processes an authenticated request:\n# In backend/vote.py def handler(event, context): try: body = json.loads(event.get(\u0026#39;body\u0026#39;, \u0026#39;{}\u0026#39;)) vote_option = body.get(\u0026#39;vote\u0026#39;) # Get user id from the authorizer context # API Gateway extracts this from the JWT and passes it to Lambda user_id = event[\u0026#39;requestContext\u0026#39;][\u0026#39;authorizer\u0026#39;][\u0026#39;jwt\u0026#39;][\u0026#39;claims\u0026#39;][\u0026#39;sub\u0026#39;] if not user_id: return {\u0026#39;statusCode\u0026#39;: 400, ...} if vote_option not in [\u0026#39;no\u0026#39;, \u0026#39;aws\u0026#39;, \u0026#39;other\u0026#39;]: return {\u0026#39;statusCode\u0026#39;: 400, ...} item = { \u0026#39;id\u0026#39;: user_id, # Use the cognito user id as the primary key \u0026#39;vote\u0026#39;: vote_option } table.put_item(Item=item) return {\u0026#39;statusCode\u0026#39;: 200, ...} except Exception as e: # ... What\u0026rsquo;s happening here?\nThe sub (subject) claim in the JWT is a unique identifier for each Cognito user API Gateway validates the JWT before the Lambda even runs If the JWT is invalid or missing, the request never reaches our Lambda We use the user\u0026rsquo;s sub as the primary key in DynamoDB, ensuring one vote per user Security is Not a Feature, It\u0026rsquo;s a Foundation Let\u0026rsquo;s talk about the security improvements we\u0026rsquo;ve made:\nManaged User Directory: We are not storing passwords. Cognito handles all password policies, hashing, and storage, following best practices. JWT Authentication: We\u0026rsquo;re using the industry standard for API authentication. The JWTs are signed by Cognito, and our API Gateway verifies this signature on every request. This prevents token tampering. Secure Token Storage: We are storing the JWT in localStorage. This is a common practice, but it\u0026rsquo;s important to be aware of the risks (like XSS attacks). For higher security applications, we could store tokens in memory and use refresh tokens to get new access tokens. HTTPS Everywhere: Our entire application, from the frontend on CloudFront to the API Gateway, enforces HTTPS. This prevents eavesdropping. Least Privilege: Our Lambda functions have fine-grained IAM roles. They can only access the DynamoDB table they need. Understanding Our Lambda Functions Let\u0026rsquo;s break down what each Lambda function does in our application:\nVote Lambda Function This function processes vote submissions from authenticated users.\nKey responsibilities:\nExtract the vote option from the request body Get the user\u0026rsquo;s unique ID from the JWT claims (provided by API Gateway) Validate the vote is one of the allowed options Store the vote in DynamoDB using the user ID as the key (prevents duplicate votes) Results Lambda Function This function retrieves and counts all votes. No authentication required - results are public!\nKey responsibilities:\nScan the entire DynamoDB table to get all votes Handle pagination (DynamoDB returns max 1MB per request) Count votes for each option using Python\u0026rsquo;s Counter Return totals as JSON: {\u0026quot;no\u0026quot;: 5, \u0026quot;aws\u0026quot;: 12, \u0026quot;other\u0026quot;: 3} Reset Lambda Function This function deletes all votes. Requires authentication to prevent abuse.\nKey responsibilities:\nScan the table to get all item IDs Use batch writer to efficiently delete items (groups of 25) Handle pagination for large datasets Return success confirmation Email Filter Lambda Function This is a special type of Lambda called a Cognito Trigger. It runs automatically before a user signs up.\nKey responsibilities:\nExtract the email from the sign-up request Check if it ends with @charlotte.edu Allow or reject the sign-up based on the domain This enforces organization-level access control The Final Result After implementing these changes, our complete authentication flow looks like this:\nWe\u0026rsquo;ve now added a robust and secure authentication layer to our serverless application, moving it from a simple demo to a more production-ready state. This is the power of leveraging managed services like AWS Cognito!\nBonus: Restricting Sign-ups to a Specific Email Domain A common requirement is to restrict application access to users from a specific organization. We can easily extend our Cognito setup to only allow sign-ups from users with a @charlotte.edu email address.\nThis is accomplished using a Cognito Pre Sign-up Lambda Trigger. This trigger fires just before Cognito creates a new user, giving us a chance to run custom validation logic.\nWhat\u0026rsquo;s a Lambda Trigger? Think of it like a hook or event listener. Cognito has several points in the user lifecycle where it can automatically invoke a Lambda function:\nPre Sign-up: Before creating a new user (we use this one!) Post Confirmation: After a user confirms their email Pre Authentication: Before signing in Post Authentication: After signing in These triggers let you customize Cognito\u0026rsquo;s behavior without modifying AWS\u0026rsquo;s code.\nStep 1: Create the Email Filter Lambda We\u0026rsquo;ll create a new Lambda function in backend/email-filter.py:\nimport json def handler(event, context): # This trigger is invoked before a user is signed up email = event[\u0026#39;request\u0026#39;][\u0026#39;userAttributes\u0026#39;].get(\u0026#39;email\u0026#39;) if email and email.endswith(\u0026#39;@charlotte.edu\u0026#39;): # Allow sign-up return event else: # Block sign-up raise Exception(\u0026#34;Only users with a @charlotte.edu email address are allowed to sign up.\u0026#34;) Step 2: Update Terraform Now, we\u0026rsquo;ll update our terraform/main.tf to create this new Lambda and associate it with our User Pool.\ndata \u0026#34;archive_file\u0026#34; \u0026#34;email_filter_zip\u0026#34; { type = \u0026#34;zip\u0026#34; source_file = \u0026#34;../backend/email-filter.py\u0026#34; output_path = \u0026#34;email-filter.zip\u0026#34; } resource \u0026#34;aws_lambda_function\u0026#34; \u0026#34;email_filter_function\u0026#34; { function_name = \u0026#34;CrackingTheCloudEmailFilter\u0026#34; role = aws_iam_role.lambda_exec_role.arn handler = \u0026#34;email-filter.handler\u0026#34; runtime = \u0026#34;python3.9\u0026#34; filename = data.archive_file.email_filter_zip.output_path source_code_hash = data.archive_file.email_filter_zip.output_base64sha256 } resource \u0026#34;aws_cognito_user_pool\u0026#34; \u0026#34;survey_user_pool\u0026#34; { name = \u0026#34;SurveyUserPool\u0026#34; auto_verified_attributes = [\u0026#34;email\u0026#34;] lambda_config { pre_sign_up = aws_lambda_function.email_filter_function.arn } } resource \u0026#34;aws_lambda_permission\u0026#34; \u0026#34;cognito_permission_email_filter\u0026#34; { statement_id = \u0026#34;AllowCognitoToInvokeEmailFilter\u0026#34; action = \u0026#34;lambda:InvokeFunction\u0026#34; function_name = aws_lambda_function.email_filter_function.function_name principal = \u0026#34;cognito-idp.amazonaws.com\u0026#34; source_arn = aws_cognito_user_pool.survey_user_pool.arn } With these changes, any attempt to sign up with an email address that does not end in @charlotte.edu will be rejected by Cognito. This is a powerful way to enforce organization-specific access control.\nKey Takeaways for Students What you\u0026rsquo;ve learned:\nServerless Authentication: No need to build your own user management system from scratch JWT Tokens: Industry-standard way to authenticate API requests Lambda Triggers: Extend AWS services with custom logic at specific lifecycle events API Gateway Authorizers: Validate tokens before requests reach your backend code Infrastructure as Code: All of this is defined in Terraform and can be deployed in minutes Real-world applications:\nStudent portals restricted to university email domains Internal company tools that only employees can access Multi-tenant SaaS applications where each organization has isolated access Mobile apps that need secure backend APIs Cost considerations:\nCognito: First 50,000 monthly active users are free Lambda: First 1 million requests per month are free DynamoDB: 25 GB of storage free API Gateway: First 1 million API calls per month are free For a student project or small application, this entire stack runs essentially for free!\nNext Steps Want to take this further? Here are some ideas:\nAdd user profiles: Store additional user data in DynamoDB Implement admin roles: Use Cognito groups to create admin users with special permissions Add MFA: Enable multi-factor authentication for extra security Social sign-in: Allow users to sign in with Google, Facebook, or other providers Password reset flow: Implement \u0026ldquo;forgot password\u0026rdquo; functionality Email customization: Customize the verification emails Cognito sends The foundation you\u0026rsquo;ve built here is production-ready and can scale to millions of users!\nGitHub Repository: https://github.com/lukelittle/adding-cognito-to-our-survey-app\n","permalink":"https://lukelittle.com/posts/2026/01/enhancing-security-adding-aws-cognito-authentication-to-your-serverless-app/","summary":"\u003cp\u003eOur serverless survey application is a great example of a modern cloud native application. It\u0026rsquo;s fast, scalable, and cost-effective. But it\u0026rsquo;s missing one critical feature: user authentication. In this post, we\u0026rsquo;ll walk through how to add robust, secure authentication using AWS Cognito.\u003c/p\u003e\n\u003ch2 id=\"why-add-authentication\"\u003eWhy Add Authentication?\u003c/h2\u003e\n\u003cp\u003eRight now, anyone can vote, and anyone can reset the entire survey. In a real-world application, we need to control access. Authentication allows us to:\u003c/p\u003e","title":"Enhancing Security: Adding AWS Cognito Authentication to Your Serverless App"},{"content":"At an AWS Road Show this fall, Darko Mesaros demoed a URL shortener he\u0026rsquo;d built in Rust called krtk.rs. Something about watching a clean, fast URL shortener just work stuck with me. I\u0026rsquo;ve built a few of these for demos since then, but I wanted to try something different this time: build one in Python with a retro 90s vibe, and let Kiro handle most of the heavy lifting.\nKiro is one of AWS\u0026rsquo;s three frontier agents announced at re:Invent 2025—autonomous AI systems that maintain context and work independently for hours. While DevOps Agent handles incident response and Security Agent conducts penetration testing, Kiro is your AI developer that takes specifications and generates production-ready code.\nThis turned into a great experiment in spec-driven AI development. Here\u0026rsquo;s what I learned about building with AI coding assistants.\nThe approach: Spec-driven development with Kiro I started by asking ChatGPT to help me write a comprehensive prompt for Kiro. The key insight: the better your specification, the better the AI output.\nHere\u0026rsquo;s what I asked for:\nYou are a senior AWS serverless engineer. Take my existing project and refactor/extend it into a URL shortener with a retro 90s website UI.\nThen I got detailed with requirements:\nFrontend: A simple 90s-style static site where authenticated users can submit URLs and get short links back. Show their created links and click counts. Auth: Amazon Cognito for create/delete/list operations. Public redirects don\u0026rsquo;t need auth. Data: DynamoDB table with link as partition key, plus url, count, creator, and created timestamp. APIs: POST /links (create), DELETE /links/{link} (delete), GET /links (list), GET /{link} (public redirect) Lambdas in Python: create_link, delete_link, list_links, redirect, plus analytics functions Visit counting had to be decoupled: redirect fires an event to Kinesis, consumer Lambda batches events and updates DynamoDB counts atomically The decoupled visit counting was critical—I\u0026rsquo;ve seen too many URL shorteners block redirects waiting for analytics writes. That\u0026rsquo;s how you turn a 50ms redirect into a 150ms redirect.\nWhat Kiro generated Kiro produced a comprehensive 15-section specification document covering functional requirements, API contracts, data models, security, and cost analysis. This spec became the foundation for everything else.\nThen it translated that into concrete design decisions:\nAnd broke the implementation into actionable tasks:\nFinally, it generated the complete architecture:\nThis is production-grade serverless: S3 + CloudFront for the frontend, API Gateway for the REST API, Lambda for compute, DynamoDB for storage, Kinesis for event streaming.\nThe visit counting architecture is particularly clever:\nUser hits a short link → redirect Lambda looks it up in DynamoDB Lambda immediately returns a 301 redirect (fast!) Then it fires an event to Kinesis (fire-and-forget, non-blocking) Kinesis batches up events A consumer Lambda processes batches and updates DynamoDB counts atomically This approach delivers redirects under 200ms and reduces DynamoDB writes by 100x. For a link getting 1,000 clicks/minute, that\u0026rsquo;s the difference between $75/month and $0.75/month just for counting.\nUnderstanding DynamoDB through AI One of the valuable learning moments came when I asked Kiro about the DynamoDB schema:\nKiro explained that DynamoDB only requires key attributes defined in Terraform—non-key attributes like count and url are dynamic. The GSI projection (projection_type = \u0026quot;ALL\u0026quot;) automatically includes everything. This kind of on-demand explanation is where AI assistants really shine.\nThe iterative refinement process During deployment, I discovered Kiro had generated all the Lambda functions, DynamoDB tables, and API Gateway configs—but initially missed the S3 bucket and CloudFront distribution for the frontend.\nThis is normal in iterative development. I pointed it out, and Kiro immediately generated:\nS3 bucket with proper access controls CloudFront distribution with Origin Access Control Bucket policies Output values for deployment This iterative feedback loop is exactly how development works—AI or not.\nConfiguration refinements The first deployment revealed Kiro had configured CloudFront with an /api/* cache behavior that forwarded headers to S3. S3 doesn\u0026rsquo;t support that pattern because:\nS3 serves static files (HTML, CSS, JS) API Gateway serves API endpoints (/links, /health, etc.) CloudFront should only cache static content Removing that behavior simplified the architecture and everything worked perfectly. AI-generated code can sometimes include extra patterns—simplifying is part of the refinement process.\nIntegration patterns The frontend initially used simulated JWT tokens while the backend had real Cognito validation configured. This kind of integration gap is common when components are generated separately.\nAdding proper Cognito authentication was straightforward:\nconst poolData = { UserPoolId: COGNITO_USER_POOL_ID, ClientId: COGNITO_CLIENT_ID }; const userPool = new AmazonCognitoIdentity.CognitoUserPool(poolData); cognitoUser.authenticateUser(authenticationDetails, { onSuccess: function (result) { authToken = result.getIdToken().getJwtToken(); showApp(); loadUserLinks(); }, onFailure: function(err) { if (err.code === \u0026#39;UserNotConfirmedException\u0026#39;) { showError(\u0026#39;Check your email to confirm your account.\u0026#39;); } else if (err.code === \u0026#39;NotAuthorizedException\u0026#39;) { showError(\u0026#39;Invalid email or password.\u0026#39;); } } }); Integration complete. System working end-to-end.\nWhat worked exceptionally well Kiro generated production-quality code across multiple areas:\nKinesis batching: The visit counting pipeline processed events exactly as designed, delivering that 100x cost reduction.\nDynamoDB schema: Properly implemented with dynamic non-key attributes and a well-designed GSI for querying by creator.\nTerraform structure: Clean, modular IaC with proper IAM roles using least-privilege permissions and good tagging for cost allocation.\nLambda functions: All eight Python functions came with error handling, structured logging, and CloudWatch metrics built in.\nCost modeling: Kiro accurately estimated ~$15/month for 100K redirects. For comparison, running this on containers would cost 10x more.\nI even asked Kiro if the system was production-ready:\nIt correctly identified what was complete and what still needed work.\nWhy spec-driven AI development works This project demonstrated several key principles:\nClear specifications produce better results: The comprehensive prompt I gave Kiro led to comprehensive, well-architected output. Garbage in, garbage out—quality specifications get quality code.\nAI excels at boilerplate: Kiro generated thousands of lines of Terraform, Lambda functions, and infrastructure configurations that would\u0026rsquo;ve taken me 10-15 hours to write manually. This is where AI delivers massive productivity gains.\nIteration is normal: Whether working with AI or human developers, you iterate. Point out gaps, refine implementations, simplify where needed. The feedback loop is fast with AI.\nArchitecture decisions matter: The Kinesis batching decision came from my prompt. AI implemented it perfectly. The human provides the \u0026ldquo;why,\u0026rdquo; the AI handles the \u0026ldquo;how.\u0026rdquo;\nLearning opportunity: Using Kiro taught me DynamoDB patterns I hadn\u0026rsquo;t used before. AI assistants can explain concepts while implementing them.\nThe final system Everything deployed and works in production:\nS3 + CloudFront hosting the 90s-themed frontend API Gateway + Lambda handling authenticated operations Cognito managing users via backend endpoints DynamoDB storing links with a creator-created GSI for queries Kinesis streaming visit events for batch processing CloudWatch logs, metrics, and alarms for observability Total cost: ~$15/month for 100K redirects. The code is open source: github.com/lukelittle/url-shortener\nThe value of AI coding assistants Working with Kiro on this project showed me why AI coding assistants are becoming essential:\n10-15 hours of coding compressed into 2-3 hours of specification and refinement. That\u0026rsquo;s a genuine productivity multiplier.\nSpec-driven development forces better architecture. Writing a comprehensive prompt made me think through requirements more carefully than I usually do for side projects.\nLearning while building. Kiro explained concepts as it generated code, turning implementation into education.\nFocus on what matters. Instead of writing boilerplate Terraform and Lambda handlers, I spent time on architecture decisions and integration patterns—the parts that actually require human judgment.\nThe 60/40 split I experienced (60% worked immediately, 40% needed refinement) is impressive for a first iteration. With better prompts and tighter feedback loops, that ratio keeps improving.\nWould I use AI for this again? Absolutely. Kiro didn\u0026rsquo;t replace my AWS knowledge—it amplified it. I still made the architectural decisions, understood the tradeoffs, and guided the implementation. But instead of spending days writing infrastructure code, I spent hours refining specifications and reviewing output.\nThis is the future: engineers focus on architecture, requirements, and integration patterns while AI handles implementation details. The tools get better every month.\nThe AI writes the code. You make sure it\u0026rsquo;s the right code.\n","permalink":"https://lukelittle.com/posts/2026/01/15-hours-of-terraform-in-3-building-with-aws-kiro/","summary":"\u003cp\u003eAt an AWS Road Show this fall, Darko Mesaros demoed a URL shortener he\u0026rsquo;d built in Rust called \u003ca href=\"https://github.com/darko-mesaros/krtk.rs\"\u003ekrtk.rs\u003c/a\u003e. Something about watching a clean, fast URL shortener just \u003cem\u003ework\u003c/em\u003e stuck with me. I\u0026rsquo;ve built a few of these for demos since then, but I wanted to try something different this time: build one in Python with a retro 90s vibe, and let Kiro handle most of the heavy lifting.\u003c/p\u003e\n\u003cp\u003eKiro is one of AWS\u0026rsquo;s three frontier agents announced at re:Invent 2025—autonomous AI systems that maintain context and work independently for hours. While DevOps Agent handles incident response and Security Agent conducts penetration testing, Kiro is your AI developer that takes specifications and generates production-ready code.\u003c/p\u003e","title":"15 Hours of Terraform in 3: Building with AWS Kiro"},{"content":"On October 20, 2025, DNS resolution failed in AWS us-east-1, and with it, a lot of DynamoDB applications went down.\nNot because DynamoDB itself failed. The service was running. Data was there. Capacity was fine. But applications couldn\u0026rsquo;t reach it because DNS queries for dynamodb.us-east-1.amazonaws.com stopped resolving correctly.\nIf you\u0026rsquo;ve ever wondered what happens when the infrastructure layer beneath your supposedly resilient database becomes unreachable—October 20 was the answer. And it wasn\u0026rsquo;t pretty.\nI\u0026rsquo;ve been thinking about that outage a lot lately, especially in the context of the survey application I built to teach students about serverless architecture. That app would have completely failed on October 20. The frontend would load from CloudFront, but every API call would hit a wall trying to reach DynamoDB in us-east-1.\nSo I decided to figure out what it would actually take to make that architecture survive a regional DNS failure. Not in theory—in practice, with real Terraform and honest tradeoffs.\nHere\u0026rsquo;s what I learned.\nWhy \u0026ldquo;Highly Available\u0026rdquo; Wasn\u0026rsquo;t Enough After the outage, I went back and reviewed how I\u0026rsquo;d configured DynamoDB for the survey app. On paper, it looked solid:\nOn-demand capacity (no throttling to worry about) SDK retries with exponential backoff Adaptive capacity for hot partitions CloudWatch alarms watching for errors This is pretty much the standard DynamoDB resilience checklist. And on October 20, none of it helped.\nBecause retries don\u0026rsquo;t work if the endpoint can\u0026rsquo;t be resolved.\nThe SDK would try to send a request, fail at DNS resolution, retry with backoff, fail again, retry longer, fail again—all hitting the same DNS wall. Exponential backoff just meant it took longer to give up.\nAdaptive capacity is irrelevant when requests never reach the service. The alarms fired, but by the time anyone saw them, users had already moved on.\nHere\u0026rsquo;s the uncomfortable realization: most DynamoDB resilience advice assumes the service is reachable. It\u0026rsquo;s optimized for throttling, hot keys, capacity planning. It doesn\u0026rsquo;t address what happens when the layer beneath DynamoDB becomes unavailable.\nAnd that\u0026rsquo;s exactly what happened on October 20.\nGlobal Tables: What They Solve (and What They Very Much Don\u0026rsquo;t) After the outage, the obvious question was: would DynamoDB Global Tables have saved us?\nGlobal Tables give you automated multi-region replication. Write to a table in us-east-1, and the data shows up in us-west-2 within a second or two. It\u0026rsquo;s designed for disaster recovery and geographic distribution.\nBut here\u0026rsquo;s what Global Tables do not give you:\nAutomatic application failover DNS independence Transparent client redirection If your application is configured to talk to dynamodb.us-east-1.amazonaws.com and that endpoint becomes unreachable, Global Tables don\u0026rsquo;t help. Your data sits in us-west-2, perfectly healthy and accessible—but your application never tries to use it.\nThis is where a lot of architects\u0026rsquo; mental models break down. They think: \u0026ldquo;I have Global Tables, so my data is replicated. I\u0026rsquo;m resilient.\u0026rdquo;\nNot quite.\nThe Hidden Assumption That Broke Everything Even with Global Tables configured, most applications still have this somewhere in the code:\ndynamodb = boto3.resource(\u0026#39;dynamodb\u0026#39;, region_name=\u0026#39;us-east-1\u0026#39;) Or this in their Terraform:\nprovider \u0026#34;aws\u0026#34; { region = \u0026#34;us-east-1\u0026#34; } The application is hard-wired to us-east-1. It knows about one region. It sends all traffic there. If DNS in that region fails, the application fails—regardless of how many replica tables exist elsewhere.\nThis is the \u0026ldquo;oh shit\u0026rdquo; moment: you replicated your data across the globe, but your application never learned to look anywhere else.\nGlobal Tables solve the data availability problem. They don\u0026rsquo;t solve the application failover problem. And on October 20, it was the second one that broke.\nAvailability vs Consistency (Suddenly This Matters) DynamoDB has two consistency modes:\nStrongly consistent reads: Always return the most recent write Eventually consistent reads: Might return slightly stale data, but stay available during replication lag In normal operations, this is mostly academic. But during a regional DNS failure, the tradeoff becomes very real.\nIf you fail over to a replica region and issue strongly consistent reads, those reads might succeed or fail depending on replication lag and table state. If you use eventually consistent reads, they\u0026rsquo;ll work—but you might see data that\u0026rsquo;s a few seconds behind.\nFor the survey app, the choice is obvious: show slightly stale vote counts. A result that\u0026rsquo;s 10 seconds behind is infinitely better than no result at all. Users can still vote. The system degrades gracefully instead of falling over.\nBut you have to make that decision before the outage, not during it. You can\u0026rsquo;t architect consistency tradeoffs while your dashboard is on fire.\nFailure Domains Nobody Models Most teams model failures like this:\nAvailability Zone goes down Service gets throttled Account gets compromised DNS typically doesn\u0026rsquo;t make the list. It\u0026rsquo;s infrastructure—it just works.\nUntil it doesn\u0026rsquo;t.\nOctober 20 revealed that regional DNS resolution is a shared fate dependency. When it fails, everything that depends on it fails together. API Gateway, Lambda, DynamoDB, S3—if those services are addressed via DNS in the affected region, you can\u0026rsquo;t reach them.\nThe survey app has this dependency everywhere: API Gateway in us-east-1 calls Lambda in us-east-1, which calls DynamoDB in us-east-1. Every single one of those calls requires DNS resolution. One DNS failure takes down the entire stack.\nA \u0026ldquo;global service\u0026rdquo; like DynamoDB doesn\u0026rsquo;t mean no regional failure modes. It means you need to understand which parts have regional dependencies—and DNS is absolutely one of them.\nWarm vs Cold Failover (Pick Your Pain) There are two ways to handle multi-region failover:\nCold failover means your secondary region exists but isn\u0026rsquo;t actively serving traffic. When the primary fails, you manually redirect traffic. This is cheaper—you\u0026rsquo;re only paying for data replication—but recovery takes longer.\nWarm failover means your secondary region is live, running infrastructure, and ready to take over immediately. Both regions serve traffic. When one fails, users barely notice. This costs more because you\u0026rsquo;re running duplicate infrastructure.\nFor the survey app, cold failover looks like:\nDynamoDB Global Table replicating votes to us-west-2 No Lambda, API Gateway, or CloudFront config for us-west-2 Manual terraform apply when things go wrong Warm failover looks like:\nFull stack deployed in both regions Route 53 health checks watching both Automatic traffic shifting when health checks fail DNS-level failures complicate this decision. With a service outage, cold failover\u0026rsquo;s longer recovery might be acceptable. But when DNS fails, even logging into the console to trigger failover might be impacted.\nWarm failover starts looking a lot more attractive when \u0026ldquo;manually fail over\u0026rdquo; might not be possible.\nHow Traffic Actually Flips (The Part Everyone Handwaves) Failover is a decision, not a default. Something has to decide when to switch regions and actually execute that switch.\nFor the survey app, there are a few options:\nApplication-level region awareness\nThe frontend JavaScript knows about both regions and tries the secondary if the primary fails:\nconst regions = [ { api: \u0026#39;https://api-us-east-1.example.com\u0026#39;, name: \u0026#39;us-east-1\u0026#39; }, { api: \u0026#39;https://api-us-west-2.example.com\u0026#39;, name: \u0026#39;us-west-2\u0026#39; } ]; async function submitVote(vote) { for (const region of regions) { try { const response = await fetch(`${region.api}/vote`, { method: \u0026#39;POST\u0026#39;, body: JSON.stringify({ vote }) }); return response; } catch (error) { console.log(`${region.name} failed, trying next`); } } throw new Error(\u0026#39;All regions failed\u0026#39;); } This works, but now your frontend code is aware of regional infrastructure. That complexity leaks all the way to the browser.\nHere\u0026rsquo;s what this application-level failover looks like in practice:\nsequenceDiagram autonumber participant Client participant Route53 as Route 53 participant APIGW_East as API Gateway\u0026lt;br/\u0026gt;(us-east-1) participant Lambda_East as Lambda\u0026lt;br/\u0026gt;(us-east-1) participant DDB_East as DynamoDB\u0026lt;br/\u0026gt;(us-east-1) participant APIGW_West as API Gateway\u0026lt;br/\u0026gt;(us-west-2) participant Lambda_West as Lambda\u0026lt;br/\u0026gt;(us-west-2) participant DDB_West as DynamoDB\u0026lt;br/\u0026gt;(us-west-2) Note over Client,DDB_West: Scenario 1: Normal Operation (us-east-1 healthy) Client-\u0026gt;\u0026gt;Route53: Request vote submission Route53-\u0026gt;\u0026gt;APIGW_East: Route to us-east-1 APIGW_East-\u0026gt;\u0026gt;Lambda_East: Invoke function Lambda_East-\u0026gt;\u0026gt;DDB_East: Write vote DDB_East--\u0026gt;\u0026gt;Lambda_East: Success Lambda_East--\u0026gt;\u0026gt;APIGW_East: 200 OK APIGW_East--\u0026gt;\u0026gt;Client: Vote recorded ✓ Note over Client,DDB_West: Scenario 2: October 20 DNS Failure (no failover) Client-\u0026gt;\u0026gt;Route53: Request vote submission Route53-xAPIGW_East: DNS resolution fails ❌ Note over Client: Request times out\u0026lt;br/\u0026gt;User sees error Note over Client,DDB_West: Scenario 3: DNS Failure with Application Failover Client-\u0026gt;\u0026gt;Route53: Request vote submission Route53-xAPIGW_East: DNS resolution fails ❌ Note over Client: Client catches error,\u0026lt;br/\u0026gt;retries us-west-2 Client-\u0026gt;\u0026gt;Route53: Retry request Route53-\u0026gt;\u0026gt;APIGW_West: Route to us-west-2 APIGW_West-\u0026gt;\u0026gt;Lambda_West: Invoke function Lambda_West-\u0026gt;\u0026gt;DDB_West: Write vote (replica) DDB_West--\u0026gt;\u0026gt;Lambda_West: Success Lambda_West--\u0026gt;\u0026gt;APIGW_West: 200 OK APIGW_West--\u0026gt;\u0026gt;Client: Vote recorded ✓ Note over DDB_East,DDB_West: Cross-region replication\u0026lt;br/\u0026gt;(\u0026lt; 1 second) The diagram illustrates three scenarios:\nNormal operation: Everything works in us-east-1 October 20 failure: DNS fails and users see errors With application failover: When us-east-1 fails, the client automatically retries with us-west-2 and succeeds Feature flags or configuration toggles\nAn external service (like LaunchDarkly) controls which region receives traffic. During an outage, ops flips the flag. This centralizes the logic but adds another dependency—and another potential failure point.\nRoute 53 health checks\nDNS-based failover that automatically routes to healthy endpoints. This works for many scenarios, but if regional DNS is failing, Route 53 lookups might also be impacted.\nThe uncomfortable truth: there\u0026rsquo;s no perfect solution. Each approach has tradeoffs around complexity, blast radius, and new failure modes.\nMaking the Survey App Actually Resilient Let\u0026rsquo;s revisit the serverless survey application. The original architecture:\nS3 + CloudFront serving the frontend API Gateway exposing REST endpoints Lambda functions (vote, results, reset) DynamoDB storing votes Everything in us-east-1 On October 20, this would have completely failed. CloudFront would serve the static site from its global cache, but every API call would hit DNS failures trying to reach API Gateway in us-east-1.\nStudents would see the survey form but couldn\u0026rsquo;t vote. The results page would spin forever. The reset function would be unreachable.\nTo make this resilient, I\u0026rsquo;d implement warm failover with application-level region awareness.\nWhat keeps working during a DNS failure:\nStatic frontend loads from CloudFront (it\u0026rsquo;s global anyway) Vote submissions succeed by failing over to us-west-2 API Results page shows counts from the replica table Eventually consistent reads mean slight lag is acceptable What degrades gracefully:\nLatency goes up for users far from the secondary region Global Table replication lag means new votes take longer to appear everywhere No strong consistency guarantees during failover What stops working:\nNothing critical (that\u0026rsquo;s the entire point) This is graceful degradation instead of complete failure.\nWhat Multi-Region DynamoDB Actually Looks Like Here\u0026rsquo;s the Terraform for making the survey app multi-region:\n# Primary region provider \u0026#34;aws\u0026#34; { alias = \u0026#34;primary\u0026#34; region = \u0026#34;us-east-1\u0026#34; } # Secondary region provider \u0026#34;aws\u0026#34; { alias = \u0026#34;secondary\u0026#34; region = \u0026#34;us-west-2\u0026#34; } # DynamoDB table with Global Tables enabled resource \u0026#34;aws_dynamodb_table\u0026#34; \u0026#34;survey_votes\u0026#34; { provider = aws.primary name = \u0026#34;survey-votes\u0026#34; billing_mode = \u0026#34;PAY_PER_REQUEST\u0026#34; hash_key = \u0026#34;id\u0026#34; stream_enabled = true stream_view_type = \u0026#34;NEW_AND_OLD_IMAGES\u0026#34; attribute { name = \u0026#34;id\u0026#34; type = \u0026#34;S\u0026#34; } # This one line enables Global Tables replica { region_name = \u0026#34;us-west-2\u0026#34; } tags = { Environment = \u0026#34;production\u0026#34; } } # Lambda in primary region resource \u0026#34;aws_lambda_function\u0026#34; \u0026#34;vote_primary\u0026#34; { provider = aws.primary function_name = \u0026#34;survey-vote\u0026#34; runtime = \u0026#34;python3.11\u0026#34; handler = \u0026#34;vote.handler\u0026#34; filename = \u0026#34;lambda/vote.zip\u0026#34; environment { variables = { TABLE_NAME = aws_dynamodb_table.survey_votes.name REGION = \u0026#34;us-east-1\u0026#34; } } } # Lambda in secondary region (same code, different region) resource \u0026#34;aws_lambda_function\u0026#34; \u0026#34;vote_secondary\u0026#34; { provider = aws.secondary function_name = \u0026#34;survey-vote\u0026#34; runtime = \u0026#34;python3.11\u0026#34; handler = \u0026#34;vote.handler\u0026#34; filename = \u0026#34;lambda/vote.zip\u0026#34; environment { variables = { TABLE_NAME = aws_dynamodb_table.survey_votes.name REGION = \u0026#34;us-west-2\u0026#34; } } } The critical piece is the replica block. That tells DynamoDB to automatically replicate to us-west-2. AWS handles the replication—you don\u0026rsquo;t write code for it.\nBut notice: you still need to deploy Lambda, API Gateway, and all the supporting infrastructure in both regions. Global Tables replicate data, not infrastructure.\nArchitecture Diagram Description The architecture flows left to right:\nFar left: Client layer\nUser browsers, mobile apps, any HTTP client Edge layer (global)\nCloudFront distribution serving static files (HTML, CSS, JS) Label: \u0026ldquo;Global Edge Network\u0026rdquo; Middle: Application layer (two parallel stacks)\nPrimary region stack (us-east-1):\nAPI Gateway endpoint Three Lambda functions (vote, results, reset) Connected to DynamoDB primary table Status indicator: normally green (\u0026ldquo;Healthy\u0026rdquo;), red during October 20 (\u0026ldquo;DNS Failed\u0026rdquo;) Secondary region stack (us-west-2):\nIdentical API Gateway endpoint Identical Lambda functions Connected to DynamoDB replica table Status indicator: green (\u0026ldquo;Healthy\u0026rdquo;) Far right: Data layer\nTwo DynamoDB tables shown side-by-side Primary table (us-east-1) Replica table (us-west-2) Bi-directional arrows between them labeled \u0026ldquo;\u0026lt; 1s replication\u0026rdquo; Both showing identical data Failover control (dashed line across diagram):\nRoute 53 health checks monitoring both regions Decision diamond: \u0026ldquo;Primary healthy?\u0026rdquo; Yes → route to us-east-1 No → route to us-west-2 Alternative path: client-side retry logic (try primary, fall back to secondary) DNS dependency markers (red warning icons):\nDNS resolution required for API Gateway in each region DNS resolution required for DynamoDB endpoints in each region These are the exact failure points October 20 exposed Bottom: Observability signals\nCloudWatch alarms in both regions Metrics: API error rate, Lambda duration, DynamoDB throttles Threshold shown: \u0026ldquo;Error rate \u0026gt; 5% for 2 minutes → consider failover\u0026rdquo; The diagram makes one thing brutally clear: when DNS fails in us-east-1, the entire primary stack becomes unreachable—but us-west-2 keeps running. Data replication keeps them in sync. Application failover keeps users working.\nTradeoffs Nobody Wants to Talk About This architecture is more complex than single-region. You\u0026rsquo;re maintaining infrastructure in two regions, managing replication, handling eventual consistency, and building failover logic.\nFor the survey app—which costs $0/month in a single region—running warm failover in two regions might cost $15/month. That\u0026rsquo;s not a lot in absolute terms, but it\u0026rsquo;s infinitely more expensive than free.\nAnd here\u0026rsquo;s an uncomfortable truth: most teams won\u0026rsquo;t actually test their failover. They\u0026rsquo;ll build it, deploy it, document it, assume it works, and never validate it until a real outage happens.\nWhen October 20 comes, they\u0026rsquo;ll discover:\nTheir failover mechanism has a bug Their health checks have false positives Their SDK client caching interferes with region switching Their observability doesn\u0026rsquo;t show which region is actually serving traffic Multi-region resilience requires ongoing operational investment. It\u0026rsquo;s not \u0026ldquo;set and forget.\u0026rdquo; You need runbooks, chaos engineering, regular failover drills, and teams who know how to operate it.\nThat\u0026rsquo;s a lot of overhead for a student survey app.\nWho Actually Needs This Not every DynamoDB workload needs multi-region failover.\nYou should build this if:\nDowntime directly costs revenue You have strict uptime SLAs (99.95%+) Users are globally distributed Regulatory requirements mandate geographic redundancy You\u0026rsquo;ve done the math: outage cost \u0026gt; infrastructure cost You\u0026rsquo;re probably over-engineering if:\nYour app is internal tooling Downtime measured in hours is tolerable You\u0026rsquo;re optimizing for shipping speed over resilience Your team doesn\u0026rsquo;t have the operational maturity to manage this complexity The survey app I built for students? It absolutely doesn\u0026rsquo;t need this. Downtime is annoying, not catastrophic. It\u0026rsquo;s a teaching tool, not a production system.\nBut if you\u0026rsquo;re running live event voting, mobile game leaderboards, or SaaS APIs that customers depend on—then yes, this makes sense.\nThe decision isn\u0026rsquo;t technical. It\u0026rsquo;s about risk tolerance, recovery time objectives, and operational burden.\nOn October 20, a lot of teams learned they\u0026rsquo;d miscalculated that risk. They assumed regional DNS wouldn\u0026rsquo;t fail.\nIt did.\nLearn from that. Design for the failures that actually happen, not the ones that feel unlikely.\n","permalink":"https://lukelittle.com/posts/2026/01/what-october-20-taught-me-about-dynamodb-and-what-it-didnt/","summary":"\u003cp\u003eOn October 20, 2025, DNS resolution failed in AWS us-east-1, and with it, a lot of DynamoDB applications went down.\u003c/p\u003e\n\u003cp\u003eNot because DynamoDB itself failed. The service was running. Data was there. Capacity was fine. But applications couldn\u0026rsquo;t reach it because DNS queries for \u003ccode\u003edynamodb.us-east-1.amazonaws.com\u003c/code\u003e stopped resolving correctly.\u003c/p\u003e\n\u003cp\u003eIf you\u0026rsquo;ve ever wondered what happens when the infrastructure layer beneath your supposedly resilient database becomes unreachable—October 20 was the answer. And it wasn\u0026rsquo;t pretty.\u003c/p\u003e","title":"What October 20 Taught Me About DynamoDB (and What It Didn't)"},{"content":"For this episode of Data Pour, I sat down with Nimish Donde—Head of Cloud Platform and Security Engineering at Truist—at Amélie\u0026rsquo;s French bakery in Charlotte. It\u0026rsquo;s a place that\u0026rsquo;s been part of the city\u0026rsquo;s fabric since 2008, growing from a single 24-hour location in NoDa (that I used to frequent during college) to four locations across Charlotte.\nLike this bakery, Charlotte\u0026rsquo;s tech scene has grown and evolved—and Nimish has been part of that transformation for the past 16 years.\nWhy Nimish Nimish is one of my closest friends and mentors. We both worked on Ally\u0026rsquo;s cloud transformation together—15 years for him at Ally, where he helped navigate a massive financial institution through one of the most significant technology shifts in banking.\nNow he\u0026rsquo;s at Truist, a much larger bank, tackling similar challenges but at an even greater scale. And what I\u0026rsquo;ve always valued about Nimish is his ability to cut through the noise. He doesn\u0026rsquo;t chase technology trends—he focuses on what actually delivers value to customers, builds trust with people, and creates platforms that work.\nThis conversation was long overdue. We grabbed French press coffee (the kind you order at Amélie\u0026rsquo;s in those big presses), settled in, and talked about everything from resiliency and chaos engineering to Charlotte\u0026rsquo;s evolution as a fintech hub.\nCharlotte: the second-largest financial hub in America One of the threads we explored early was Charlotte itself. It\u0026rsquo;s the second-largest financial hub in the United States—home to Bank of America, Wells Fargo, Truist, Ally, and a growing number of financial enterprises setting up operations.\nOver the past two decades, Nimish has watched the city transform from just a banking hub into a tech-financial hub—the birthplace of fintechs, a magnet for cloud talent, and a city where people now want to move for their careers, not just pass through on their way to New York or Silicon Valley.\nCharlotte showed grit after the 2008 financial crisis. The city came back stronger, the tech scene expanded, and the talent pipeline from local universities (UNC Charlotte, Wake Forest, UNC, USC) started feeding directly into these financial institutions.\nEven Amélie\u0026rsquo;s mirrors that resilience. This location shut down during the pandemic but came back in 2023—booming again, just like the city itself.\nCloud platforms as products, not projects Nimish\u0026rsquo;s role at Truist breaks down into three core tenets:\nDeveloper experience and enablement — Lower the barrier to entry so more teams can build on the platform Resiliency and reliability — Ensure customers (internal developers) get maximum value AI and data gravity — Enable the platform to accelerate AI journeys The key insight here: cloud platforms are products, not operational systems.\nToo many organizations treat cloud as infrastructure to manage. Nimish treats it as a product with customers—and those customers are the developers building applications that serve end users. Everything flows from that mindset.\nResiliency is a mindset, not a metric We spent a significant chunk of the conversation on resiliency—something both of us are obsessed with.\nNimish\u0026rsquo;s definition: resiliency is a mindset, not a destination.\nGone are the days when you measured platform quality by uptime percentages. Resiliency today is about the experiences you build for customers. It\u0026rsquo;s about how you respond under stress. It\u0026rsquo;s about building systems that anticipate failure, learn from it, and continuously improve.\nHe referenced the SRE evolution and Google SRE\u0026rsquo;s famous motto: \u0026ldquo;Hope is not a strategy.\u0026rdquo; That philosophy shaped how Nimish approaches resiliency today.\nChaos engineering as a first-class citizen One of the most interesting parts of our discussion was around chaos engineering.\nRegulators—especially in Europe with legislation like DORA—are no longer accepting tabletop exercises as proof of resiliency. They want evidence. They want to see that under stress, your systems still deliver value.\nChaos engineering isn\u0026rsquo;t just a nice-to-have anymore. It\u0026rsquo;s becoming a compliance requirement. It\u0026rsquo;s how you prove resiliency in production. And for banks, that shift is profound.\nNimish believes chaos engineering will become a first-class citizen in the audit process—just like security compliance scores today. You\u0026rsquo;ll need to validate your platforms\u0026rsquo; resilience continuously, not just once a year.\nThe 80/20 rule: technology is the easy 20% A theme that came up repeatedly: technology is the easy part.\nWhether it\u0026rsquo;s on-prem systems or cloud platforms, the binary is only 20% of the challenge. The other 80%? People.\nBringing people along on transformation. Building influence. Convincing risk partners, internal audit, regulators, and developers that this is the right model. Creating feedback loops. Enabling self-service while embedding guardrails.\nNimish has mastered that 80%. And he\u0026rsquo;s clear about it: if you can\u0026rsquo;t influence people, your platform will fail—even if the technology is perfect.\nData gravity and AI: you can\u0026rsquo;t skip the foundation We talked extensively about data gravity—the idea that data needs a central home, properly cataloged and governed, before AI can deliver real value.\nNimish\u0026rsquo;s take: garbage in, garbage out.\nYou can\u0026rsquo;t unlock AI value if your data is in disparate locations, inconsistent, or inaccessible. Executive sponsorship is critical. You need investment in data warehousing, enrichment, cataloging, and lineage before you start training large language models.\nHis advice for enterprises: eat your own dog food.\nIf you\u0026rsquo;re the cloud platform team, build your own data warehouse. Use the platform you\u0026rsquo;re providing to developers. Give them transparency into their usage. Lower the barrier to entry by shining a light on dark spaces where people wouldn\u0026rsquo;t normally pay attention.\nThat transparency accelerates the entire organization\u0026rsquo;s AI journey.\nGuardrails as code: the tiered cake model Nimish shared his tiered cake model for building secure, compliant platforms—an analogy that resonated deeply (especially sitting in a French bakery):\nBase layer: Landing zones with preventative and detective controls (policies, SCPs, IAM boundaries) Second layer: Data visibility—a warehouse that tells the story of what\u0026rsquo;s happening on your platform Third layer: Proactive controls—CI/CD, Infrastructure as Code, reusable templates, inner source Cherry on top: Human interaction—communities of practice, office hours, self-service tools Each layer has guardrails baked in. And when packaged together, you can hand the whole thing to auditors and regulators with full transparency on ingredients, processes, and governance.\nThe result? Developers move fast. Compliance is embedded. And the enterprise is protected.\nAdvice for young engineers: be resilient, adaptable, and comfortable with ambiguity I asked Nimish what he\u0026rsquo;d tell a young engineer aspiring to leadership.\nHis answer: three things matter most:\nResilience — Bounce back. Learn from failure. Keep going. Adaptability — The tech changes constantly. Your ability to evolve matters more than what you know today. Comfort with ambiguity — Leadership isn\u0026rsquo;t clean. You won\u0026rsquo;t always have the answer. Learn to navigate uncertainty. And most importantly: have empathy. Understand people. Use that as an accelerator.\nTechnology skills get you in the door. Soft skills keep you there.\nWhy this conversation mattered Nimish and I have known each other for years, but this was the first time we sat down and recorded a full conversation about the work we\u0026rsquo;ve been doing. It felt less like an interview and more like two friends catching up—talking through transformation, resiliency, Charlotte\u0026rsquo;s evolution, and where the industry is heading.\nI\u0026rsquo;m grateful he made time to do this. If you work in cloud, fintech, or platform engineering—or if you\u0026rsquo;re just curious about what it takes to build systems that scale in highly regulated environments—this episode is worth your time.\nWatch the full episode This blog is just the teaser. If you want the full story—including our takes on agentic AI, the growth of Charlotte\u0026rsquo;s microbrewery scene (which somehow correlates with the tech scene), and Nimish\u0026rsquo;s upcoming hike to Machu Picchu—watch the episode.\nAnd if you\u0026rsquo;re ever in Charlotte, grab a salted caramel brownie at Amélie\u0026rsquo;s. Nimish has been a fan for 15 years. I\u0026rsquo;m partial to the pistachio macaron.\nWatch the full Data Pour episode here:\nhttps://youtu.be/3xuco4R8EHE\n","permalink":"https://lukelittle.com/posts/2026/01/data-pour-with-nimish-donde-resiliency-data-gravity-and-building-cloud-platforms-that-scale/","summary":"\u003cp\u003eFor this episode of Data Pour, I sat down with Nimish Donde—Head of Cloud Platform and Security Engineering at Truist—at Amélie\u0026rsquo;s French bakery in Charlotte. It\u0026rsquo;s a place that\u0026rsquo;s been part of the city\u0026rsquo;s fabric since 2008, growing from a single 24-hour location in NoDa (that I used to frequent during college) to four locations across Charlotte.\u003c/p\u003e\n\u003cp\u003eLike this bakery, Charlotte\u0026rsquo;s tech scene has grown and evolved—and Nimish has been part of that transformation for the past 16 years.\u003c/p\u003e","title":"Data Pour with Nimish Donde: Resiliency, Data Gravity, and Building Cloud Platforms That Scale"},{"content":"At re:Invent 2025, AWS CEO Matt Garman announced something that made me stop and actually pay attention during a keynote—which doesn\u0026rsquo;t happen often.\nHe introduced frontier agents: AI systems that don\u0026rsquo;t just help you write code or answer questions. They work autonomously for hours or days, maintaining context, investigating problems, and making decisions without you holding their hand.\nThree agents got announced:\nKiro - your AI developer AWS Security Agent - your AI security engineer AWS DevOps Agent - your AI operations engineer This isn\u0026rsquo;t another coding assistant that autocompletes your Lambda functions. This is AWS betting that AI agents can handle the kind of multi-hour incident investigations that currently wake up humans at 2 AM.\nLet\u0026rsquo;s be real: AI-assisted incident response isn\u0026rsquo;t new. PagerDuty, Datadog, Dynatrace, and a dozen startups have been doing \u0026ldquo;pull operational data into an LLM and suggest fixes\u0026rdquo; for years. What makes AWS DevOps Agent different is the depth of integration into the AWS control plane and the architectural pattern it represents.\nYou can watch the frontier agents announcement teaser here: https://www.youtube.com/watch?v=fMQfzwS0prQ\nBut I wanted to dig deeper and figure out what this actually means for teams running production systems on AWS.\nWhat makes an agent \u0026ldquo;frontier-class\u0026rdquo; AWS uses the term \u0026ldquo;frontier agent\u0026rdquo; to mean something specific. It\u0026rsquo;s not just GPT-4 with AWS API access.\n1. Autonomous goal-directed behavior\nTraditional AI: \u0026ldquo;Hey ChatGPT, what might cause high Lambda errors?\u0026rdquo;\nFrontier agent: \u0026ldquo;Investigate this Lambda error spike\u0026rdquo; → agent figures out how\nYou give it an objective, it decomposes the problem, forms hypotheses, collects evidence, and executes—without asking you for step-by-step guidance.\n2. Multi-agent coordination\nDevOps Agent doesn\u0026rsquo;t work alone. When investigating an incident, it spawns specialized sub-agents—one analyzing logs, another reconstructing the deployment timeline, a third mapping topology. These agents run concurrently, investigating multiple hypotheses simultaneously and coordinating across AWS accounts. It\u0026rsquo;s less \u0026ldquo;AI assistant\u0026rdquo; and more \u0026ldquo;AI team.\u0026rdquo;\n3. Long-running independent operation\nHere\u0026rsquo;s the paradigm shift: it works for hours without constant human intervention.\nTraditional AI forgets everything when you close the chat. Frontier agents maintain persistent context, remember your infrastructure, learn from past incidents, and pick up where they left off after restarts.\nWhen your Lambda error alarm goes off at 2 AM, DevOps Agent can investigate for 30 minutes, form a hypothesis, collect evidence, and have a diagnosis ready by the time you wake up and check Slack.\nHow it actually works DevOps Agent integrates with your existing monitoring tools—it doesn\u0026rsquo;t replace them.\nOn the observability side, DevOps Agent integrates natively with CloudWatch and can pull data from Datadog, Dynatrace, New Relic, and Splunk. If you\u0026rsquo;re using custom monitoring tools, you can build integrations via Model Context Protocol (MCP) servers—AWS\u0026rsquo;s standard for extending agent capabilities.\nFor incident coordination, there\u0026rsquo;s built-in support for ServiceNow and PagerDuty, plus Slack for real-time updates. Pretty much any tool with webhooks can be integrated into the workflow.\nDevOps Agent can be triggered three ways: automatically when a CloudWatch alarm fires (fully autonomous response), manually through the web UI when you want to investigate something specific, or on a schedule for proactive analysis—think nightly scans looking for anomalies before they become incidents.\nWhen an alert fires—say, Lambda errors spike at 2 AM—here\u0026rsquo;s what happens:\ngraph TB Alert[CloudWatch Alarm Fires] --\u0026gt; Orchestrator[Investigation Orchestrator] Orchestrator --\u0026gt; Topo[Topology Sub-Agent\u0026lt;br/\u0026gt;Maps dependencies] Orchestrator --\u0026gt; Telem[Telemetry Sub-Agent\u0026lt;br/\u0026gt;Analyzes metrics/logs] Orchestrator --\u0026gt; Deploy[Deployment Sub-Agent\u0026lt;br/\u0026gt;Checks recent changes] Topo --\u0026gt; RCA[Root Cause Analysis] Telem --\u0026gt; RCA Deploy --\u0026gt; RCA RCA --\u0026gt; Slack[Post to Slack #incidents] RCA --\u0026gt; Ticket[Create ServiceNow ticket] The clever part is the application topology map. DevOps Agent builds and maintains an intelligent map of your entire system—which Lambda functions call which APIs, which services depend on which databases, when each component was last deployed and by whom. It tracks cross-account dependencies and even external dependencies like third-party APIs, SaaS integrations, and CDNs.\nWhen an incident happens, this topology becomes invaluable. The agent can immediately identify blast radius (what\u0026rsquo;s affected by this outage?), trace dependency chains (if the API is down, what upstream services caused it? what downstream services are impacted?), and correlate timing (there was a deploy 15 minutes ago—is that when this started?).\nThe investigation loop Once triggered, DevOps Agent enters an iterative loop:\nGenerate hypotheses based on alert type, topology, recent changes Collect evidence by querying logs, metrics, traces, configs Correlate patterns across time, services, accounts Assess confidence in each hypothesis Recommend mitigation or continue investigating Learn from outcome to improve future investigations This keeps going until it reaches high confidence in the root cause or exhausts reasonable paths.\nWhat separates this from dumb rule-based systems: it doesn\u0026rsquo;t just pattern-match. It reasons about your infrastructure.\nAgent Spaces and IAM permission boundaries Everything starts with an Agent Space—the workspace where the agent operates and the IAM permission boundary defining what it can access.\nYou can structure Agent Spaces multiple ways:\nPer-application: One space per critical app Per-team: One space per on-call team Centralized: One space in monitoring account observing everything Here\u0026rsquo;s what makes this not just \u0026ldquo;magic AI with root access\u0026rdquo;: DevOps Agent uses explicit, auditable IAM trust relationships.\nThe Agent Space role trust policy:\nPrincipal: aidevops.amazonaws.com (not some opaque service) Conditions: SourceAccount and SourceArn bound to your specific AgentSpace Permissions: Standard IAM policies you control You can audit exactly what DevOps Agent accessed, when, and why. It\u0026rsquo;s not a black box.\nMulti-account setup (the real production pattern) Production incidents rarely happen in a single AWS account. You have workload accounts, shared services accounts, security accounts, monitoring accounts.\nDevOps Agent supports this natively via External Account Associations:\ngraph TD subgraph Monitor[\u0026#34;Monitoring Account\u0026#34;] AgentSpace[DevOps Agent Space] end subgraph Workload1[\u0026#34;Workload Account 1\u0026#34;] Role1[IAM Role\u0026lt;br/\u0026gt;ReadOnly + Logs] end subgraph Workload2[\u0026#34;Workload Account 2\u0026#34;] Role2[IAM Role\u0026lt;br/\u0026gt;ReadOnly + Logs] end AgentSpace --\u0026gt;|Cross-account trust| Role1 AgentSpace --\u0026gt;|Cross-account trust| Role2 Create your Agent Space in a central monitoring account, associate it with workload accounts via cross-account IAM roles, and let it investigate incidents spanning account boundaries.\nThis is how you do AWS at scale.\nDeploying it (Terraform example) AWS provides Terraform resources (aws_devopsagent_agentspace, aws_devopsagent_association) and CDK constructs.\nBasic Terraform setup:\nresource \u0026#34;aws_devopsagent_agentspace\u0026#34; \u0026#34;main\u0026#34; { name = \u0026#34;production-monitoring\u0026#34; agent_role { create_role = true role_name = \u0026#34;DevOpsAgentSpaceRole\u0026#34; } enable_web_app = true # Optional UI } resource \u0026#34;aws_devopsagent_association\u0026#34; \u0026#34;workload\u0026#34; { agent_space_id = aws_devopsagent_agentspace.main.id external_account { account_id = \u0026#34;987654321098\u0026#34; role_arn = \u0026#34;arn:aws:iam::987654321098:role/DevOpsAgentWorkloadRole\u0026#34; } } In each workload account, create a role that trusts your Agent Space:\nresource \u0026#34;aws_iam_role\u0026#34; \u0026#34;devops_agent_workload\u0026#34; { name = \u0026#34;DevOpsAgentWorkloadRole\u0026#34; assume_role_policy = jsonencode({ Principal = { AWS = \u0026#34;arn:aws:iam::123456789012:role/DevOpsAgentSpaceRole\u0026#34; } }) } resource \u0026#34;aws_iam_role_policy_attachment\u0026#34; \u0026#34;read\u0026#34; { role = aws_iam_role.devops_agent_workload.name policy_arn = \u0026#34;arn:aws:iam::aws:policy/ReadOnlyAccess\u0026#34; } Connect your monitoring tools (Data API keys, GitHub tokens) through the AWS Console.\nImportant: AWS explicitly says Terraform resources may change before GA. Pin your provider versions.\nTesting it AWS provides test scenarios. I recommend running these before connecting production systems.\nTest 1: Lambda error investigation\nDeploy a Lambda that intentionally throws errors:\nimport random def lambda_handler(event, context): errors = [ \u0026#34;Simulated database timeout\u0026#34;, \u0026#34;Test API rate limit\u0026#34;, \u0026#34;Validation error\u0026#34; ] raise Exception(f\u0026#34;Test: {random.choice(errors)}\u0026#34;) Create a CloudWatch alarm, trigger it, watch DevOps Agent:\nDetect the spike Analyze logs Check deployment timeline Identify root cause Recommend fixes Test 2: EC2 CPU spike\nDeploy an EC2 instance, run a CPU stress test, trigger an alarm, watch it correlate with recent changes and recommend auto-scaling.\nWhat\u0026rsquo;s not ready yet (the honest limitations) 1. us-east-1 only DevOps Agent is currently only available in us-east-1.\nIf you have data residency requirements (GDPR, finance, healthcare), this is a blocker. Cross-region investigations require routing everything through us-east-1.\nMitigation: Deploy Agent Space in us-east-1, use cross-account associations to observe other regions. AWS will probably expand regions post-GA.\n2. Investigation vs action It\u0026rsquo;s unclear whether DevOps Agent can execute remediation or just recommend it.\nThe documentation emphasizes \u0026ldquo;investigations,\u0026rdquo; \u0026ldquo;recommendations,\u0026rdquo; \u0026ldquo;mitigation suggestions\u0026rdquo;—not \u0026ldquo;auto-rollback\u0026rdquo; or \u0026ldquo;auto-scale.\u0026rdquo;\nMy read: GA will probably support both:\nInvestigation-only mode (default): analyze → recommend → human executes Action mode (opt-in): execute pre-approved actions within guardrails For regulated industries, you\u0026rsquo;ll live in investigation-only mode. For fast-moving startups, action mode might be tempting.\n3. Integration maturity Integrations exist for CloudWatch, Datadog, Dynatrace, New Relic, Splunk, GitHub, GitLab, ServiceNow, PagerDuty—but they\u0026rsquo;re first-generation.\nMissing:\nOpenTelemetry native support ArgoCD, Flux, Spinnaker Opsgenie, Incident.io AppDynamics, Elastic APM Good news: Model Context Protocol (MCP) support means you can build custom integrations.\n4. Learning curve DevOps Agent builds its topology map over time. Early investigations might be less accurate.\nMitigation:\nRun test investigations to let it learn Tag resources consistently Document dependencies explicitly Should you actually use this? Use it if:\nYou\u0026rsquo;re heavily invested in AWS Your team is drowning in incident response toil You have operational maturity (monitoring, tagging, CI/CD) You\u0026rsquo;re comfortable with preview-phase tech Wait if:\nYou need multi-region support now You require deterministic pricing Your incident response is already highly optimized You need production SLAs (preview = no SLAs) Key insight: DevOps Agent amplifies good practices and exposes bad ones. If your infrastructure is poorly tagged, deployments aren\u0026rsquo;t tracked, and metrics are scattered, it\u0026rsquo;ll struggle. But if you have solid foundations, it can be transformative.\nMy honest take This is the future of operations. Not because AI replaces engineers, but because it handles undifferentiated heavy lifting.\nThe question isn\u0026rsquo;t whether agentic operations are coming—they\u0026rsquo;re here. The question is whether you\u0026rsquo;ll be ready when GA drops.\nIf you\u0026rsquo;re experimenting with this or have questions, reach out on LinkedIn. The technology is moving fast, and we\u0026rsquo;re all figuring it out together.\nResources:\nAWS DevOps Agent User Guide Terraform Sample Repo AWS Frontier Agents Overview ","permalink":"https://lukelittle.com/posts/2026/01/your-ai-on-call-engineer-inside-aws-devops-agent/","summary":"\u003cp\u003eAt re:Invent 2025, AWS CEO Matt Garman announced something that made me stop and actually pay attention during a keynote—which doesn\u0026rsquo;t happen often.\u003c/p\u003e\n\u003cp\u003eHe introduced \u003cstrong\u003efrontier agents\u003c/strong\u003e: AI systems that don\u0026rsquo;t just help you write code or answer questions. They work autonomously for hours or days, maintaining context, investigating problems, and making decisions without you holding their hand.\u003c/p\u003e\n\u003cp\u003eThree agents got announced:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eKiro\u003c/strong\u003e - your AI developer\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAWS Security Agent\u003c/strong\u003e - your AI security engineer\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAWS DevOps Agent\u003c/strong\u003e - your AI operations engineer\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThis isn\u0026rsquo;t another coding assistant that autocompletes your Lambda functions. This is AWS betting that AI agents can handle the kind of multi-hour incident investigations that currently wake up humans at 2 AM.\u003c/p\u003e","title":"Your AI On-Call Engineer: Inside AWS DevOps Agent"},{"content":"For this episode, I visited the PORTAL building at UNC Charlotte to sit down with Dr. Mohamed Shehab—a professor whose mobile development course keeps showing up in conversations with students as one of the most impactful experiences of their degree.\nI\u0026rsquo;ve heard it repeatedly from students I\u0026rsquo;ve hired: \u0026ldquo;Dr. Shehab\u0026rsquo;s course changed how I think about building software.\u0026rdquo; When you hear that kind of feedback consistently, you have to ask what he\u0026rsquo;s doing differently.\nWhy students remember this course I originally connected with Dr. Shehab after noticing the pattern. Current students, recent grads, people years into their careers—they all talked about his mobile dev class the way you talk about a course that actually stuck.\nThe answer is pretty straightforward: it\u0026rsquo;s hands-on, industry-aligned, and constantly evolving.\nDr. Shehab started teaching mobile development in 2008, right when the iPhone App Store launched. He got lucky with timing, but what kept the course relevant was the approach: flip the lecture material to video, use class time for hands-on work, teach students to ship working apps—not just pass exams.\nThe tech changes constantly. Updates every few months. So the course has to change with it. Students learn how to build, deploy, and debug real applications. They encounter the same decision-making tradeoffs developers face in production.\nThat\u0026rsquo;s the alignment. Students walk out with job-ready experience, not just academic knowledge.\nEmbracing AI (and changing how you evaluate) We spent a lot of time talking about AI in the classroom. Dr. Shehab\u0026rsquo;s take is pragmatic: the technology is here, students are using it, developers are using it—you have to adapt.\nBut it creates real challenges for educators. How do you evaluate student understanding when they can generate code with a chatbot?\nDr. Shehab\u0026rsquo;s approach: change how you evaluate. Don\u0026rsquo;t just check if the app runs—have students demo it and explain how it works. Ask questions. Make them walk through their code. If they used GPT to build it but can understand and articulate the solution, that\u0026rsquo;s acceptable. If they can\u0026rsquo;t explain it, they get half credit.\nHe\u0026rsquo;s also experimenting with GitHub Copilot in the development environment—teaching students to use AI as an assistive tool, not a replacement for understanding.\nThe goal isn\u0026rsquo;t to block AI. It\u0026rsquo;s to teach students how to use it effectively while still building foundational skills.\nBuild, build, build I asked what advice he gives students trying to break into the job market—especially when entry-level roles feel harder to land and AI is changing expectations.\nHis answer was clear: build things.\nDon\u0026rsquo;t just do assignments. Build real projects. Put them on GitHub. Deploy them. Show your work.\nThe biggest challenge right now is for early talent. If you\u0026rsquo;re graduating and want to stand out, you need a strong portfolio. You need to show up to employers with something you\u0026rsquo;ve built—not expecting them to hold your hand and train you from scratch.\nCommunication matters too. AI can do a lot of things, but it can\u0026rsquo;t be you. Being a good collaborator, explaining things clearly, presenting confidently—those skills still matter.\nAnd get involved outside the classroom. Go to meetups. Join student orgs. Use the maker spaces. Network. Ask questions at recruiting events—even if you think the question is stupid. Have presence.\nDr. Shehab\u0026rsquo;s motto: build, build, build. No hand-waving. Features built = points. No features = no points.\nThe maker space movement One of the more interesting parts of our conversation was the emphasis on maker spaces. UNC Charlotte has invested heavily—3D printers, CNC machines, fabrication labs. Students can access them for free.\nDr. Shehab\u0026rsquo;s point: the university isn\u0026rsquo;t going to teach you everything in the classroom. A lot of learning happens outside—in maker spaces, student orgs, side projects. Open yourself to other experiences. Print something. Build something. See what\u0026rsquo;s possible.\nWhen companies interview students and see that breadth of experience—someone who\u0026rsquo;s not just focused on one narrow skill but has built, experimented, and solved real problems—that opens doors.\nBuilding stronger industry partnerships We also talked about what industry can do to be better partners with academia.\nMore collaboration. More joint projects. More events on campus exposing students to the ecosystem. More feedback to faculty on what skills actually matter in the market.\nDr. Shehab mentioned that UNC Charlotte is exploring focused talent pipelines—where companies have input on curriculum and certificate programs. Containerization certificates. Visualization certificates. AI certificates. Programs designed not just for traditional students but for employees or potential employees.\nThe challenge is cultural. A lot of academics are traditional: teach a class, do research, go home. But exposing faculty to industry—and integrating those relationships—can boost relevance for everyone.\nWhy I wanted him on the show I\u0026rsquo;m incredibly grateful to Dr. Shehab for making time to be part of The Data Pour. Seeing the level of impact he\u0026rsquo;s had on students is what motivated me to get more involved at UNC Charlotte—to help create stronger pathways between academia and industry.\nI\u0026rsquo;m excited to continue working together in the future.\nWatch the episode If you want the full conversation—including the parts on the PORTAL incubator, startup culture, and what Dr. Shehab sees as the future of CS education—watch the YouTube episode here:\nhttps://www.youtube.com/watch?v=vZOze-a_JTU\u0026amp;t=564s\n","permalink":"https://lukelittle.com/posts/2026/01/data-pour-with-dr.-mohamed-shehab-build-build-build/","summary":"\u003cp\u003eFor this episode, I visited the PORTAL building at UNC Charlotte to sit down with Dr. Mohamed Shehab—a professor whose mobile development course keeps showing up in conversations with students as one of the most impactful experiences of their degree.\u003c/p\u003e\n\u003cp\u003eI\u0026rsquo;ve heard it repeatedly from students I\u0026rsquo;ve hired: \u0026ldquo;Dr. Shehab\u0026rsquo;s course changed how I think about building software.\u0026rdquo; When you hear that kind of feedback consistently, you have to ask what he\u0026rsquo;s doing differently.\u003c/p\u003e","title":"Data Pour with Dr. Mohamed Shehab: Build, Build, Build"},{"content":"Back in November, I was preparing for the Cracking the Cloud presentation at UNC Charlotte. I needed a way to explain how the cloud fundamentally changed what\u0026rsquo;s possible on the internet—not through abstract concepts, but through something students could immediately relate to.\nThat\u0026rsquo;s when I remembered Thomas Game Docs.\nIf you\u0026rsquo;ve never heard of her: she\u0026rsquo;s a YouTuber who makes incredibly well-produced video essays about video games. And she sometimes runs surveys asking her audience things like \u0026ldquo;Who\u0026rsquo;s the LEAST popular Pokémon?\u0026rdquo; or \u0026ldquo;Who\u0026rsquo;s the LEAST popular Animal Crossing villager?\u0026rdquo;\nThese aren\u0026rsquo;t small surveys. They get millions of responses.\nI helped with the backend for the Pokémon survey—a Flask app on Heroku. But the technology stack wasn\u0026rsquo;t the interesting part. What mattered was that hosting something like this doesn\u0026rsquo;t require infrastructure expertise anymore.\nShe didn\u0026rsquo;t need to buy servers, configure databases, or hire a DevOps team.\nTwenty years ago, hosting a survey that could handle 50,000 votes meant:\nBuy or rent physical servers Set up database infrastructure Configure load balancers Plan for capacity (and hope you got it right) Deal with outages, scaling issues, and hardware failures All of that would cost thousands of dollars—and that\u0026rsquo;s before you wrote a single line of code.\nToday? You build it with Lambda, API Gateway, and DynamoDB. You deploy it with Terraform. And unless traffic gets truly ridiculous, it costs you basically nothing.\nThe cloud didn\u0026rsquo;t just make infrastructure cheaper. It made building things accessible.\nThat\u0026rsquo;s what I wanted students to understand. Not that AWS has a lot of services. But that those services remove the barriers that used to keep people from building.\nThe demo: a survey students could actually participate in To drive the point home, I didn\u0026rsquo;t just talk about Thomas Game Docs surveys.\nI had the students take one.\nAt the start of the presentation, I pulled up a simple survey asking about their exposure to AWS:\nHave you used AWS before? Have you deployed something to the cloud? They voted. They saw the results update in real-time. And then I showed them exactly how it worked—with no servers running, no databases to manage, and no ongoing costs to worry about.\nThat survey? It\u0026rsquo;s the same repo I\u0026rsquo;m writing about now: cracking-the-cloud.\nThe architecture (small, but real) Here\u0026rsquo;s the final shape of the system:\nS3 hosts the static frontend (HTML, CSS, JS) CloudFront sits in front for HTTPS, caching, and global delivery API Gateway exposes a REST API Lambda handles business logic (vote, results, reset) DynamoDB stores votes IAM wires permissions together Terraform defines everything No servers. No containers. No databases to patch. No stateful nonsense.\nJust managed services doing exactly what they\u0026rsquo;re good at.\nThe request flow looks like this:\nUser clicks a button → JavaScript calls the API → API Gateway invokes Lambda → Lambda writes to DynamoDB → Response goes back to the browser.\nSimple. Explicit. Observable.\nWhy static frontend + API (on purpose) I didn\u0026rsquo;t use React.\nI didn\u0026rsquo;t use Next.js.\nI didn\u0026rsquo;t use server-side rendering.\nNot because those tools are bad—they\u0026rsquo;re not. But because I wanted you to see what\u0026rsquo;s actually happening.\nWhen you open the frontend code, you can immediately see:\nwhere the API URL lives how a POST request is formed what the response looks like how the browser handles the data No build steps. No transpilation. No abstractions hiding what\u0026rsquo;s really going on.\nOnce you understand how a browser talks to an API using vanilla JavaScript, then you can add React, TypeScript, and all the modern tooling. But you\u0026rsquo;ll know what those tools are doing for you—not just that they work.\nThe frontend\u0026rsquo;s job here is to show you the fundamentals, not teach you the latest framework.\nThe Lambdas (three, on purpose) There are three Lambda functions:\nVote (backend/vote.py) – Processes vote submissions Results (backend/results.py) – Retrieves vote counts Reset (backend/reset.py) – Clears all data Could this be one Lambda with a switch statement?\nAbsolutely.\nDid I do that?\nAbsolutely not.\nEach function has:\none responsibility one API route one IAM policy This lets students see how permissions map to behavior.\nThe vote function can write, but not delete The results function can read, but not write The reset function can delete, but nothing else You don\u0026rsquo;t need a lecture on least privilege when the code makes it obvious.\nHow voting actually works When a student clicks a vote button, here\u0026rsquo;s the journey that request takes through the serverless stack:\nsequenceDiagram participant User as 👤 User Browser participant S3 as 🪣 S3 + CloudFront participant APIG as 🌐 API Gateway participant Lambda as ⚡ vote.py participant DDB as 🗄️ DynamoDB User-\u0026gt;\u0026gt;S3: GET /vote.html S3--\u0026gt;\u0026gt;User: HTML + JavaScript Note over User: User clicks vote button\u0026lt;br/\u0026gt;sessionId generated (UUID)\u0026lt;br/\u0026gt;stored in sessionStorage User-\u0026gt;\u0026gt;APIG: POST /vote\u0026lt;br/\u0026gt;{ \u0026#34;vote\u0026#34;: \u0026#34;aws\u0026#34;, \u0026#34;sessionId\u0026#34;: \u0026#34;abc123\u0026#34; } APIG-\u0026gt;\u0026gt;Lambda: Invoke vote function Note over Lambda: Validate sessionId exists\u0026lt;br/\u0026gt;Validate vote in [\u0026#39;no\u0026#39;, \u0026#39;aws\u0026#39;, \u0026#39;other\u0026#39;] Lambda-\u0026gt;\u0026gt;DDB: PutItem\u0026lt;br/\u0026gt;{ id: \u0026#34;abc123\u0026#34;, vote: \u0026#34;aws\u0026#34; } Note over DDB: Overwrites if sessionId\u0026lt;br/\u0026gt;already voted\u0026lt;br/\u0026gt;(allows vote changes) DDB--\u0026gt;\u0026gt;Lambda: Success Lambda--\u0026gt;\u0026gt;APIG: 200 OK\u0026lt;br/\u0026gt;{ \u0026#34;message\u0026#34;: \u0026#34;Vote recorded\u0026#34; } APIG--\u0026gt;\u0026gt;User: Response Note over User: JavaScript displays\u0026lt;br/\u0026gt;\u0026#34;Vote recorded!\u0026#34; message The critical piece here is the sessionId. It\u0026rsquo;s a random UUID generated in the browser and stored in sessionStorage—which means it persists for the current tab but disappears when you close the browser.\nThis gives us:\nOne vote per browser session – you can\u0026rsquo;t spam-click the vote button Vote changes allowed – if you vote \u0026ldquo;No experience\u0026rdquo; and change your mind, the second vote overwrites the first (DynamoDB\u0026rsquo;s PutItem does this automatically) Privacy by default – no accounts, no tracking, no persistent identifiers Simple anti-spam – good enough for a teaching demo Could someone bypass this by opening incognito windows? Yes. Is that fine for a teaching app? Also yes. The point is showing how to prevent duplicate votes, not building a production election system.\nHere\u0026rsquo;s what vote.py actually looks like:\ndef handler(event, context): # Parse the incoming request body = json.loads(event.get(\u0026#39;body\u0026#39;, \u0026#39;{}\u0026#39;)) vote_option = body.get(\u0026#39;vote\u0026#39;) session_id = body.get(\u0026#39;sessionId\u0026#39;) # Validate inputs if not session_id: return { \u0026#39;statusCode\u0026#39;: 400, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;message\u0026#39;: \u0026#39;Session ID is required\u0026#39;}) } if vote_option not in [\u0026#39;no\u0026#39;, \u0026#39;aws\u0026#39;, \u0026#39;other\u0026#39;]: return { \u0026#39;statusCode\u0026#39;: 400, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;message\u0026#39;: \u0026#39;Invalid vote option\u0026#39;}) } # Store the vote table.put_item(Item={ \u0026#39;id\u0026#39;: session_id, \u0026#39;vote\u0026#39;: vote_option }) return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;headers\u0026#39;: {\u0026#39;Access-Control-Allow-Origin\u0026#39;: \u0026#39;*\u0026#39;}, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;message\u0026#39;: \u0026#39;Vote recorded successfully\u0026#39;}) } Twenty lines of code. No ORM. No database migrations. No connection pooling. Just write to DynamoDB and return a response.\nHow results actually work The results page is where students first encounter the concept of scanning a database:\nsequenceDiagram participant User as 👤 User Browser participant S3 as 🪣 S3 + CloudFront participant APIG as 🌐 API Gateway participant Lambda as ⚡ results.py participant DDB as 🗄️ DynamoDB User-\u0026gt;\u0026gt;S3: GET /results.html S3--\u0026gt;\u0026gt;User: HTML + JavaScript + Chart.js Note over User: Page loads\u0026lt;br/\u0026gt;JavaScript calls API User-\u0026gt;\u0026gt;APIG: GET /results APIG-\u0026gt;\u0026gt;Lambda: Invoke results function Lambda-\u0026gt;\u0026gt;DDB: Scan table\u0026lt;br/\u0026gt;ProjectionExpression=\u0026#39;vote\u0026#39; Note over DDB: Returns all vote values\u0026lt;br/\u0026gt;[\u0026#39;aws\u0026#39;, \u0026#39;no\u0026#39;, \u0026#39;aws\u0026#39;, \u0026#39;other\u0026#39;, ...] DDB--\u0026gt;\u0026gt;Lambda: Page 1 of results Note over Lambda: Check for LastEvaluatedKey\u0026lt;br/\u0026gt;(pagination if table \u0026gt; 1MB) loop While LastEvaluatedKey exists Lambda-\u0026gt;\u0026gt;DDB: Scan with ExclusiveStartKey DDB--\u0026gt;\u0026gt;Lambda: Next page of results end Note over Lambda: Count votes using Counter\u0026lt;br/\u0026gt;{ \u0026#39;no\u0026#39;: 15, \u0026#39;aws\u0026#39;: 42, \u0026#39;other\u0026#39;: 8 } Lambda--\u0026gt;\u0026gt;APIG: 200 OK\u0026lt;br/\u0026gt;{ \u0026#34;no\u0026#34;: 15, \u0026#34;aws\u0026#34;: 42, \u0026#34;other\u0026#34;: 8 } APIG--\u0026gt;\u0026gt;User: Response Note over User: Chart.js renders\u0026lt;br/\u0026gt;vote counts as bar chart The interesting part here is pagination. DynamoDB\u0026rsquo;s Scan operation returns a maximum of 1MB of data per request. If your table is larger than that, you get a LastEvaluatedKey in the response, which you use to fetch the next page.\nHere\u0026rsquo;s what that looks like in code:\ndef handler(event, context): # First scan response = table.scan(ProjectionExpression=\u0026#39;vote\u0026#39;) items = response.get(\u0026#39;Items\u0026#39;, []) # Keep scanning if there\u0026#39;s more data while \u0026#39;LastEvaluatedKey\u0026#39; in response: response = table.scan( ProjectionExpression=\u0026#39;vote\u0026#39;, ExclusiveStartKey=response[\u0026#39;LastEvaluatedKey\u0026#39;] ) items.extend(response.get(\u0026#39;Items\u0026#39;, [])) # Extract just the vote values votes = [item[\u0026#39;vote\u0026#39;] for item in items] # Count them vote_counts = Counter(votes) # Return with defaults for zero-vote options return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;headers\u0026#39;: {\u0026#39;Access-Control-Allow-Origin\u0026#39;: \u0026#39;*\u0026#39;}, \u0026#39;body\u0026#39;: json.dumps({ \u0026#39;no\u0026#39;: vote_counts.get(\u0026#39;no\u0026#39;, 0), \u0026#39;aws\u0026#39;: vote_counts.get(\u0026#39;aws\u0026#39;, 0), \u0026#39;other\u0026#39;: vote_counts.get(\u0026#39;other\u0026#39;, 0) }) } Students immediately see three concepts:\nScanning costs – you\u0026rsquo;re reading the entire table, which is fine for 100 votes but would be expensive for 10 million Pagination handling – real-world data doesn\u0026rsquo;t fit in one response Aggregation happens in code – DynamoDB doesn\u0026rsquo;t have COUNT(*) GROUP BY vote, so you pull the data and count it yourself This naturally leads to questions like \u0026ldquo;how would you make this more efficient?\u0026rdquo; which is exactly where you want students\u0026rsquo; brains to go.\nHow reset actually works The reset function is the most dangerous one in the app—and also the most instructive:\nsequenceDiagram participant User as 👤 User Browser participant S3 as 🪣 S3 + CloudFront participant APIG as 🌐 API Gateway participant Lambda as ⚡ reset.py participant DDB as 🗄️ DynamoDB User-\u0026gt;\u0026gt;S3: GET /reset.html S3--\u0026gt;\u0026gt;User: HTML + JavaScript Note over User: User clicks\u0026lt;br/\u0026gt;\u0026#34;Reset All Votes\u0026#34; button\u0026lt;br/\u0026gt;(⚠️ Destructive operation) User-\u0026gt;\u0026gt;APIG: POST /reset APIG-\u0026gt;\u0026gt;Lambda: Invoke reset function Lambda-\u0026gt;\u0026gt;DDB: Scan table\u0026lt;br/\u0026gt;ProjectionExpression=\u0026#39;id\u0026#39; Note over DDB: Only return IDs\u0026lt;br/\u0026gt;(need keys to delete) DDB--\u0026gt;\u0026gt;Lambda: All item IDs loop While LastEvaluatedKey exists Lambda-\u0026gt;\u0026gt;DDB: Scan with ExclusiveStartKey DDB--\u0026gt;\u0026gt;Lambda: More IDs end Note over Lambda: Batch delete in groups of 25\u0026lt;br/\u0026gt;(DynamoDB batch write limit) loop For each batch of 25 items Lambda-\u0026gt;\u0026gt;DDB: BatchWriteItem\u0026lt;br/\u0026gt;Delete items DDB--\u0026gt;\u0026gt;Lambda: Batch delete success end Note over Lambda: All items deleted\u0026lt;br/\u0026gt;Table is now empty Lambda--\u0026gt;\u0026gt;APIG: 200 OK\u0026lt;br/\u0026gt;{ \u0026#34;message\u0026#34;: \u0026#34;Survey reset\u0026#34; } APIG--\u0026gt;\u0026gt;User: Response Note over User: \u0026#34;All votes deleted!\u0026#34; This introduces batch operations:\ndef handler(event, context): # Scan for all IDs (we only need keys to delete) scan_response = table.scan(ProjectionExpression=\u0026#39;id\u0026#39;) items = scan_response.get(\u0026#39;Items\u0026#39;, []) # Handle pagination while \u0026#39;LastEvaluatedKey\u0026#39; in scan_response: scan_response = table.scan( ProjectionExpression=\u0026#39;id\u0026#39;, ExclusiveStartKey=scan_response[\u0026#39;LastEvaluatedKey\u0026#39;] ) items.extend(scan_response.get(\u0026#39;Items\u0026#39;, [])) # Delete all items using batch writer if items: with table.batch_writer() as batch: for item in items: batch.delete_item(Key={\u0026#39;id\u0026#39;: item[\u0026#39;id\u0026#39;]}) return { \u0026#39;statusCode\u0026#39;: 200, \u0026#39;headers\u0026#39;: {\u0026#39;Access-Control-Allow-Origin\u0026#39;: \u0026#39;*\u0026#39;}, \u0026#39;body\u0026#39;: json.dumps({\u0026#39;message\u0026#39;: \u0026#39;Survey reset successfully\u0026#39;}) } The batch_writer() context manager is doing a lot of hidden work:\nGroups deletes into batches of 25 (DynamoDB\u0026rsquo;s limit) Automatically retries failed operations Handles throttling gracefully Only commits when the context exits Students don\u0026rsquo;t need to know all of that on day one, but they can see that deleting 100 items doesn\u0026rsquo;t require 100 API calls.\nImportant note: In a real app, you\u0026rsquo;d absolutely add authentication and authorization here. This function is intentionally unprotected for teaching purposes—it demonstrates the mechanics of batch operations without the complexity of auth flows.\nDynamoDB (intentionally unsexy) The DynamoDB table is boring by design.\nPartition key (id) Simple attributes (vote) No GSIs No streams No TTL magic Why?\nBecause the lesson isn\u0026rsquo;t \u0026ldquo;DynamoDB is infinite and weird.\u0026rdquo;\nThe lesson is:\nYou can persist state without running a database.\nOnce students are comfortable, then you add:\nsecondary indexes conditional writes access patterns cost modeling But not on day one.\nTerraform as the real curriculum Here\u0026rsquo;s the quiet truth: The Terraform is the most important part of this project.\nStudents don\u0026rsquo;t learn AWS by clicking around the console. They learn AWS by reading infrastructure definitions and realizing: \u0026ldquo;Oh—that\u0026rsquo;s what connects to that.\u0026rdquo;\nThis repo forces them to see how API Gateway connects to Lambda (aws_api_gateway_integration), how Lambda permissions work (aws_lambda_permission), how CloudFront talks to S3 (origin_access_identity), how outputs become frontend configuration (cloudfront_domain).\nThey can delete everything and recreate it in minutes. That alone teaches more than most cloud courses.\nHow does this compare to the Heroku version? Remember the Flask app I built for Thomas Game Docs\u0026rsquo; Pokémon survey?\nI monitored that deployment closely. Every time she dropped an announcement on social media—Twitter, YouTube community posts, etc.—I watched the app response time blow up. We were constantly aware that we were one viral tweet away from needing to manually scale the Heroku dyno or upgrade the database.\nWith this serverless version? There wouldn\u0026rsquo;t have been a hiccup.\nLambda would have spun up as many concurrent executions as needed. API Gateway would have handled the traffic without breaking a sweat. DynamoDB would have throttled gracefully and auto-scaled. CloudFront would have cached the static assets globally.\nNo monitoring dashboards.\nNo capacity planning.\nNo \u0026ldquo;should we upgrade now or wait?\u0026rdquo; decisions.\nNo watching metrics at 2 AM when a post goes viral.\nThe infrastructure would have scaled to meet demand and then scaled back down when traffic dropped. And the bill would have stayed under $5 for the entire campaign.\nThat\u0026rsquo;s the difference between \u0026ldquo;serverless\u0026rdquo; and \u0026ldquo;server-you-manage-less.\u0026rdquo;\nCosts (because someone always asks) S3: pennies CloudFront: free tier Lambda: free tier API Gateway: free tier DynamoDB: free tier Total monthly cost for light usage: effectively $0.\nWhich matters, because students shouldn\u0026rsquo;t need a credit card panic attack to learn cloud fundamentals.\nIf you want to fork it The entire project is open source and designed to be broken, modified, and rebuilt:\ngithub.com/lukelittle/cracking-the-cloud\nHow this project can be used to teach cloud The beauty of this baseline is that every extension becomes a teaching moment. The answer to nearly every \u0026ldquo;Can I add\u0026hellip;?\u0026rdquo; question is: yes.\n\u0026ldquo;Can I add another survey question?\u0026rdquo; Yes. Modify the DynamoDB schema and update the frontend. Students learn about schema evolution and backwards compatibility.\n\u0026ldquo;Can I add authentication?\u0026rdquo; Yes. Add Cognito, modify the Lambda to verify JWT tokens, update the frontend to handle login flows. Students learn about identity providers, token validation, and authorization.\n\u0026ldquo;Can I track who voted when?\u0026rdquo; Yes. Add timestamps to DynamoDB items, maybe stream changes to S3 for analytics. Students learn about audit trails and data retention.\n\u0026ldquo;Can I swap DynamoDB for RDS?\u0026rdquo; Yes. But now you need VPCs, security groups, connection pooling, and Lambda cold start considerations. Students learn why DynamoDB was the right choice for this use case.\n\u0026ldquo;Can I add email notifications?\u0026rdquo; Yes. Give a Lambda permission to use SES, trigger it from DynamoDB Streams. Students learn about event-driven architecture and service integration.\n\u0026ldquo;Can I add a CI/CD pipeline?\u0026rdquo; Yes. Add GitHub Actions, IAM roles with OIDC, and S3 sync logic. Students learn about deployment automation and security best practices.\nBecause the baseline is so small, every addition is visible. Every new service has a before-and-after moment. This is where the app stops being a demo and starts being a scaffold—students can extend it in any direction and immediately see what changes.\nChange the question. Add auth. Add metrics. Rip it apart. That\u0026rsquo;s how you build instincts. Not by memorizing services, but by wiring them together and watching what happens.\n","permalink":"https://lukelittle.com/posts/2025/12/pok%C3%A9mon-surveys-serverless-architecture-and-teaching-students-to-build-on-aws/","summary":"\u003cp\u003eBack in November, I was preparing for the Cracking the Cloud presentation at UNC Charlotte. I needed a way to explain how the cloud fundamentally changed what\u0026rsquo;s possible on the internet—not through abstract concepts, but through something students could immediately relate to.\u003c/p\u003e\n\u003cp\u003eThat\u0026rsquo;s when I remembered \u003cstrong\u003eThomas Game Docs\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eIf you\u0026rsquo;ve never heard of her: she\u0026rsquo;s a YouTuber who makes incredibly well-produced video essays about video games. And she sometimes runs surveys asking her audience things like \u0026ldquo;Who\u0026rsquo;s the LEAST popular Pokémon?\u0026rdquo; or \u0026ldquo;Who\u0026rsquo;s the LEAST popular Animal Crossing villager?\u0026rdquo;\u003c/p\u003e","title":"Pokémon Surveys, Serverless Architecture, and Teaching Students to Build on AWS"},{"content":"For this episode of Data Pour, I sat down with Lucas Ward—one of my SWE managers and a senior technical manager at Ippon—to talk about what it actually looks like to lead people in a consulting practice while the industry is shifting under our feet.\nWe filmed at the Charlotte Beer Garden, which feels like the most Charlotte setting possible: more taps than you can count, plus the steady soundtrack of loud cars and motorcycles rolling by. We both went with an OMB beer (Mecktoberfest) because… you kind of have to.\nIf you want the full conversation, watch the YouTube episode—this post is the background and the highlights.\nWhy I wanted Lucas on the show Lucas isn\u0026rsquo;t just a strong engineer—he\u0026rsquo;s the kind of leader who can translate. In consulting, that\u0026rsquo;s the job: you\u0026rsquo;re constantly moving between client expectations, delivery realities, and the growth of the people on your team.\nAnd Lucas has lived both sides.\nHe broke into tech the hard way—after years in restaurants—then got his start at a small Charlotte business where he had to wear every hat imaginable: shipping code, running systems in production, taking late-night calls when something broke, and learning by doing because there wasn\u0026rsquo;t anyone else to do it.\nThat \u0026ldquo;wear every hat\u0026rdquo; foundation shows up in how he leads today.\nPeople leadership in consulting is a different sport One of the most real parts of our conversation was the dynamic of managing people when you aren\u0026rsquo;t sitting next to them every day.\nIn a traditional org, you see the work constantly. In consulting, your direct report might be on another client, another team, another tech lead. Sometimes you\u0026rsquo;re getting signal through check-ins and feedback loops rather than direct observation—which makes coaching and performance conversations trickier (and honestly, more important to handle thoughtfully).\nLucas talked about what it means to guide technical folks who are laser-focused on hard skills—while also helping them build the soft skills that determine whether they actually grow: communication, delivery, stakeholder management, and the ability to explain why the work matters.\nTechnical skills get you a seat at the table. The rest is what keeps you there.\n\u0026ldquo;Consultant\u0026rdquo; = solve the problem and explain the value Lucas gave a definition I loved:\nA consultant is someone who can make the square peg fit in the round hole—and then explain how they did it to the people who care about the peg fitting.\nThat\u0026rsquo;s the work. Not just building. Translating. Connecting technical decisions to business outcomes. Helping stakeholders understand what\u0026rsquo;s happening and why it matters.\nEnterprise vs midsize: don\u0026rsquo;t write off the \u0026ldquo;breadth\u0026rdquo; engineers We also dug into how midsize clients operate differently than large enterprises—and what each can learn from the other.\nEnterprises have specialization and depth. Midsize companies often have engineers with ridiculous breadth because they\u0026rsquo;ve had to take products through the entire lifecycle—build, deploy, operate, support.\nLucas\u0026rsquo;s advice to enterprise hiring managers was simple: don\u0026rsquo;t dismiss candidates from smaller companies. They may not have the same \u0026ldquo;one deep slice\u0026rdquo; experience, but they often bring systems thinking and end-to-end ownership that\u0026rsquo;s hard to teach.\nAI readiness: everybody wants AI, but not everyone is ready This came up a lot:\nLarge enterprises are thinking about data strategy, fine-tuning, governance, and ROI.\nMidsize companies often want AI, but their data is siloed, inconsistent, and not operationally prepared for \u0026ldquo;real\u0026rdquo; AI use cases.\nAnd Lucas made a key point: AI readiness is incremental—and most of what you do to become \u0026ldquo;AI ready\u0026rdquo; is just good engineering anyway. Clean data, lineage, security controls, sustainable platforms, well-architected foundations. AI just forces the conversation.\nAlso: not every problem is an LLM problem. Traditional ML is having a comeback for a reason—it\u0026rsquo;s often cheaper, easier to operationalize, and more effective for specific tasks.\nCareer advice Lucas would give (and what he\u0026rsquo;d tell you) Lucas\u0026rsquo;s advice for people trying to break into tech right now:\nStick to it.\nDon\u0026rsquo;t let the AI noise psych you out. There are plenty of companies nowhere near \u0026ldquo;AI-first\u0026rdquo; that still need great engineers. Take the shot, even if it\u0026rsquo;s not the shiny job. Keep side projects going. Show your work. Stay learning. Don\u0026rsquo;t give up.\nAnd the advice he\u0026rsquo;d give himself 10 years ago was even better:\nIf you\u0026rsquo;re stuck, challenge your assumptions. Step back. Look elsewhere. There are more solutions than you think—most people just burn time spinning in one lane.\nWatch the episode This one is for anyone navigating leadership, consulting, or just trying to stay sane while AI changes the shape of the industry.\nWatch the full Data Pour episode with Lucas Ward here: https://www.youtube.com/watch?v=732vznyGndQ\u0026amp;t=1011s\n","permalink":"https://lukelittle.com/posts/2025/12/data-pour-with-lucas-ward-consulting-people-leadership-and-staying-curious-in-the-ai-era/","summary":"\u003cp\u003eFor this episode of Data Pour, I sat down with Lucas Ward—one of my SWE managers and a senior technical manager at Ippon—to talk about what it actually looks like to lead people in a consulting practice while the industry is shifting under our feet.\u003c/p\u003e\n\u003cp\u003eWe filmed at the Charlotte Beer Garden, which feels like the most Charlotte setting possible: more taps than you can count, plus the steady soundtrack of loud cars and motorcycles rolling by. We both went with an OMB beer (Mecktoberfest) because… you kind of have to.\u003c/p\u003e","title":"Data Pour with Lucas Ward: Consulting, People Leadership, and Staying Curious in the AI Era"},{"content":"AWS re:Invent 2025 keynotes felt like AWS repeating the same move they pulled 15 years ago—democratizing something that used to be gated behind massive budgets and specialized teams. This time it\u0026rsquo;s AI.\nI published a full write-up on Ippon\u0026rsquo;s blog breaking down the core thread running through the keynote: scale as the prerequisite for \u0026ldquo;access,\u0026rdquo; abundance via custom silicon (Trainium), and the shift from \u0026ldquo;AI as a feature\u0026rdquo; to \u0026ldquo;AI as a production capability.\u0026rdquo;\nThe big idea that stuck with me: it\u0026rsquo;s not just models—AWS is packaging hard-won operational and security expertise into agents (DevOps + Security) so more teams can build safely without needing a 30-person specialist squad.\nAWS has clearly gone all in on agentic AI, and it\u0026rsquo;s worth watching the keynote to understand where the industry is going.\nIf you\u0026rsquo;re trying to separate signal from noise after re:Invent, this is the lens I\u0026rsquo;d use—and the question I\u0026rsquo;d ask: now that the barriers keep dropping, what are you actually going to build?\nRead the full article here: https://blog.ippon.tech/takeaways-from-aws-reinvent-keynotes-2025\n","permalink":"https://lukelittle.com/posts/2025/12/aws-reinvent-2025-democratizing-ai-again/","summary":"\u003cp\u003eAWS re:Invent 2025 keynotes felt like AWS repeating the same move they pulled 15 years ago—democratizing something that used to be gated behind massive budgets and specialized teams. This time it\u0026rsquo;s AI.\u003c/p\u003e\n\u003cp\u003eI published a full write-up on Ippon\u0026rsquo;s blog breaking down the core thread running through the keynote: scale as the prerequisite for \u0026ldquo;access,\u0026rdquo; abundance via custom silicon (Trainium), and the shift from \u0026ldquo;AI as a feature\u0026rdquo; to \u0026ldquo;AI as a production capability.\u0026rdquo;\u003c/p\u003e","title":"AWS re:Invent 2025: Democratizing AI (Again)"},{"content":"We filmed this episode of Data Pour at People\u0026rsquo;s Market in Myers Park, Charlotte—grabbed drinks, hit record, and got into the kind of conversation that happens when two people who grew up in \u0026ldquo;big data\u0026rdquo; start comparing notes on where AI is actually headed.\nFunny thing is—this bottle shop closed down a week after we filmed. Perfect metaphor for tech, honestly: everything feels stable until it isn\u0026rsquo;t.\nWhy James James Barney is one of my closest friends and mentors. I\u0026rsquo;ve known him for a long time, and he\u0026rsquo;s been the person I go to when I want an AI take that isn\u0026rsquo;t hype and isn\u0026rsquo;t fear—just reality.\nHe\u0026rsquo;s spent the last decade-plus inside big enterprises, helping teams evaluate what\u0026rsquo;s real, what scales, and what\u0026rsquo;s just noise. Right now, he\u0026rsquo;s helping lead AI initiatives at a major insurance company—meaning his day job is basically: everyone wants the next big thing… which parts of this are going to stick?\nOur shared roots: fintech, big data, and invisible systems We both started our tech careers in big data, and a lot of our growth came from the same kind of pressure cooker—fintech needs.\nIf you\u0026rsquo;ve worked in regulated financial environments, you know the deal: huge volumes of data, constant scrutiny, and systems that have to be accurate, auditable, and fast. That world forces you to get good at the unsexy stuff—digesting, categorizing, transforming, and moving data reliably.\nIn the episode, we talk about the era when we were building Kafka infrastructure before half the managed conveniences existed. It was messy, duct-tape engineering—but it\u0026rsquo;s also the foundation for how so many modern \u0026ldquo;invisible\u0026rdquo; systems work today.\nThe point: AI is new… but the patterns aren\u0026rsquo;t One of the threads we pull on: AI feels new to everyone, but the underlying enterprise challenges rhyme with what we already lived through with cloud and big data.\nThe model is only part of the story. The real value shows up when you add:\nclean, governed data real context integrations into systems that matter tools that can retrieve or act (not just \u0026ldquo;generate text\u0026rdquo;) We also get into why copilots can feel underwhelming in enterprises: the public tools people use every day are \u0026ldquo;fully layered.\u0026rdquo; Inside a company, you\u0026rsquo;re often starting with the plain base layer—and you have to build the rest.\nWatch the full episode This blog is just the background and the framing—if you want the full story (and the full vibe), watch the YouTube episode. It\u0026rsquo;s a real conversation between two people who\u0026rsquo;ve been in the trenches, trying to describe what\u0026rsquo;s actually happening in AI without turning it into marketing copy or doomposting.\nLink to the episode: https://youtu.be/ibtZckIQ_JI\n","permalink":"https://lukelittle.com/posts/2025/12/data-pour-with-james-barney-from-big-data-to-ai/","summary":"\u003cp\u003eWe filmed this episode of Data Pour at People\u0026rsquo;s Market in Myers Park, Charlotte—grabbed drinks, hit record, and got into the kind of conversation that happens when two people who grew up in \u0026ldquo;big data\u0026rdquo; start comparing notes on where AI is actually headed.\u003c/p\u003e\n\u003cp\u003eFunny thing is—this bottle shop closed down a week after we filmed. Perfect metaphor for tech, honestly: everything feels stable until it isn\u0026rsquo;t.\u003c/p\u003e\n\u003ch2 id=\"why-james\"\u003eWhy James\u003c/h2\u003e\n\u003cp\u003eJames Barney is one of my closest friends and mentors. I\u0026rsquo;ve known him for a long time, and he\u0026rsquo;s been the person I go to when I want an AI take that isn\u0026rsquo;t hype and isn\u0026rsquo;t fear—just reality.\u003c/p\u003e","title":"Data Pour with James Barney: from Big Data to AI"},{"content":"Recently I came across another engineer\u0026rsquo;s personal blog — clean layout, good typography, that \u0026ldquo;I actually finish my side projects\u0026rdquo; energy — and it pushed me to finally build one of my own.\nI picked Hugo because I like Go, and because using Jekyll in 2025 feels like opting into pain. I briefly considered Ghost, remembered it either requires paying Ghost or hosting Ghost, and closed the tab. And since I\u0026rsquo;m \u0026ldquo;the AWS guy,\u0026rdquo; it felt morally necessary to deploy the whole thing on AWS. Maybe I\u0026rsquo;d even use Kiro if I felt extra fancy.\nFor context: Hugo and Jekyll are both static site generators—tools that convert Markdown files into HTML at build time. Jekyll, written in Ruby, was the OG choice for GitHub Pages and still powers thousands of blogs. But it\u0026rsquo;s slow, requires managing Ruby dependencies, and feels dated. Hugo, written in Go, is blazingly fast (builds in milliseconds), has a single binary with zero dependencies, and handles large sites without breaking a sweat. Both produce the same outcome—static HTML you can throw on a CDN—but Hugo does it in a fraction of the time and with far less friction.\nSo I wrote some Terraform, vibe-coded a theme, deployed it, and immediately remembered that I don\u0026rsquo;t build personal websites very often.\nThe Architecture Before diving into the problems I hit, here\u0026rsquo;s what the final architecture looks like.\nIt\u0026rsquo;s a classic serverless static site setup: Route 53 handles DNS, pointing the domain to a CloudFront distribution. CloudFront sits in front of an S3 bucket that stores all the static files—HTML, CSS, JavaScript, images. The bucket is private; CloudFront is the only thing allowed to read from it via an Origin Access Identity.\nHere\u0026rsquo;s where it gets slightly more interesting: attached to CloudFront is a CloudFront Function—a lightweight JavaScript function that runs at the edge, modifying incoming requests before they hit the origin. This function handles two things: redirecting www.lukelittle.com to the apex domain, and rewriting clean URLs (like /posts/my-article) to their actual paths (/posts/my-article/index.html).\nSSL certificates come from AWS Certificate Manager and are automatically attached to CloudFront. The whole thing is defined in Terraform, and deployment happens via GitHub Actions—push to main, Hugo builds the site, the output gets synced to S3, and CloudFront\u0026rsquo;s cache gets invalidated.\nThe request flow is straightforward: user hits the domain → Route 53 resolves it to CloudFront → CloudFront invokes the edge function → function rewrites the request if needed → CloudFront fetches from S3 (or serves from cache) → content gets delivered globally from the nearest edge location.\nClean. Simple. Serverless. Costs about $0.50/month for the Route 53 hosted zone. Everything else fits in free tier.\nNow, the problems.\nFirst realization: www doesn\u0026rsquo;t redirect itself I wanted www.lukelittle.com → lukelittle.com.\nSimple enough — except CloudFront doesn\u0026rsquo;t have .htaccess or Apache-style rewrite configs.\nThe fix is a CloudFront Function on viewer-request:\nsequenceDiagram participant User as 👤 User Browser participant CF as ☁️ CloudFront participant Func as ⚡ CloudFront Function User-\u0026gt;\u0026gt;CF: GET https://www.lukelittle.com/ CF-\u0026gt;\u0026gt;Func: Viewer Request Event Note over Func: Check host header\u0026lt;br/\u0026gt;host === \u0026#39;www.lukelittle.com\u0026#39; Func--\u0026gt;\u0026gt;CF: 301 Redirect Note over Func: Location: https://lukelittle.com/ CF--\u0026gt;\u0026gt;User: 301 Moved Permanently User-\u0026gt;\u0026gt;CF: GET https://lukelittle.com/ Note over User: Browser follows redirect CF--\u0026gt;\u0026gt;User: 200 OK (homepage) Here\u0026rsquo;s the actual CloudFront Function code:\nfunction handler(event) { var request = event.request; var host = request.headers.host.value; // Redirect www to apex domain if (host === \u0026#39;www.lukelittle.com\u0026#39;) { return { statusCode: 301, statusDescription: \u0026#39;Moved Permanently\u0026#39;, headers: { location: { value: \u0026#39;https://lukelittle.com\u0026#39; + request.uri } } }; } return request; } CloudFront Functions are perfect for this because they run at the edge, cost almost nothing (you get 2 million free invocations per month), and don\u0026rsquo;t require Lambda, bucket changes, or origin rewrites. They execute in under a millisecond, which means your redirect happens before the user even realizes they typed www.\nThat part was easy. The next part was not.\nSecond realization: CloudFront does not assume index.html Hugo outputs directories like:\n/posts/ /posts/index.html Apache and Nginx automatically serve index.html when you access /posts/.\nCloudFront does not. It will happily 404 unless you rewrite the URI yourself.\nHere\u0026rsquo;s what\u0026rsquo;s happening under the hood:\nsequenceDiagram participant User as 👤 User Browser participant CF as ☁️ CloudFront participant Func as ⚡ CloudFront Function participant S3 as 🪣 S3 Bucket User-\u0026gt;\u0026gt;CF: GET https://lukelittle.com/posts/cracking-the-cloud CF-\u0026gt;\u0026gt;Func: Viewer Request Event Note over Func: URI: /posts/cracking-the-cloud\u0026lt;br/\u0026gt;No extension detected\u0026lt;br/\u0026gt;!uri.includes(\u0026#39;.\u0026#39;) Func-\u0026gt;\u0026gt;Func: Append /index.html Note over Func: New URI:\u0026lt;br/\u0026gt;/posts/cracking-the-cloud/index.html Func--\u0026gt;\u0026gt;CF: Modified Request CF-\u0026gt;\u0026gt;S3: GetObject\u0026lt;br/\u0026gt;/posts/cracking-the-cloud/index.html S3--\u0026gt;\u0026gt;CF: HTML Content CF--\u0026gt;\u0026gt;User: 200 OK (post content) Note over CF: Cache for 1 hour So I added this logic to the same CloudFront Function:\nfunction handler(event) { var request = event.request; var uri = request.uri; var host = request.headers.host.value; // Redirect www to apex domain if (host === \u0026#39;www.lukelittle.com\u0026#39;) { return { statusCode: 301, statusDescription: \u0026#39;Moved Permanently\u0026#39;, headers: { location: { value: \u0026#39;https://lukelittle.com\u0026#39; + uri } } }; } // Append index.html for clean URLs if (uri.endsWith(\u0026#39;/\u0026#39;)) { request.uri += \u0026#39;index.html\u0026#39;; } else if (!uri.includes(\u0026#39;.\u0026#39;)) { request.uri += \u0026#39;/index.html\u0026#39;; } return request; } Now CloudFront behaves like a normal web server circa 2008. Hugo pages immediately started working.\nFor a moment.\nWhere everything went off the rails I tried to add www.lukelittle.com as an alternate domain name on the CloudFront distribution.\nCloudFront refused.\nThe error claimed it was already associated with another distribution.\nIt wasn\u0026rsquo;t — at least not in any AWS account I currently have access to.\nSo I did what any rational engineer does:\nGoogled Re-Googled Asked ChatGPT Deleted the distribution Recreated the distribution Repeated the cycle Began questioning my past life choices After four hours, I accepted defeat and temporarily upgraded my support plan.\nThe answer was unexpected:\nwww.lukelittle.com was attached to a CloudFront distribution — in another AWS account.\nWhich account?\nI have no clue.\nPossibilities include:\nsome forgotten sandbox from 2017 a leftover test account an old attempt at this blog I completely wiped from memory a parallel universe Support asked me to add a TXT record to prove I owned the domain. I added it in Route 53, they cleared the stale binding, and immediately everything began working exactly as expected.\nHow deployment actually works Once I got the infrastructure sorted, I needed a deployment pipeline. GitHub Actions + OIDC federation makes this dead simple:\nsequenceDiagram participant Dev as 👨‍💻 Developer participant GH as GitHub participant GHA as GitHub Actions participant IAM as 🔑 IAM Role participant S3 as 🪣 S3 Bucket participant CFront as ☁️ CloudFront Dev-\u0026gt;\u0026gt;GH: git push origin main GH-\u0026gt;\u0026gt;GHA: Trigger workflow Note over GHA: hugo --minify GHA-\u0026gt;\u0026gt;GHA: Build static site GHA-\u0026gt;\u0026gt;IAM: AssumeRoleWithWebIdentity Note over IAM: OIDC Federation\u0026lt;br/\u0026gt;Verify GitHub token IAM--\u0026gt;\u0026gt;GHA: Temporary credentials GHA-\u0026gt;\u0026gt;S3: aws s3 sync ./public s3://bucket/ Note over S3: Upload HTML, CSS, JS\u0026lt;br/\u0026gt;--delete flag removes old files S3--\u0026gt;\u0026gt;GHA: Sync complete GHA-\u0026gt;\u0026gt;CFront: CreateInvalidation --paths \u0026#34;/*\u0026#34; Note over CFront: Clear edge cache\u0026lt;br/\u0026gt;Force fresh content CFront--\u0026gt;\u0026gt;GHA: Invalidation ID Note over Dev,CFront: Deployment complete ✅\u0026lt;br/\u0026gt;New content live globally Here\u0026rsquo;s what makes this beautiful:\nNo long-lived credentials. GitHub Actions uses OIDC to assume an IAM role, gets temporary credentials that expire in an hour, and those credentials only work for this specific repo. If someone compromises the GitHub Actions environment, they get access for 60 minutes max—and only to deploy this blog. Not exactly a treasure trove.\nHugo builds in ~200ms. Static site generators are fast when your entire site fits in memory. No database queries, no server-side rendering, just Markdown → HTML and done.\nS3 sync is smart. It only uploads files that changed. New post? Upload one file. Tweak CSS? Upload one file. CloudFront invalidation clears the edge cache, so every visitor gets fresh content within seconds.\nGlobal deployment in under a minute. Push to main → build → sync → invalidate → live. The entire pipeline runs faster than most people can brew coffee.\nAnd now the site actually exists Despite writing constantly — deep dives, rants, slides, Data Pour episodes — I\u0026rsquo;ve never had a single place to put any of it. Everything has been scattered across GitHub repos, Slack threads, LinkedIn posts, and random folders.\nThis site fixes that.\nHugo prerenders everything.\nCloudFront serves it globally.\nThere\u0026rsquo;s no backend.\nNo patching.\nNo maintenance.\nJust HTML, a CDN, and vibes.\nThe whole stack:\nHugo for static site generation S3 for object storage CloudFront for CDN + edge functions Route 53 for DNS ACM for SSL certificates GitHub Actions for CI/CD Terraform for infrastructure as code Total monthly cost: ~$0.50 for Route 53 hosted zone. Everything else fits in free tier.\nThe unexpected benefit Honestly, getting stuck for a few hours was probably good for me. I had to slow down, re-read documentation, and remember exactly how CloudFront, ACM, and Route 53 interact — instead of relying on half-remembered muscle memory.\nBuilding this site reminded me why I got into infrastructure in the first place: you can take a bunch of managed services, wire them together thoughtfully, and end up with something that just works. No servers to patch, no databases to tune, no midnight pages about memory leaks.\nJust a website that loads fast, costs nothing, and requires zero maintenance.\nNot the night I planned, but not wasted either.\nAnyway — the blog is live now. Hopefully the next update doesn\u0026rsquo;t require another round of DNS archaeology or CloudFront forensics.\nWant to see the code? The entire infrastructure setup is open source:\ngithub.com/lukelittle/homepage\n","permalink":"https://lukelittle.com/posts/2025/12/i-tried-to-deploy-a-simple-website-on-aws.-it-became-a-full-blown-side-quest./","summary":"\u003cp\u003eRecently I came across another engineer\u0026rsquo;s personal blog — clean layout, good typography, that \u0026ldquo;I actually finish my side projects\u0026rdquo; energy — and it pushed me to finally build one of my own.\u003c/p\u003e\n\u003cp\u003eI picked Hugo because I like Go, and because using Jekyll in 2025 feels like opting into pain. I briefly considered Ghost, remembered it either requires paying Ghost or hosting Ghost, and closed the tab. And since I\u0026rsquo;m \u0026ldquo;the AWS guy,\u0026rdquo; it felt morally necessary to deploy the whole thing on AWS. Maybe I\u0026rsquo;d even use Kiro if I felt extra fancy.\u003c/p\u003e","title":"I Tried to Deploy a Simple Website on AWS. It Became a Full-Blown Side Quest."},{"content":"When I talk to students — especially STEM students — one question keeps coming up: \u0026ldquo;Is AI going to take my job?\u0026rdquo; They ask it jokingly, but you can tell they\u0026rsquo;re serious. What they want is assurance. They want to know there\u0026rsquo;s a light at the end of the tunnel. That the late nights, the debt, the effort, and the hope they\u0026rsquo;ve poured into their degree will amount to something real.\nAnd I rarely give a quick answer, because the truth is uncomfortable: maybe. It\u0026rsquo;s not the certainty anyone wants, but it\u0026rsquo;s honest.\nFor a long time, I danced around that answer. I hedged. I emphasized the optimistic parts, hoping they would land softer. But last week, during his final re:Invent keynote, Werner Vogels showed us how to talk about this moment with clarity instead of avoidance.\nHe walked the audience through the history of software development — from punch cards and COBOL to structured programming, object orientation, distributed systems, the cloud, and now agentic AI. In every era, developers were convinced the sky was falling. And in some ways, it was. Punch-card clerks disappeared. Mainframe operators disappeared. Entire roles vanished because something faster, cheaper, and more scalable came along.\nBut one role never disappeared: the builder — the person who understands systems, solves real-world problems, thinks in abstractions, and brings ideas to life.\nThat\u0026rsquo;s the part people forget. Technology doesn\u0026rsquo;t just destroy. Technology democratizes. It opens doors. And it does so in waves.\nEvery wave has displaced someone. Every wave has taken something away. But every wave has also created something bigger for the people willing to evolve. Builders adapted. They reskilled. They moved up the stack. They went from pushing paper to writing code, from writing code to designing systems. The pattern is always the same: a wave arrives, jobs shift, and the builders who lean in rise with it.\nI\u0026rsquo;ve lived through enough of these waves to see the pattern clearly.\nThe first wave I experienced was the World Wide Web. Before the web, we logged into CompuServe, posted on message boards, browsed clunky online catalogs. And then came the browser — Netscape, Mosaic, whatever arrived on your screen first. Suddenly anyone with a little HTML could put a page online for the entire world to see. The web was so small in those days that there were literal lists of websites. That\u0026rsquo;s how new it was.\nAnd that was the first time technology felt like something you could create with, not just consume. The web democratized publishing and expression. It changed how we interacted with the world — Amazon instead of bookstores, eBay instead of classifieds — and it changed what it meant to be a builder. You didn\u0026rsquo;t need a printing press or a distribution network. You needed curiosity and a willingness to learn a markup language.\nThen came the cloud. Suddenly anyone with a credit card and a vision could become a global service provider. You didn\u0026rsquo;t need a server room. You didn\u0026rsquo;t need procurement approvals. You didn\u0026rsquo;t need capital. The cloud collapsed the distance between idea and impact. Entire companies were born in dorm rooms because the playing field had leveled.\nAnd now we\u0026rsquo;re in the third wave: agentic AI. Tools that can generate code, test it, reason about it, orchestrate systems, and solve problems alongside you — without you manually writing every line. You don\u0026rsquo;t need specialized hardware or a custom model. You can build on top of general-purpose models, wrap your logic behind a Model Context Protocol, and stand up entire services in hours. Once again, the gap between imagination and execution has narrowed.\nThis is what Vogels calls the Renaissance Developer — someone who cultivates depth and breadth. Someone who goes deep enough to solve hard problems but understands enough of adjacent systems to see the whole picture. Someone who stays curious, thinks in systems, communicates clearly, takes ownership, and adapts when the next wave comes.\nBecause there will always be a next wave.\nSo when students ask whether AI will take their job, I finally know how to answer. The question is slightly off. The better question is: What can I build with AI?\nIf you define yourself by a tool — a language, a framework, a stack — then yes, you\u0026rsquo;re in trouble. Tools come and go. They always have. The developers who clung to punch cards got left behind. The ones who refused to learn the web got left behind. The ones who ignored the cloud got left behind. And the ones who dismiss AI will get left behind too.\nBut if you define yourself by what you create — if you define yourself as a builder — then AI isn\u0026rsquo;t a threat. AI is another wave of democratization. AI is a force multiplier. AI is the next thing that expands what\u0026rsquo;s possible for people willing to evolve.\nThe path forward isn\u0026rsquo;t about predicting the next dominant framework or clinging to one narrow specialty. The path forward is about building. That\u0026rsquo;s the one thing AI cannot replace: the human who knows what should exist and has the drive to bring it to life.\nSo will AI take your job? Maybe. Probably, if the job is defined purely by a tool.\nBut will AI take away your ability to create, adapt, imagine, build? No. It won\u0026rsquo;t. It can\u0026rsquo;t.\nIf you lean into what it means to be a Renaissance Developer, then AI doesn\u0026rsquo;t close the door on your future. AI blows it wide open.\n","permalink":"https://lukelittle.com/posts/2025/12/ai-uncertainty-and-the-rise-of-the-renaissance-developer/","summary":"\u003cp\u003eWhen I talk to students — especially STEM students — one question keeps coming up: \u0026ldquo;Is AI going to take my job?\u0026rdquo; They ask it jokingly, but you can tell they\u0026rsquo;re serious. What they want is assurance. They want to know there\u0026rsquo;s a light at the end of the tunnel. That the late nights, the debt, the effort, and the hope they\u0026rsquo;ve poured into their degree will amount to something real.\u003c/p\u003e","title":"AI, Uncertainty, and the Rise of the Renaissance Developer"},{"content":"I was invited to speak at UNC Charlotte recently to the computer science programs, and I approached the session, Cracking the Cloud, with a very specific goal: to give students a realistic, actionable way to stand out in a job market that\u0026rsquo;s becoming more competitive every year. Companies are slowing early-career hiring. AI is reshaping workflows. Expectations for junior talent are rising, not shrinking. Students can sense this shift, but many don\u0026rsquo;t know what to do with that reality. That\u0026rsquo;s where the conversation begins.\nThe message I shared with them is simple: even as the job market tightens, students today have an unprecedented advantage that previous generations didn\u0026rsquo;t — they can build real, meaningful things from their dorm rooms. Modern cloud platforms have removed the resource barriers that once kept students from gaining hands-on experience. They don\u0026rsquo;t need servers, budget approvals, or someone in authority to greenlight their ideas. With nothing more than a laptop and the AWS Free Tier, they can deploy applications, experiment with architectures, analyze data, and work with AI. The tools that used to be locked behind enterprise doors are now available to anyone willing to try.\nThis democratization of technology matters because it gives students something incredibly valuable: the ability to build experience before they have experience. When the market tightens, that distinction becomes decisive. Employers want graduates who can contribute, who understand how modern systems work, and who have touched real tools — not just read about them in a textbook. Students who start building early walk into interviews with confidence and momentum. Students who wait often find themselves starting from behind.\nThat\u0026rsquo;s why I encourage students to pursue an early cloud certification. It\u0026rsquo;s not about collecting badges. It\u0026rsquo;s about giving them a foundation that unlocks everything else. A certification gets them past HR filters, but more importantly, it teaches them the vocabulary and mental models they need to become actual builders. Once they have that baseline, the cloud stops feeling abstract. They can dive in, deploy something real, break things, fix them, and learn through doing. That\u0026rsquo;s where real growth happens.\nTo help them get started, I shared a simple 30-60-90 plan:\nIn the first 30 days, learn the fundamentals. The AWS Cloud Practitioner exam is a gentle entry point that teaches students how the cloud works and how the pieces fit together.\nBy 60 days, build something small. A website, an API, a basic data pipeline — anything that forces them to make architectural decisions and use real services.\nBy 90 days, connect with people. Attend meetups, talk with alumni, share what they built, and start forming the relationships that will carry them into internships and full-time roles.\nThis plan works because it\u0026rsquo;s simple, it\u0026rsquo;s doable, and it converts uncertainty into direction. Students don\u0026rsquo;t need to wait for permission. They just need a starting point.\nWhat Comes Next The response to the talk made one thing clear: students are hungry for guidance, community, and hands-on support. So naturally, I bought CrackingTheCloud.com, and I intend to do something meaningful with it. The vision is bigger than a single event or a single presentation. I want to build a space — both digital and physical — where early-career engineers can learn from industry, explore real tools, and develop confidence long before graduation.\nOver the coming months, I plan to start hosting meetups in Charlotte and Richmond, bringing together students, engineers, hiring managers, and community leaders. The goal isn\u0026rsquo;t to create another generic networking group. It\u0026rsquo;s to create an environment where students can ask real questions, see real demos, get unfiltered career advice, and make connections that matter. When students can sit across from practitioners, hear how the industry works, and get direct feedback on their projects, the gap between classroom theory and real-world engineering gets much smaller.\nAnd this matters — not just for students, but for the entire ecosystem. When we invest in early talent, we improve the pipeline for every company in the region. We reduce onboarding burden. We strengthen local tech communities. And we help students step into the industry with confidence, clarity, and purpose. The cloud may have democratized the tools, but it\u0026rsquo;s up to us to democratize the guidance.\nCracking the Cloud started as a talk, but it won\u0026rsquo;t end there. There\u0026rsquo;s a real opportunity to build something lasting — a community that helps students navigate a challenging market, access modern tools, and become the kind of builders the industry needs. And if the energy from UNC Charlotte is any indication, this is only the beginning.\n","permalink":"https://lukelittle.com/posts/2025/12/cracking-the-cloud-preparing-students-to-build-in-a-tougher-job-market/","summary":"\u003cp\u003eI was invited to speak at UNC Charlotte recently to the computer science programs, and I approached the session, Cracking the Cloud, with a very specific goal: to give students a realistic, actionable way to stand out in a job market that\u0026rsquo;s becoming more competitive every year. Companies are slowing early-career hiring. AI is reshaping workflows. Expectations for junior talent are rising, not shrinking. Students can sense this shift, but many don\u0026rsquo;t know what to do with that reality. That\u0026rsquo;s where the conversation begins.\u003c/p\u003e","title":"Cracking the Cloud: Preparing Students to Build in a Tougher Job Market"},{"content":"I recently stepped into the role of host for Ippon\u0026rsquo;s Data Pour series, but before I officially took over, our CTO, Andy Lamora, sat me down at Wooden Robot Brewery in Charlotte and interviewed me on camera. It was a great atmosphere — good beer, good weather — and then a handful of cameras appeared, all pointed directly at me. I\u0026rsquo;m still getting used to public speaking, so the entire setup felt a little awkward. Beer helps, but only so much.\nThe conversation itself was wide-ranging in the best way. Andy and I talked through my background and how I somehow ended up moving from baking to banking to cloud consulting. It\u0026rsquo;s not a linear story, but that\u0026rsquo;s the point — careers rarely follow a straight line, and most of the interesting stuff happens in the detours. We dug into some of the pivots that shaped my path, the things that pulled me deeper into cloud and DevOps work, and what I\u0026rsquo;ve learned hopping between industries and roles along the way.\nWe eventually moved into the topics I spend most of my time thinking about: cloud platforms, data, and where AI fits into all of it. I shared some opinions on agentic systems, the future of operational automation, and why building a strong platform data foundation matters long before you start layering in intelligent tooling. It wasn\u0026rsquo;t a rehearsed script — just two people talking through where the industry is heading, what excites us, and what still makes us pause.\nFilming at Wooden Robot added an unexpected layer to the experience. Between the smell of beer brewing, the coffee bar behind us, and the general energy of the place, the whole conversation felt more natural than anything you\u0026rsquo;d get in a studio. That\u0026rsquo;s the vibe I want to bring into Season 4: real conversations, in real environments, with people who actually build and think about this stuff every day.\nIf you want to watch me awkwardly navigate my first time on camera — before I technically became the host — the episode is here: https://www.youtube.com/watch?v=ZyoRAQsS9c0\nSeason 4 has since wrapped. Every episode is on the Talks page.\n","permalink":"https://lukelittle.com/posts/2025/11/taking-the-mic-on-the-data-pour/","summary":"\u003cp\u003eI recently stepped into the role of host for Ippon\u0026rsquo;s Data Pour series, but before I officially took over, our CTO, Andy Lamora, sat me down at Wooden Robot Brewery in Charlotte and interviewed me on camera. It was a great atmosphere — good beer, good weather — and then a handful of cameras appeared, all pointed directly at me. I\u0026rsquo;m still getting used to public speaking, so the entire setup felt a little awkward. Beer helps, but only so much.\u003c/p\u003e","title":"Taking the Mic on The Data Pour"},{"content":"On the morning of October 20, AWS us-east-1 services were degraded—in particular, DNS services for DynamoDB. Most of us didn\u0026rsquo;t find out from monitoring alerts or dashboards. We found out because the apps on our phones stopped working.\nThat\u0026rsquo;s the reality of modern infrastructure incidents: they often surface as user-facing failures long before the official root cause analysis lands in your inbox.\nI wrote about this outage for Ippon Technologies, focusing on three critical aspects that go beyond just understanding what broke:\nWhat Actually Happened This wasn\u0026rsquo;t a full region going dark—it was a DNS problem at a foundational layer. DNS (Domain Name System) is the internet\u0026rsquo;s address book. When DNS breaks, everything that depends on it breaks too.\nIt\u0026rsquo;s like someone removing all the street signs in a city overnight—your services are still there, but nothing can find its way.\nFor many teams, it meant increased error rates, intermittent failures, and retry storms. Not catastrophic downtime, but the kind of disruption that floods support tickets and frustrates users.\nHaving the Leadership Conversation This is where many engineers struggle. You know the issue was upstream. You know it\u0026rsquo;s AWS\u0026rsquo;s infrastructure. But leadership doesn\u0026rsquo;t care about the cloud provider—they care about impact and what you\u0026rsquo;re doing about it.\nThe key is framing it as a dependency visibility problem, not a blame game:\n\u0026ldquo;We were affected by a regional outage in AWS\u0026rsquo;s us-east-1 region due to DNS issues. This exposed areas where we\u0026rsquo;re overly dependent on single-region infrastructure. We\u0026rsquo;re using this to map our regional dependencies, prioritize applications by criticality, and identify where we need fallback logic and multi-region routing.\u0026rdquo;\nThat\u0026rsquo;s ownership. That\u0026rsquo;s a path forward.\nThe Hard Questions You Need to Answer The article explores four critical questions every team should be able to answer confidently:\nWhich of your apps run in us-east-1? Which rely on DynamoDB? Which of those are Tier 1 or customer-facing? Which of those have active-active failover across regions? Most teams can\u0026rsquo;t answer these questions. Not because they\u0026rsquo;re negligent—but because cloud estates grow organically. Services get deployed. Teams change. Documentation drifts. Before you know it, you\u0026rsquo;re running critical workloads on infrastructure patterns that no one fully understands anymore.\nMaking Resilience Visible The full article dives into:\nStructured risk assessment using tools like AWS Resilience Hub to define applications and assess risk against RTO/RPO targets Chaos engineering with AWS Fault Injection Service (FIS) to validate that resilience isn\u0026rsquo;t just theoretical Cultural shifts to prioritize resilience alongside feature delivery Practical next steps for mapping dependencies and building observability around failure modes Why This Matters Today\u0026rsquo;s outage wasn\u0026rsquo;t catastrophic—but it was loud enough to get everyone\u0026rsquo;s attention. It revealed real architectural risks that often go unnoticed until they become outages.\nOutages like this are reminders, not just disruptions. They\u0026rsquo;re opportunities to begin conversations across architecture, risk, and engineering teams about what resilience really means for your organization.\nNot in terms of making everything indestructible, but in making risk visible, decisions intentional, and recovery predictable.\n👉 Read the full article on the Ippon blog for detailed guidance on communicating with leadership, assessing your infrastructure, and building measurable resilience practices.\nThe goal isn\u0026rsquo;t perfection—it\u0026rsquo;s visibility, intention, and readiness for when the next incident inevitably arrives.\n","permalink":"https://lukelittle.com/posts/2025/10/learning-from-the-october-20-aws-outage-questions-every-team-should-ask/","summary":"\u003cp\u003eOn the morning of October 20, AWS us-east-1 services were degraded—in particular, DNS services for DynamoDB. Most of us didn\u0026rsquo;t find out from monitoring alerts or dashboards. We found out because the apps on our phones stopped working.\u003c/p\u003e\n\u003cp\u003eThat\u0026rsquo;s the reality of modern infrastructure incidents: they often surface as user-facing failures long before the official root cause analysis lands in your inbox.\u003c/p\u003e\n\u003cp\u003eI wrote about this outage for \u003ca href=\"https://blog.ippon.tech/explaining-the-october-20-aws-outage-to-leadership/\"\u003eIppon Technologies\u003c/a\u003e, focusing on three critical aspects that go beyond just understanding what broke:\u003c/p\u003e","title":"Learning from the October 20 AWS Outage: Questions Every Team Should Ask"},{"content":" I\u0026rsquo;m Luke Little. My career started in baking, moved into banking, and ended up in cloud consulting, and I wouldn\u0026rsquo;t trade any of the detours. Today I\u0026rsquo;m Head of Intelligent Data and Cloud at Ippon Technologies, where I lead our cloud, data, and AI infrastructure work. I build cloud and AI systems for banking and capital markets, and I spend a lot of my spare time helping engineers, teams, and students level up. I\u0026rsquo;m also pursuing my MBA at Duke\u0026rsquo;s Fuqua School of Business (Class of 2028), and I serve on the UNC Charlotte Alumni Board of Directors.\nThis site is where I write about what I\u0026rsquo;ve learned, what I\u0026rsquo;m learning, and what I think matters.\nWhat I write about Regulated systems on AWS. Architectures for the rules that keep markets safe: pre-trade risk controls, margin monitoring, best execution, and order lifecycle reporting. Start with the Regulated Markets on AWS series. Applied AI and agents. Hands-on builds with Amazon Bedrock, MCP, and agentic coding tools, with a focus on running them safely in environments where compliance matters, including guardrails and governance strategies for AI in regulated industries. See everything tagged AI. Teaching the next generation of builders. How students can get real experience before their first job: certifications, cloud projects, and a public portfolio. See posts tagged education and career. Community. I\u0026rsquo;m the lead organizer of the Richmond AWS User Group, I give talks at universities, and I hosted Season 4 of Ippon\u0026rsquo;s The Data Pour. The Talks page has all of it. Invite me I enjoy speaking with students, meetups, and engineering teams about cloud careers, AWS architecture in regulated industries, and building with AI. If that sounds useful for your group, bug me on LinkedIn.\nGet in touch The best way to reach me is to bug me on LinkedIn. My code is on GitHub, and you can follow along via RSS.\n","permalink":"https://lukelittle.com/about/","summary":"\u003cfigure class=\"about-photo\"\u003e\n  \u003cimg src=\"/images/luke.jpg\" alt=\"Luke Little mid-kick on a steep cobblestone street, with mountains behind him\" width=\"928\" height=\"1400\"\u003e\n\u003c/figure\u003e\n\u003cp\u003eI\u0026rsquo;m Luke Little. My career started in baking, moved into banking, and ended up in cloud consulting, and I wouldn\u0026rsquo;t trade any of the detours. Today I\u0026rsquo;m Head of Intelligent Data and Cloud at Ippon Technologies, where I lead our cloud, data, and AI infrastructure work. I build cloud and AI systems for banking and capital markets, and I spend a lot of my spare time helping engineers, teams, and students level up. I\u0026rsquo;m also pursuing my MBA at Duke\u0026rsquo;s Fuqua School of Business (Class of 2028), and I serve on the UNC Charlotte Alumni Board of Directors.\u003c/p\u003e","title":"About"},{"content":"Last updated: September 2026. This is a now page.\nWriting: finishing the Regulated Markets on AWS series. Learning: working through my MBA at Duke (Class of 2028). Speaking: fresh off Cracking the Cloud at UNC Charlotte in September. ","permalink":"https://lukelittle.com/now/","summary":"\u003cp\u003e\u003cem\u003eLast updated: September 2026. This is a \u003ca href=\"https://nownownow.com/about\"\u003enow page\u003c/a\u003e.\u003c/em\u003e\u003c/p\u003e\n\u003c!-- TODO(Luke): add a \"Building\" line if you want one --\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eWriting:\u003c/strong\u003e finishing the \u003ca href=\"/series/regulated-markets-on-aws/\"\u003eRegulated Markets on AWS\u003c/a\u003e series.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eLearning:\u003c/strong\u003e working through my MBA at Duke (Class of 2028).\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSpeaking:\u003c/strong\u003e fresh off \u003ca href=\"/talks/\"\u003eCracking the Cloud\u003c/a\u003e at UNC Charlotte in September.\u003c/li\u003e\n\u003c/ul\u003e","title":"Now"},{"content":"Most of what I build starts as a demo for a talk, a class, or a question I couldn\u0026rsquo;t stop thinking about. Everything here has a write-up, and most of it has code you can deploy yourself.\nRegulated markets on AWS Pre-trade risk controls (SEC Rule 15c3-5)\nStreaming market access controls with Kafka and Spark, inspired by the Knight Capital incident.\nWrite-up · Code\nMore in the Regulated Markets on AWS series.\nAI and agents Vinyl collection chatbot with FastMCP\nA serverless MCP server that lets an AI agent query a record collection.\nWrite-up · Code\nGitHub PR reviewer with Bedrock Agents\nAn agent that reviews pull requests using Bedrock Agents and action groups.\nWrite-up\nAWS cost optimization agent\nA Bedrock agent that reads Cost Explorer data and recommends ways to cut your bill.\nWrite-up\nCompany knowledge bot\nSlack plus Bedrock Knowledge Bases for instant answers from your internal docs.\nWrite-up\nServerless URL shortener, built with Kiro\nRoughly 15 hours of Terraform done in 3 with an agentic coding tool.\nWrite-up · Code\nTeaching Cracking the Cloud survey app\nA small serverless Pokémon survey app I use to show students what the cloud can do.\nWrite-up · Code\nAdding Cognito authentication to the survey app\nThe follow-up: securing the same app with Amazon Cognito.\nWrite-up · Code\nStudent branding starter\nA template for students to launch a personal site and online presence.\nWrite-up · Code\nThis site lukelittle.com\nHugo on S3 and CloudFront, deployed by GitHub Actions with Terraform, for about $3 a year.\nWrite-up · Code\\\n","permalink":"https://lukelittle.com/projects/","summary":"\u003cp\u003eMost of what I build starts as a demo for a talk, a class, or a question I couldn\u0026rsquo;t stop thinking about. Everything here has a write-up, and most of it has code you can deploy yourself.\u003c/p\u003e\n\u003ch2 id=\"regulated-markets-on-aws\"\u003eRegulated markets on AWS\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003ePre-trade risk controls (SEC Rule 15c3-5)\u003c/strong\u003e\u003cbr\u003e\nStreaming market access controls with Kafka and Spark, inspired by the Knight Capital incident.\u003cbr\u003e\n\u003ca href=\"/posts/2026/02/designing-pre-trade-risk-controls-on-aws-sec-rule-15c3-5/\"\u003eWrite-up\u003c/a\u003e · \u003ca href=\"https://github.com/lukelittle/sec-15c3-5-market-access-controls-example\"\u003eCode\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eMore in the \u003ca href=\"/series/regulated-markets-on-aws/\"\u003eRegulated Markets on AWS\u003c/a\u003e series.\u003c/p\u003e\n\u003ch2 id=\"ai-and-agents\"\u003eAI and agents\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eVinyl collection chatbot with FastMCP\u003c/strong\u003e\u003cbr\u003e\nA serverless MCP server that lets an AI agent query a record collection.\u003cbr\u003e\n\u003ca href=\"/posts/2026/01/fastmcp-and-the-vinyl-collection-chatbot-serverless-agentic-ai-in-action/\"\u003eWrite-up\u003c/a\u003e · \u003ca href=\"https://github.com/lukelittle/rawsug-fastmcp-demo\"\u003eCode\u003c/a\u003e\u003c/p\u003e","title":"Projects"},{"content":"I like talking about cloud careers, AWS architecture, and building with AI, especially with students and early-career engineers. If you\u0026rsquo;d like me to speak at your school, meetup, or team, bug me on LinkedIn.\nSpeaker kit Organizing an event? Everything you need to promote a talk is here.\nDownload headshot Short bio\nLuke Little is Head of Intelligent Data and Cloud at Ippon Technologies, where he leads cloud, data, and AI infrastructure work. He is the lead organizer of the Richmond AWS User Group and an MBA candidate at Duke\u0026rsquo;s Fuqua School of Business.\nLonger bio\nLuke Little is Head of Intelligent Data and Cloud at Ippon Technologies, where he leads cloud, data, and AI infrastructure work, much of it for banking and capital markets. His career runs from baking to banking to cloud consulting, and he writes at lukelittle.com about building regulated systems on AWS and running AI safely in production. Luke is the lead organizer of the Richmond AWS User Group, hosted Season 4 of Ippon\u0026rsquo;s podcast The Data Pour, and regularly speaks to computing students about breaking into cloud careers. He serves on the UNC Charlotte Alumni Board of Directors and is pursuing an MBA at Duke\u0026rsquo;s Fuqua School of Business.\nTopics I speak on\nCracking the Cloud. How students and early-career engineers can stand out with certifications, real projects, and relationships, even in a tough market. AI guardrails and governance. How to run generative AI safely in regulated industries: guardrails, governance design, and audit trails. Agentic AI on AWS. Building and operating AI agents with MCP and Amazon Bedrock, usually as a live demo. Regulated systems on AWS. Streaming architectures for SEC and FINRA rules like market access controls and margin monitoring. Talks Cracking the Cloud (2026 edition)\nUNC Charlotte · September 2026\nBack at UNC Charlotte, this time as a member of the alumni board, on how students can stand out in a tough market with certifications, projects, and relationships.\nAI guardrails and governance in regulated industries\nRichmond AWS User Group · August 2026\nHow guardrails wrap model invocations, and how to design governance strategies for AI where compliance matters.\nCracking the Cloud: How AWS Certifications Can Launch Your Career\nVirginia Commonwealth University, Department of Computer Science · March 2026\nA session for CS seniors on standing out with certifications, self-directed projects, and real relationships.\nWrite-up\nFastMCP on AWS: a live demo\nRichmond AWS User Group · February 2026\nAn unrehearsed, AI-assisted build of an MCP server on AWS, hiccups included.\nWrite-up · Code and prompt\nCracking the Cloud: Preparing Students to Build in a Tougher Job Market\nUNC Charlotte, Computer Science · December 2025\nA 30-60-90 day plan for students to get hands-on cloud experience before their first job.\nWrite-up · Materials\nPodcast: The Data Pour (Season 4 host) I hosted Season 4 of Ippon\u0026rsquo;s The Data Pour: real conversations about cloud, data, and AI with the people who build it.\nEpisode Guest Links S4E1 Me, interviewed by Andy Lamora Write-up · Watch S4E2 James Barney Write-up · Watch S4E3 Lucas Ward Write-up · Watch S4E4 Dr. Mohamed Shehab Write-up · Watch S4E5 Nimish Donde Write-up · Watch Community I\u0026rsquo;m the lead organizer of the Richmond AWS User Group. Recent meetups include AI guardrails and governance (August 2026), FastMCP on AWS and a session on Kiro.\n","permalink":"https://lukelittle.com/talks/","summary":"\u003cp\u003eI like talking about cloud careers, AWS architecture, and building with AI, especially with students and early-career engineers. If you\u0026rsquo;d like me to speak at your school, meetup, or team, \u003ca href=\"https://www.linkedin.com/in/lucaslittle/\"\u003ebug me on LinkedIn\u003c/a\u003e.\u003c/p\u003e\n\u003ch2 id=\"speaker-kit\"\u003eSpeaker kit\u003c/h2\u003e\n\u003cp\u003eOrganizing an event? Everything you need to promote a talk is here.\u003c/p\u003e\n\u003cfigure class=\"speaker-photo\"\u003e\n  \u003cimg src=\"/images/luke-headshot.jpg\" alt=\"Headshot of Luke Little\" width=\"780\" height=\"1000\"\u003e\n  \u003cfigcaption\u003e\u003ca href=\"/images/luke-headshot.jpg\" download\u003eDownload headshot\u003c/a\u003e\u003c/figcaption\u003e\n\u003c/figure\u003e\n\u003cp\u003e\u003cstrong\u003eShort bio\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLuke Little is Head of Intelligent Data and Cloud at Ippon Technologies, where he leads cloud, data, and AI infrastructure work. He is the lead organizer of the Richmond AWS User Group and an MBA candidate at Duke\u0026rsquo;s Fuqua School of Business.\u003c/p\u003e","title":"Talks \u0026 Media"}]