Site Reliability and Cloud Infrastructure

Nate Fletcher

AWS, Azure, Terraform, observability, migration, incident response, and production operations for enterprise SaaS.

I work where cloud infrastructure, reliability, automation, and customer-facing production support meet. My path from L3 support and product engineering into SRE gives me a practical bias toward systems that can be operated, documented, and handed off cleanly.

Jacksonville, FL nate@natefl.com

Engineering Work

Production initiatives across cloud migration, security controls, observability, storage modernization, and repeatable operations.

Security platform

Centralized AWS WAFv2 Management Platform

Built a Git-driven Terraform and Terragrunt platform for managing AWS WAFv2 across multi-account, multi-region workloads.

  • Supports EKS environments and standalone ALB-backed products across US and Canada deployments.
  • Centralized shared rules, managed rule groups, rate limits, ALB associations, logging, and global IP blocking.
  • Built GitHub Actions planning/deployment with OIDC authentication and dependency-aware change detection.

Cloud migration

On-Premises to AWS Migration

Led hands-on migration work from datacenter infrastructure to AWS EC2 across customer-facing SaaS environments.

  • Coordinated EC2 migration, DNS preparation, load balancer updates, application configuration, and validation.
  • Worked across customer-facing applications and internal tools moving from datacenter infrastructure to AWS.
  • Supported production cutovers and post-migration cleanup.

Automation

Trust Secret Expiration Notifier

Replaced an Azure Logic Apps monitoring workflow with an AWS-native EventBridge and Lambda service.

  • Queries private SQL Server databases weekly for trust secrets expiring within 30 days.
  • Uses VPC-connected Lambda, Secrets Manager, CloudWatch Logs, SNS dead-letter handling, and reusable Terraform.
  • Keeps SDLC and production deployments isolated while using the same module and application logic.

Database automation

DDL Automation Pipeline

Built a biweekly schema extraction and version-control workflow for SQL Server databases.

  • Covers ~2,038 databases across 8 EC2 SQL Server instances and 5 RDS instances.
  • Replaced production-impacting SMO extraction with direct T-SQL catalog queries.
  • Stores extracted schema output in S3 and commits version-controlled records to GitHub.

Role Fit

A fit for hands-on reliability, DevOps, cloud infrastructure, and platform roles that need practical production experience, automation, and clear operational handoffs.

Roles

Site Reliability Engineer
DevOps Engineer
Cloud Infrastructure Engineer
Platform Engineer
Cloud Operations Engineer
Infrastructure Automation Engineer

How I Help

Turn repeated operational work into versioned infrastructure, scripts, runbooks, and deployment workflows.

Troubleshoot production incidents across applications, databases, identity, networking, and cloud services.

Keep reliability work grounded in customer impact, not just platform abstraction.

Use AI-assisted tools to speed up planning, code review, debugging, and documentation while keeping source review, testing, and production judgment in place.

Technical Range

Tools and Work Areas

Infrastructure, observability, automation, production operations, and AI-assisted engineering work.

AWS Azure Terraform Terragrunt GitHub Actions Docker ECS Fargate Lambda S3 DMS CloudWatch New Relic Kinesis Firehose WAF ALB EC2 DataSync SQL Server Python PowerShell Bash Incident response Runbooks Codex AWS Kiro Claude Skills files

Contact

Useful for reliability and cloud infrastructure roles

This page is intended for roles that need practical cloud operations, infrastructure automation, observability, incident response, migration support, and production-grade documentation.

Contact Nate