Platform & DevOps Engineer
Platform Engineer with 5+ years of experience designing and operating cloud-native infrastructure that stays reliable at scale — provisioning environments as code with Terraform, orchestrating workloads on Kubernetes, and directing Claude Code as a plan-reviewed agent on live infrastructure changes. I've built that discipline into tooling too: an AI PR reviewer running across ~80 repositories that catches real infrastructure issues before they merge.
▸Resilience and repeatability aren't an afterthought — they're the platform. Every agent-drafted diff gets the same review bar as my own.
A snapshot of the real pipeline, not a live feed — click a node for the plan → diff → review behind it.
terraform plan surfaced the conflict before apply — caught in review.Claude Code as a plan-reviewed agent on live Terraform/Helm changes — plus an AI PR reviewer (GitHub MCP, 4 specialized skills) and an incident-diagnostic workflow, both self-built
Cloud-native infrastructure across EC2, VPC, IAM, S3, and Load Balancers
Containerized app delivery with Docker images, Compose, and lifecycle management
Production workload orchestration — Deployments, Services, Ingress, Gateway API, Secrets
Infrastructure as Code with reusable modules and remote state backends
End-to-end automation pipelines built on GitHub Actions
Production observability with Datadog and CloudWatch — dashboards, alerting, and direct on-call paging
Automation scripting and tooling in Python, Bash, and Linux
Cloud Native Computing Foundation (CNCF)
Led a Terraform/Helm migration at vAuto from a single shared ingress-nginx config to per-team Gateway API + ALB routing — directed Claude Code through a plan-first, diff-reviewed process, shipped in ~2 weeks across a 6-person team.
Standardized a multi-tenant AWS platform across 6 EKS clusters and 50+ microservices at vAuto using reusable Terraform modules, serving about 15 engineering teams on a consistent provisioning model.
Built an AI PR reviewer with Claude Code and GitHub MCP across ~80 repositories — ~30 infra PRs a week, catching ~3 real issues weekly (IAM widening, drift, unsafe Helm changes) before merge. Backed by four specialized Claude Code skills: Terraform plan review, Kubernetes manifest linting, RBAC audit, and Gateway API/Helm validation. Cut review turnaround from ~1 day to same-day and saves ~5 hours a week.
Built an AI workflow that correlates Kubernetes pod state, GitHub/deployment history, and Terraform changes to point at the most likely production root cause — a human still drives the diagnosis and fix from there.
Diagnosed a full customer-facing outage in ~15 minutes — Datadog alert to a misrouted ALB listener rule to fix — on the cluster from the Gateway API migration above.
Open to Platform and DevOps Engineering roles — let's connect and build something reliable together.