Moses Sarpong

Platform & DevOps Engineer

Platform Engineer with 5+ years of experience designing and operating cloud-native infrastructure that stays reliable at scale — provisioning environments as code with Terraform, orchestrating workloads on Kubernetes, and directing Claude Code as a plan-reviewed agent on live infrastructure changes. I've built that discipline into tooling too: an AI PR reviewer running across ~80 repositories that catches real infrastructure issues before they merge.

5+years experience
CKADcertified
50+services in production
15 minalert → resolution
engineer.yaml
apiVersion: profile/v1
kind: Engineer
metadata:
name: moses-sarpong
location: remote-ready
spec:
experience: 5+ years
certifications:
- CKAD
focus:
- agentic-ai
- kubernetes
- terraform
- aws
- ci-cd
status: Open to opportunities

▸Resilience and repeatability aren't an afterthought — they're the platform. Every agent-drafted diff gets the same review bar as my own.

kind: Infrastructure

Live architecture

A snapshot of the real pipeline, not a live feed — click a node for the plan → diff → review behind it.

$ theme --set
🔀 GitHub OPERATIONAL
$branch → PR → review
Situation
Every infrastructure change — Terraform, Helm, CI config — starts as a branch, never a direct edit to main.
Action
Claude Code drafts the diff when I'm directing it as an agent; an AI PR reviewer I built takes the first pass — GitHub MCP, ~80 repos, ~30 infra PRs a week — then I review every line before it merges, same bar whether the agent wrote it or I did.
Result
A clean audit trail — every change traceable to a PR, a review, and a reason. The reviewer catches ~3 real issues a week (IAM widening, drift, unsafe Helm changes) before they merge.
→
⚙️ CI/CD DEPLOYED
$build → test → promote → gate
Situation
Manual deploys don't scale past a couple of services, and catching problems after production is expensive.
Action
GitHub Actions builds and tests in dev, auto-promotes to staging, and gates production behind manual approval. Validation — Terraform plan/validate, app tests, image checks, deployment health checks — happens as early as possible.
Result
Staging ships daily, production about twice a week, with OIDC federation instead of static AWS keys in the pipeline.
→
🏗️ Terraform OPERATIONAL
$plan → apply (agent-directed)
Situation
Directed Claude Code to add a VPC endpoint as part of the Gateway API migration (see Kubernetes).
Action
The agent didn't check what already existed and wrote a duplicate. terraform plan surfaced the conflict before apply — caught in review.
Result
Zero blast radius — never touched real infra. Changed the workflow afterward: the agent now checks existing state before writing a new resource.
→
☁️ AWS OPERATIONAL
$aws sts assume-role (OIDC)
Situation
CI/CD used long-lived static AWS access keys — a standing credential-exposure risk.
Action
Replaced them with OIDC federation. GitHub Actions now assumes a scoped IAM role via the cluster's OIDC provider on each run instead of storing a secret.
Result
No long-lived AWS credentials left in CI/CD, across the whole multi-tenant platform.
→
☸️ Kubernetes DEPLOYED
$helm upgrade --install
Situation
One shared, centrally-controlled ingress-nginx config routed every team's traffic — a routing problem for one team could affect all of them.
Action
Led a ~2-week migration to Gateway API + ALB, giving each team its own routing rules. Planned it stage by stage with Claude Code, diff reviewed before anything applied, coordinated with a 6-person team.
Result
Routing issues now stay scoped to the team that caused them. 6-cluster multi-tenant EKS platform, 50+ microservices, ~15 engineering teams, all on Terraform and Helm.
→
📊 Monitoring OPERATIONAL
$datadog alert → root-cause → resolve
Situation
Customers couldn't reach the application — a full external-facing outage, on the ALB from the Gateway API migration.
Action
A Datadog alert paged me directly. Reviewed the ALB's listener and routing rules and found requests being routed incorrectly.
Result
Fixed in about 15 minutes, alert to resolution. The dashboards and alerts themselves are self-built — I configure the thresholds and page routing. Since then I've also built an AI incident-diagnostic workflow that correlates pod state, deployment history, and Terraform changes to point at the likely root cause — a human still drives the fix.
kind: Skills

Technical toolkit

# agentic-ai
🤖

Agentic AI

Claude Code as a plan-reviewed agent on live Terraform/Helm changes — plus an AI PR reviewer (GitHub MCP, 4 specialized skills) and an incident-diagnostic workflow, both self-built

# aws
☁️

AWS

Cloud-native infrastructure across EC2, VPC, IAM, S3, and Load Balancers

# docker
🐳

Docker

Containerized app delivery with Docker images, Compose, and lifecycle management

# kubernetes
☸️

Kubernetes

Production workload orchestration — Deployments, Services, Ingress, Gateway API, Secrets

# terraform
🏗️

Terraform

Infrastructure as Code with reusable modules and remote state backends

# ci-cd
⚙️

CI/CD

End-to-end automation pipelines built on GitHub Actions

# monitoring
📊

Monitoring

Production observability with Datadog and CloudWatch — dashboards, alerting, and direct on-call paging

# scripting
💻

Programming

Automation scripting and tooling in Python, Bash, and Linux

kind: Certifications

Credentials

CKAD

Certified Kubernetes Application Developer

Cloud Native Computing Foundation (CNCF)

kind: Projects

Recent work

$kubectl apply -f gateway-api-migration.yaml

Gateway API Migration

Led a Terraform/Helm migration at vAuto from a single shared ingress-nginx config to per-team Gateway API + ALB routing — directed Claude Code through a plan-first, diff-reviewed process, shipped in ~2 weeks across a 6-person team.

$terraform apply -var-file=aws-infra.tfvars

EKS Platform on Terraform

Standardized a multi-tenant AWS platform across 6 EKS clusters and 50+ microservices at vAuto using reusable Terraform modules, serving about 15 engineering teams on a consistent provisioning model.

$claude-code run ai-pr-review

AI PR Reviewer

Built an AI PR reviewer with Claude Code and GitHub MCP across ~80 repositories — ~30 infra PRs a week, catching ~3 real issues weekly (IAM widening, drift, unsafe Helm changes) before merge. Backed by four specialized Claude Code skills: Terraform plan review, Kubernetes manifest linting, RBAC audit, and Gateway API/Helm validation. Cut review turnaround from ~1 day to same-day and saves ~5 hours a week.

$claude-code run incident-diagnostics

AI Incident-Diagnostic Workflow

Built an AI workflow that correlates Kubernetes pod state, GitHub/deployment history, and Terraform changes to point at the most likely production root cause — a human still drives the diagnosis and fix from there.

$datadog alert → root-cause → resolve

ALB Incident Response

Diagnosed a full customer-facing outage in ~15 minutes — Datadog alert to a misrouted ALB listener rule to fix — on the cluster from the Gateway API migration above.

kind: Contact

Get in touch

Open to Platform and DevOps Engineering roles — let's connect and build something reliable together.