Seventeen years.
One operating principle.
Make the system reliable. Make the failure understandable.
Make the recovery repeatable.

I started in 2008 supporting servers and V-SAT links at Chevron’s Escravos Gas-to-Liquids facility in Warri - Delta State. The work covered server infrastructure, LAN/WAN networks and field operations.
In Riyadh I led reliability work for an ambulance dispatch platform. That experience shaped the way I approach observability, runbooks and incident response. The person responding to an alert needs to understand the service and know what to do next.
Today my platform work brings Kubernetes, delivery automation and security together. I publish infrastructure packages with architecture notes and recovery guidance. I document validation boundaries so another engineer can assess what is ready and what still needs testing.
My frontier AI work applies the same discipline to evaluation. I build failure scenarios and independent verifiers that test whether coding agents can solve real engineering problems.
I’m based in North Carolina and looking for remote senior SRE or platform engineering work on systems where reliability matters.
The work behind the approach.
Remote
Snorkel AI
Expert Contributor · Software Engineering & Frontier AITask authoring and review associated with Senior SWE-Bench and Terminal-Bench 2.0, 2.1 and 3.0. Reproducible Docker environments with independent verifiers. Repository-scale review and grading calibration.
Remote
Greenstand
DevSecOps EngineerKubernetes and Argo CD delivery for TreeTracker. Security checks embedded into CI/CD. Multi-organization Fabric infrastructure and a Keycloak–Fabric identity bridge merged upstream.
Riyadh, Saudi Arabia
SAHAB
Senior Site Reliability EngineerReliability engineering for the Ahyeha emergency medical dispatch platform. Highly available infrastructure, incident command, disaster recovery exercises and production-readiness reviews.
Earlier: DevOps engineering at Stockbridge LLC (2017–2022) and Dantos Group (2014–2017). IT engineering at Chevron’s Escravos Gas-to-Liquids facility (2008–2014).
Where I contribute.
Kubernetes & platform engineering
Cluster lifecycle, networking, storage, RBAC, network policies and autoscaling. Bare-metal and cloud platforms. CKA certified.
Reliability & incident response
SLOs, alert routing, metrics, logs and distributed tracing. Root cause analysis, disaster recovery and capacity planning.
Security engineering
Image and IaC scanning, supply-chain checks and SBOMs. Identity, PKI, secrets lifecycle and runtime security.
Delivery & infrastructure as code
Argo CD, Flux, GitHub Actions and Jenkins. Helm, Kustomize, Terraform, OpenTofu and Ansible.
Systems & code
Python, Bash, TypeScript and SQL. Linux administration, distributed systems, RAFT, replication and highly available databases.