← All case studiesCase study 01 / Ahyeha

Reliability when every second counts.

Site reliability at SAHAB for Ahyeha’s digital ambulance transport and emergency medical services in Saudi Arabia.

Role
Senior Site Reliability Engineer
Organization
SAHAB · Riyadh
Scope / period
June 2022 – November 2023
Explore Ahyeha

The context

Ahyeha brings ambulance transport and emergency medical coverage into one digital platform for people and organizations across Saudi Arabia. The experience includes online requests, real-time tracking and flexible online payment.

Its services cover patient transport and medical coverage for events. The product is designed to make access to care easier through a secure, user-focused digital experience.

My work at SAHAB covered the operational infrastructure behind that experience. Reliability meant keeping dispatch services available and making incidents understandable to the people responsible for restoring them.

Illustrative reliability flow: dispatch traffic through ingress to services with database and observability branches.DispatchIngressServicesDatabaseTelemetry
Simplified functional view of the reliability scope. This is an explanatory diagram rather than a deployment inventory.

My contribution

Designed and maintained highly available cloud and Kubernetes environments. The scope included load balancing, DNS, autoscaling, deployment automation and database high availability.

Built metrics, logging and alerting coverage across distributed services. Led incident command and root cause analysis so operational issues could be investigated across service boundaries.

Drove disaster recovery exercises, capacity planning and production-readiness reviews before major releases.

Engineering decisions

01. Prepare before release

Production-readiness reviews connected deployment work with the operating conditions the service would face.

02. Make alerts actionable

Sharper alert routing and practical runbooks helped the on-call team identify the affected service and choose a recovery path.

03. Close the incident loop

Post-incident actions connected root cause analysis with the next reliability improvement.

The outcome

Improved incident response through clearer alert routing, actionable runbooks and tracked follow-up work. The focus was a platform the team could operate and recover with confidence.

Tools & evidence

KubernetesDockerCI/CDDNSLoad balancingDatabase HAMonitoringIncident response

The public product site describes Ahyeha’s services. Engineering responsibilities here are drawn from my professional experience.

NEXT CASE STUDYSnorkel AI benchmarks
Let's start a conversation

What are you trying
to keep running?

Have a role or a project in mind?
Tell me what you’re working on.

North Carolina, USAOpen to remote opportunities

Send me a message

Minimum 10 characters.

Privacy Policy