Search Jobs

Search by job, company or skills

Senior Site Reliability Engineer (Azure)

Senior Site Reliability Engineer (Azure)

Salt
  • Posted 19 hours ago
  • Be among the first 10 applicants

Job Description

Location: Abu Dhabi, UAE

Employment: 12-months initial (contract to perm)

We are seeking a hands-on Senior DevOps / Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications, workflow platforms and AI-enabled products.

This is not primarily an infrastructure-administration role. You will work directly with software engineers to create reliable, secure and automated paths from code to production.

You will own deployment automation, runtime reliability, observability, infrastructure-as-code and operational readiness across applications that integrate with critical enterprise systems.

What You Will Own

Platform & Infrastructure

* Build and maintain cloud infrastructure using Infrastructure as Code.

* Design secure, repeatable environments across development, test, staging and production.

* Manage containerized workloads and Kubernetes-based deployments where appropriate.

* Define standard application deployment patterns for backend, frontend and AI services.

* Implement secure secrets and configuration management.

* Support network, identity and connectivity requirements for enterprise integrations.

CI/CD & Developer Productivity

* Build automated CI/CD pipelines.

* Standardize build, test, security scanning and deployment processes.

* Automate environment provisioning and configuration.

* Reduce manual deployment steps and production configuration drift.

* Work closely with engineering teams to improve release frequency and reliability.

Reliability & Observability

* Establish logging, metrics, tracing and alerting.

* Define service-level indicators and operational thresholds.

* Build dashboards for system health and application performance.

* Implement incident-response and production-support practices.

* Design for graceful degradation, retries, failover and recovery.

* Lead root-cause analysis of production incidents.

Security & Operational Controls

* Implement least-privilege access and secure deployment patterns.

* Support auditability of infrastructure and production changes.

* Integrate security checks into delivery pipelines.

* Work with security and infrastructure teams to meet enterprise control requirements.

Resilience

* Support business continuity and disaster-recovery design.

* Define backup, restore and recovery procedures.

* Test operational recovery rather than relying solely on documented plans.

Required Experience

* 6+ years in DevOps, SRE, platform engineering or cloud infrastructure.

* Strong production experience with Azure, AWS or GCP; Azure strongly preferred.

* Docker and Kubernetes.

* Infrastructure as Code using Terraform, Bicep, Pulumi or equivalent.

* CI/CD using Azure DevOps, GitHub Actions, GitLab CI or similar.

* Strong Linux and networking fundamentals.

* Observability tooling and distributed-system troubleshooting.

* Secure secrets, identity and access-management patterns.

* Production incident-management experience.

* Scripting/programming capability in Python, Go, Bash or equivalent.

Strong Advantage

* Azure Kubernetes Service.

* Azure Service Bus, API Management, Key Vault and related Azure services.

* Enterprise integration platforms.

* SAP-connected environments.

* AI/LLM application deployment.

* Regulated or government environments.

* High-availability and disaster-recovery architecture.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Identity and access-management

Infrastructure as Code

Distributed-system troubleshooting

CI CD

GitLab CI

GitHub Actions

Observability tooling

Bicep

Networking fundamentals

About Company