SRE (Site Reliability Engineer)
- Posted 12 days ago
- Be among the first 10 applicants
Job Description
Responsible for ensuring application and infrastructure reliability, availability, performance, and operational efficiency through monitoring, automation, and incident management.
Required Skills *
Linux administration, cloud platforms (AWS/Azure/GCP), Kubernetes & Docker, monitoring and observability, scripting (Python/Bash), CI/CD, networking, incident management, automation, troubleshooting, and reliability engineering practices (SLA/SLO/SLI).
Required Experience / Knowledge Areas *
Required Skills *
Linux administration, cloud platforms (AWS/Azure/GCP), Kubernetes & Docker, monitoring and observability, scripting (Python/Bash), CI/CD, networking, incident management, automation, troubleshooting, and reliability engineering practices (SLA/SLO/SLI).
Required Experience / Knowledge Areas *
- Strong experience in Linux/Unix administration, troubleshooting, and system performance tuning.
- Hands-on expertise with cloud platforms such as AWS, Azure, or GCP.
- Proficiency in Kubernetes, Docker, automation, and CI/CD pipelines.
- Experience with monitoring, observability, incident management, and root cause analysis.
- Preferred experience in banking, financial services, fintech, or enterprise application
- Strong understanding of networking, security, and infrastructure management.
- Experience in incident management, root cause analysis (RCA), and problem resolution.
- Knowledge of high availability, disaster recovery, SLA/SLO/SLI, and performance optimization.
- Experience supporting banking, financial services, fintech, or enterprise application environments is preferred.
More Info
Key Skills
monitoring and observability
cloud platforms
CI CD
reliability engineering practices
SLI
SLO
