Cloud & Infrastructure Site Reliability Engineer (SRE)
Role Overview
We are looking for a Cloud & Infrastructure Site Reliability Engineer (SRE) to build, operate, automate, and continuously improve highly available cloud-native infrastructure powering SaaS platform. You will work across public cloud, Kubernetes, compute, storage, networking, observability, and infrastructure automation - with a strong focus on reliability, scalability, security, and reducing manual operations. The ideal candidate combines strong infrastructure fundamentals with a software engineering and automation mindset - someone who proactively identifies operational issues and engineers permanent solutions rather than relying on repetitive manual processes.
Key Responsibilities
- Cloud & Infrastructure Operations: Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP - including compute, storage, networking, and containers. Handle provisioning, upgrades, patching, and capacity management; resolve connectivity, DNS, and load-balancing issues.
- Site Reliability Engineering: Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response, root-cause analysis, and post-incident reviews. Automate away recurring operational work and improve system resilience through redundancy and automated recovery.
- Kubernetes & Microservices: Deploy, manage, and troubleshoot containerized microservices on Kubernetes (EKS, AKS, GKE), including ingress, storage, secrets, and cluster health.
- Infrastructure as Code & Automation: Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash; build reusable IaC modules integrated into CI, CD.
- Monitoring & Observability: Implement monitoring, logging, tracing, and alerting (Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, etc.) with meaningful dashboards and low-noise alerting.
- Networking & Security: Troubleshoot cloud networking (VPC, VNet, DNS, VPN, firewalls, load balancers); apply IAM, RBAC, secrets, and least-privilege security practices.
- Incident & Production Support: Participate in on-call rotations, own incidents end-to-end, and maintain runbooks and automated remediation procedures.
- Cost & Performance Optimization: Identify underutilized resources and support FinOps initiatives (rightsizing, scheduling, lifecycle management) while balancing cost with reliability.
Required Skills
- Strong hands-on experience with at least one major public cloud - AWS, Azure, or GCP.
- Solid understanding of Linux administration and infrastructure fundamentals.
- Experience with Kubernetes and containerized microservices.
- Experience with Infrastructure as Code (Terraform and, or Ansible).
- Strong automation, scripting skills in Python and Bash (PowerShell a plus).
- Good understanding of networking, DNS, load balancing, firewalls, and cloud routing.
- Experience with monitoring, logging, alerting, and observability platforms.
- Understanding of CI, CD, Git, APIs, and modern DevOps practices.
- Experience troubleshooting distributed production environments.
- Understanding of security, IAM, RBAC, secrets, certificates, and infrastructure security.
- Familiarity with database operations a plus.
Preferred Skills
Experience & Qualifications
- Experience operating hybrid-cloud or multi-cloud environments.
- Experience with VMware, Nutanix, OpenShift, or similar private-cloud technologies.
- Experience building automated remediation or self-healing workflows.
- Familiarity with AIOps, event correlation, anomaly detection, and AI-assisted operations.
- Understanding of FinOps and cloud cost optimization.
- Experience with ITSM platforms such as ServiceNow.
- Cloud, Kubernetes, Linux, or Terraform certifications (AWS, Azure, GCP) are a strong plus. SRE Mindset
- Automate before repeating, and engineer permanent fixes instead of resolving the same incident twice.
- Treat infrastructure and operational configuration as code.
- Design for failure and recovery, and measure reliability using data.
- Build reusable automation and operational standards.
- Reduce operational toil and unnecessary tickets.
- Continuously improve reliability, performance, security, and cost efficiency.
- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70%.
- SRE Engineer - 2–6 years of relevant Cloud, Infrastructure, DevOps, SRE experience.
- Cloud certifications (AWS, Azure, or GCP) are a strong plus. Roles and responsibilities
- Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP. Handle provisioning, upgrades, patching, and capacity management.
- Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response.
- Automate away recurring operational work.
- Deploy, manage, and troubleshoot containerized microservices on Kubernetes.
- Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash.
- Implement monitoring, logging, tracing, and alerting.
- Troubleshoot cloud networking and apply security practices.
- Participate in on-call rotations and own incidents end-to-end.
- Identify underutilized resources and support FinOps initiatives. Experience and education
- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70% Location Remote