

Search by job, company or skills

Location: Bengaluru
Internal Team: OnPrem
Years of experience: 2-5
About the Role
We are seeking an SRE to manage release operations and infrastructure reliability across 100+ Kubernetes clusters. You will drive automated deployments, build observability, troubleshoot complex network/system issues, and create production runbooks to minimize downtime. In addition to core reliability work, you will contribute directly to internal coding projects, build custom tooling, and assist customers during escalations.
Key Responsibilities
Requirements
Must-Have
Nice-to-Have
Job ID: 151897293
Skills:
Elk, Prometheus, Networking Concepts, Grafana, Datadog, Docker, Terraform, Shell scripting, Python, Azure DevOps, AWS, Cloudformation, Bash, New Relic, Jenkins, Git, Gcp, Ansible, Incident Management, Splunk, Azure, Kubernetes, GitHub Actions, Linux Unix administration, Production Support, Rca, GitLab CI
Skills:
Storm, Cassandra, Prometheus, Kafka, Terraform, Docker, Elasticsearch, Shell scripting, Postgres, Gitlab, Python, AWS, Rust, Cloudformation, Redis, Jenkins, Cloudwatch, Gcp, Linux, Ansible, Spark, Kubernetes, Go, Flink, ArangoDB, GitHub Actions, Stackdriver
Skills:
Splunk, Prometheus, Aws S3, Kubernetes, Python, Bash, Grafana, Terraform, Docker, Go
Skills:
Monitoring Tools, Cloud Infrastructure, Devops, Aws, Automation, SRE
Skills:
Gcp, Datadog, Prometheus, Azure, Terraform, Grafana, Jenkins, Ansible, GitHub Actions, AI-Ops, GCP Operations Suite, Azure Monitor