Responsibilities
- Manage and maintain cloud infrastructure across AWS and Azure environments.
- Provision, configure, and monitor AWS services, including EC2 S3 EBS, EFS, IAM, VPC, CloudWatch, and Security Groups.
- Design, implement, and maintain CI/CD pipelines using Jenkins and other automation tools.
- Automate application deployment, testing, monitoring, and rollback processes.
- Build, optimise, and manage Docker-based deployments for development and production environments.
- Monitor infrastructure health, system performance, application uptime, and resource utilisation.
- Implement logging, monitoring, and alerting solutions to improve platform observability.
- Troubleshoot infrastructure, deployment, and networking issues across environments.
- Implement cloud security best practices, access controls, SSL/TLS certificates, firewall policies, and secrets management.
- Create and maintain infrastructure documentation, deployment guides, runbooks, and standard operating procedures.
- Collaborate with backend, frontend, QA, and ML teams to streamline deployment and operational workflows.
- Support AI/ML workloads running on cloud infrastructure, including GPU-based workloads.
- Participate in incident management, root cause analysis, and disaster recovery planning.
- Continuously identify opportunities for automation, reliability improvements, and cost optimisation.
Requirements
- 2-3 years of experience in DevOps, Cloud Engineering, or Site Reliability Engineering (SRE).
- Strong hands-on experience with AWS services: EC2 S3 EBS, EFS, IAM, VPC, CloudWatch, Security Groups
- Working knowledge of Microsoft Azure services.
- Strong experience with Jenkins for CI/CD automation.
- Hands-on experience with Docker and containerised deployments.
- Experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
- Strong Linux system administration skills.
- Experience with monitoring and logging tools.
- Strong understanding of networking concepts, including DNS, SSL/TLS, load balancers, reverse proxies, and firewalls.
- Familiarity with Python, Bash, or Shell scripting.
- Knowledge of Infrastructure as Code (IaC) concepts.
- Strong documentation and process management skills.
- Understanding of cloud security best practices and infrastructure hardening.
- Excellent debugging, troubleshooting, and problem-solving skills.
Good To Have
- Experience with Kubernetes.
- Experience with Terraform, Pulumi, or similar infrastructure as code tools.
- Experience managing GPU-based workloads.
- Experience with PostgreSQL, Redis, and distributed systems.
- Experience supporting AI/ML infrastructure and production workloads.
- Knowledge of backup, disaster recovery, and high-availability architectures.
- Exposure to FastAPI, microservices, or media processing platforms.
- Experience optimising cloud infrastructure costs across AWS and Azure.
This job was posted by Subrina Ahoy Lai from NeuralGarage.