Search by job, company or skills

  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Role: Platform SRE + AI

Experience: 7–10 yrs

Location: Bangalore

Senior Platform SRE:

  • Owns reliability and performance of the platform end to end — leads incident response, drives cluster and runtime tuning standards, and sets direction for the AI/MCP platform.
  • Expected to mentor and influence vendor and stakeholder decisions.
  • Looking for a strong AWS & Kubernetes (EKS) expert to manage production-scale cloud infrastructure, platform reliability, performance tuning, and automation.
  • Experience with AI/LLM platform operations, monitoring, troubleshooting, and cloud-native environments is highly preferred.
  • This is a deep platform and performance-engineering role — owning production Kubernetes, AWS infrastructure, and the reliability of LLM-backed and agentic services in a regulated environment.
  • It goes well beyond standard SRE: we are looking for an engineer who tunes clusters, diagnoses memory and runtime behaviour, and operates AI platform workloads at production scale

Mandatory Skills:

  • Deep AWS expertise — End-to-end production deployment on AWS is a must — EC2, ECS, EKS, Lambda, S3, IAM, RDS, API Gateway, VPC, EFS, SNS, SQS, EventBridge, CodeBuild.
  • Kubernetes (deep & mandatory) — Production EKS ownership — cluster configuration (CPU/ RAM sizing, node groups, autoscaling), workload fine-tuning (requests/limits, HPA/VPA, eviction policies), and hands-on Helm chart management.
  • Memory expertise — Container vs. runtime memory models, OOMKill diagnosis, cgroup behaviour, JVM/Python heap tuning, and database buffer-pool configuration.
  • Serverless fine-tuning — Lambda memory/concurrency/cold-start optimization, ECS/Fargate task sizing, and serverless cost-performance tradeoffs.
  • Database operations — Running and tuning RDS or equivalent in production; StatefulSet-based DB deployments in K8s a strong plus.
  • AI platform & MCP environment — Hands-on deployment or operation of LLM-backed services, MCP servers, or agentic pipelines on cloud infrastructure.

Core Requirements:

  • 5–10 years in software development, DevOps, SRE or production support roles.
  • Working knowledge of GCP and/or Azure — multi-cloud integrations, platform differences, hybrid workloads.
  • CI/CD with GitHub Actions, Jenkins or GitLab.
  • Scripting proficiency in Python and/or Bash; YAML fluency.
  • Terraform and/or CloudFormation for infrastructure provisioning.
  • ITIL framework across Incident, Change, Problem and CAPA management.
  • Monitoring with Splunk and/or Grafana — including infra-level resource and memory dashboards.
  • ServiceNow and JIRA; strong ITSM discipline.
  • Bachelor's degree or equivalent practical experience

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151958243

Similar Jobs

Bengaluru, India

Skills:

Aws LambdaS3RDSCloudformationPrometheusGrafanaElk StackCloudwatchTerraformAnsibleECSHelmKubernetesPythonApi GatewayAWSKarpenterChefGoIstioEKS

Bengaluru, India

Skills:

ElkPrometheusGrafanaJenkinsTerraformPythonKubernetesInfrastructure as CodeGKEGoAKSEKSGitLab CIGitHub ActionsOpenTelemetryArgoCD

Bengaluru, India

Skills:

ARM TemplatesOracle DatabasePrometheusGrafanaRmanDockerTerraformAshMicrosoft AzureAwrPythonData GuardAzure DevOpsPowerShellBashAddmOemKubernetesKey VaultGoLog AnalyticsAzure RBACGitHub ActionsDefender for CloudMicrosoft SentinelBicepApplication InsightsManaged IdentitiesAzure PolicyAzure MonitorPrivate Endpoints

Bengaluru

Skills:

GithubSamlPostgreSQLPythonAws

Bengaluru, India

Skills:

SQL ServerElkGrafanaDatadogArmPrometheusMySQLCloudformationOracleKubernetesBashPythonTerraformPostgreSQLGCP Cloud SQLAWS RDS AuroraAlloyDBGoAzure Database ServicesAzure SQL MICI CD pipelines

Beware of Scammers

We don’t charge money for job offers