Search by job, company or skills

SRE Lead DBaaS Platform

Early Applicant
  • Posted 2 months ago
  • Be among the first 10 applicants

Job Description

Job Title: SRE Lead – DBaaS Platform

Role Overview

We are seeking an experienced Site Reliability Engineering (SRE) Lead to strengthen

production reliability ownership for our Database-as-a-Service (DBaaS) platform. This role

will bring hyperscaler-grade (RDS-level) operational expertise to drive deep product

debugging, reliability engineering, and Dev collaboration across cloud-native database

services.

The SRE Lead will own platform stability, availability, performance, and incident excellence

across Azure/AWS/GCP-hosted database workloads.

Location :- Hyderabad

Department :- Customer Success

Reporting :- Senior Director Customer Success/SRE

Key Responsibilities

  • Production Reliability Ownership

Own end-to-end reliability, availability, and performance of the DBaaS platform.

Define and enforce SLIs, SLOs, and SLAs across all supported database engines.

Lead production incident response (P1/P2), RCAs, and long-term resilience

improvements.

Drive error budget governance with Engineering and Product teams.

  • Hyperscaler-Level Operational Excellence

Bring RDS/Cloud SQL/Azure SQL Managed Instance operational patterns into the

platform.

Implement automation-first operations (self-healing, auto-remediation, failover

orchestration).

Standardize HA/DR architectures across multi-region deployments.

Improve backup reliability, replication integrity, and failover predictability.

  • Deep Product Debugging & Dev Collaboration

Partner with Product Engineering for deep database engine-level debugging.

Troubleshoot complex performance bottlenecks (IO, CPU, locking, replication lag).

Support root cause analysis involving cloud infrastructure, storage, networking, and

database internals.

Influence platform architecture for operability and reliability.

  • Observability & Reliability Engineering

Build unified observability across DBaaS (metrics, logs, traces).

Define golden signals for database reliability.

Improve proactive anomaly detection and capacity forecasting.

Drive chaos testing and resilience validation practices.

  • Automation & Platform Hardening

Lead reliability automation (runbooks code).

Improve provisioning, patching, upgrade, and scaling reliability.

Standardize configuration management and drift detection.

Enhance security posture aligned to enterprise compliance needs.

  • DevOps & Platform Governance

Champion SRE best practices across engineering teams.

Establish production readiness review frameworks.

Define release reliability gates for DBaaS components.

Mentor junior SREs and build a reliability-first culture.

Technical Requirements

Cloud Platforms (Mandatory – Multi-Cloud Preferred)

Deep hands-on experience with:

  • AWS RDS / Aurora
  • Azure SQL MI / Azure Database Services
  • GCP Cloud SQL / AlloyDB

Strong understanding of cloud networking, storage, IAM, HA architectures.

Database Expertise

Strong operational knowledge of:

  • Oracle
  • PostgreSQL
  • MySQL
  • SQL Server

Experience handling large-scale production databases (TB+ workloads).

Performance tuning, replication troubleshooting, and backup recovery validation.

SRE & Platform Skills

Strong scripting: Python / Bash / Go.

Infrastructure as Code (Terraform / ARM / CloudFormation).

CI/CD pipelines and release automation.

Observability stack (Prometheus, Grafana, ELK, Datadog, etc.).

Kubernetes exposure preferred.

Leadership Expectations

10+ years overall experience, 5+ in SRE/Platform roles.

Prior experience in hyperscaler environments or cloud-native SaaS products.

Strong incident leadership and executive communication skills.

Ability to influence cross-functional stakeholders.

Experience Building And Leading SRE Teams Preferred.

Success Metrics (First 12 Months)

Reduction in P1/P2 incidents by X%.

Improved MTTR by X%.

Defined SLO framework implemented across all DBaaS services.

Automation coverage >70% of repeat operational tasks.

Zero critical audit non-compliance findings.

Why Join Us

Opportunity to build hyperscaler-grade DBaaS reliability.

Direct impact on mission-critical enterprise workloads.

Multi-cloud platform engineering exposure.

High visibility role working with Product, Engineering, and Leadership.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 147523021

Beware of Scammers

We don’t charge money for job offers