Ideal candidate for this role is a Senior DevOps / Platform Engineer with 6+ years of experience and strong hands-on expertise in managing on-premises, self-hosted infrastructure and CI/CD platforms. The candidate should have solid production experience with OpenShift, Jenkins/CloudBees, Prometheus, Grafana, ELK, and Linux, along with strong Python backend development and system-design skills. The role requires someone who can independently design, operate, troubleshoot, and optimize infrastructure while building reliable, scalable, and highly available platforms.
Responsibilities:
- Design and implement CI/CD pipelines using Jenkins/CloudBees, incorporating artifact management and code-quality gates.
- Architect, administer, and maintain CI/CD tooling infrastructure, including Jenkins controllers and agents, artifact repositories, and code-quality servers.
- Manage Jenkins/CI-CD platform upgrades, plugins, configurations, backups, high availability, and access controls.
- Operate and maintain OpenShift clusters in production, including capacity planning, upgrades, RBAC, networking, ingress, and workload optimization.
- Own infrastructure monitoring using Prometheus and Grafana, including PromQL, dashboards, SLOs, and alert tuning.
- Set up and maintain ELK for centralized application logging, including log parsing, indexing, and retention.
- Drive resource and infrastructure optimization by analyzing utilization, rightsizing workloads, and providing recommendations to stakeholders.
- Develop custom backend services, REST APIs, and integrations using Python.
- Design and maintain PR-DR capabilities for critical infrastructure, including replication, failover/failback, and periodic DR drills.
- Manage supporting databases, including installation, backup/restore, replication, monitoring, and performance tuning.
- Troubleshoot issues across the technology stack, including CI/CD failures, OpenShift, storage, networking, and infrastructure issues.
- Develop and maintain technical runbooks and documentation.
- Mentor engineers and contribute to improving platform engineering practices.
Requirements:
- 6+ years of experience in DevOps, Platform Engineering, or a related role.
- Strong hands-on experience administering OpenShift in production environments.
- Strong experience with Jenkins/CloudBees, including declarative pipelines and Groovy shared libraries.
- Good understanding of CI/CD platform architecture, including HA, distributed build agents, and scalability.
- Hands-on experience implementing and managing PR-DR, including replication, failover/failback, and RTO/RPO requirements.
- Database administration experience with PostgreSQL, MySQL, or similar, including backups, restores, and replication.
- Experience managing artifact repositories and code-quality gates within CI/CD pipelines.
- Strong hands-on experience with Prometheus, Grafana, and ELK.
- Strong Python development experience, including backend services, REST APIs, and integrations.
- Strong system-design skills with an understanding of scalability, reliability, and failure handling.
- Good experience with Linux administration and shell scripting.
- Working knowledge of networking, TLS, secrets management, and security hardening.
- Exposure to Infrastructure as Code and GitOps, such as Ansible, Helm, Argo CD, or similar tools.
- Strong troubleshooting and analytical skills with the ability to work across infrastructure and application layers.