Location: Bengaluru
Job Title: Senior Engineer
Budget: Upto INR 90 lacs (Fixed) + (Variable), based on experience and fit
Industry: B2B SaaS Product Company (Prior experience with IT Product company is a must)
About the Role
We are looking for a hands-on Senior MLOps / LLMOps Engineer with 8 to 12 years of experience to help operationalize Machine Learning and Generative AI capabilities across enterprise products and platforms.
This role is focused on MLOps and LLMOps, including model lifecycle management, deployment automation, prompt and model governance, evaluation pipelines, observability, monitoring, controlled rollout, release governance, and production support.
You will help establish the operational foundation that enables AI Engineers, Data Scientists, Platform Teams, and Product Engineering teams to deploy, monitor, evaluate, govern, and support AI/ML and LLM-powered capabilities in production through secure, repeatable, and auditable operational practices.
The ideal candidate should have strong experience with Cloud Engineering, DevOps, MLOps, Azure Machine Learning, MLflow, Azure OpenAI, CI/CD automation, AI governance, monitoring, and production support, along with programming experience in Python and/or .NET/C#.
Key Responsibilities:
ML Ops Engineering:
- Build and maintain ML Ops pipelines for model packaging, validation, registration, deployment, monitoring, rollback, and lifecycle management.
- Standardize model tracking, versioning, approval workflows, promotion gates, release evidence, and lifecycle governance using Azure Machine Learning, MLflow, or similar platforms.
- Support deployment of model endpoints, inference services, batch scoring jobs, and AI/ML service artifacts across environments.
- Create repeatable deployment patterns for moving models from experimentation to production.
- Implement CI/CD and release automation for AI/ML workloads using Azure DevOps, GitHub Actions, or similar tools.
- Support controlled promotion of models across Development, Test, Staging, and Production environments.
- Partner with Data Scientists and AI teams to ensure models are production-ready.
- Define operational runbooks, rollback procedures, support processes, and release checklists.
LLM Ops Engineering:
- Build and maintain LLM Ops practices for prompt lifecycle management, versioning, testing, deployment, rollback, and governance.
- Manage operational workflows for prompts, model configurations, evaluation datasets, retrieval configurations, and AI service settings.
- Support evaluation and monitoring of LLM outputs for correctness, relevance, groundedness, hallucination risk, latency, safety, and cost.
- Implement prompt and model comparison workflows across environments.
- Support safe rollout strategies using feature flags, canary deployments, A/B testing, approval gates, and human-in-the-loop controls.
- Track token usage, latency, failed requests, prompt regressions, retrieval quality, and operational trends.
- Help establish Responsible AI controls including auditability, traceability, explainability, privacy, security, and compliance.
AI Governance & Operational Excellence:
- Establish standardized operational practices for AI/ML and LLM systems.
- Ensure models, prompts, datasets, deployment artifacts, and configurations remain version-controlled and traceable.
- Capture release evidence including model versions, prompt versions, approvals, testing, rollback plans, and production readiness.
- Support separation of duties and approval workflows for production deployments.
- Collaborate with Security, Governance, Architecture, Platform, Privacy, Risk, and Audit teams.
- Develop reusable standards for AI governance, production readiness, incident response, and operational excellence.
Monitoring & Observability:
- Build evaluation pipelines for both traditional ML models and LLM-powered applications.
- Implement monitoring for model performance, service health, latency, failures, cost, output quality, and operational reliability.
- Create dashboards, alerts, traces, and reporting for AI services.
- Monitor model drift, prompt regressions, retrieval quality, and production anomalies.
- Support incident management, root cause analysis, and continuous operational improvements.
Production Readiness:
- Define and operationalize AI/ML production readiness standards.
- Implement automated quality gates for model validation, prompt regression testing, security checks, and deployment approvals.
- Support phase-gated rollout strategies for AI capabilities.
- Ensure regulated or sensitive AI use cases comply with governance, privacy, and security requirements before production deployment.
Cloud & DevOps:
- Work extensively with Azure Machine Learning, Azure OpenAI, Azure AI Services, Azure Storage, Azure App Services, Azure Functions, AKS, and related Azure services.
- Automate deployments, configuration management, and environment promotion.
- Support containerized deployments using Docker and Kubernetes.
- Integrate AI operational workflows with enterprise CI/CD, monitoring, logging, audit, and incident management platforms.
- Ensure AI workloads meet enterprise standards for reliability, scalability, security, observability, and cost optimization.
Cross-Functional Collaboration:
- Partner with Data Scientists, AI Engineers, Platform Engineering, Product Engineering, Security, Architecture, Governance, and Product Management teams.
- Guide teams in operationalizing AI solutions from experimentation to production.
- Participate in sprint planning, release planning, production readiness reviews, and operational governance forums.
- Provide technical leadership on MLOps, LLMOps, deployment automation, monitoring, governance, and operational excellence.
Required Qualifications:
- 8-12 years of experience in MLOps, DevOps, Cloud Engineering, Platform Engineering, AI Operations, or related disciplines.
- Strong experience deploying, monitoring, and supporting production AI/ML platforms.
- Hands-on experience with Azure Machine Learning, MLflow, model registries, and model lifecycle management.
- Experience with Azure OpenAI, Azure AI Services, OpenAI APIs, or similar LLM platforms.
- Strong understanding of LLMOps, prompt lifecycle management, AI observability, token usage, hallucination monitoring, and safe deployment practices.
- Programming experience in Python and/or .NET/C#.
- Experience with Azure DevOps, GitHub Actions, or similar CI/CD platforms.
- Strong knowledge of Docker, Kubernetes, cloud-native deployments, monitoring, logging, dashboards, alerting, and operational support.
- Familiarity with secure release practices, audit logging, access controls, governance, and production operations.
- Strong collaboration and communication skills.
Preferred Qualifications:
- Experience operationalizing Generative AI applications including LLMs, RAG, embeddings, chatbots, summarization, classification, document processing, and workflow automation.
- Experience with Azure AI Search, Azure AI Document Intelligence, Azure Functions, Azure App Services, and AKS.
- Experience with ML frameworks such as Scikit-learn, XGBoost, PyTorch, or TensorFlow.
- Exposure to Semantic Kernel, LangChain, LlamaIndex, AutoGen, CrewAI, or similar orchestration frameworks.
- Experience with evaluation frameworks, prompt regression testing, AI quality dashboards, and golden datasets.
- Exposure to Terraform, RBAC, secrets management, audit logging, and cloud security best practices.
- Experience working in regulated industries such as Healthcare, Insurance, FinTech, or Financial Services.
- Knowledge of Responsible AI, privacy, compliance, and governance practices.