Search Jobs

Search by job, company or skills

Lead Machine Learning Engineer

Lead Machine Learning Engineer

weekday (yc w21)
Early Applicant
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

This role is for one of our clients
Industry: Software Development

Seniority level: Mid-Senior level

Min Experience: 9+ years

Location: Bengaluru

JobType: full-time

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

We are seeking a hands-on Lead Machine Learning Engineer to design, build, and scale production-grade Generative AI and Machine Learning applications. The role will focus on developing AI-powered assistants, retrieval and reasoning systems, agentic workflows, document intelligence, decision-support solutions, and intelligent automation capabilities that improve productivity, service quality, customer experience, and business outcomes.

This is a technical leadership role for an engineer who has moved beyond experimentation and prototypes and has proven experience taking AI applications through the last mile into production. You will be responsible for ensuring AI systems are reliable, observable, secure, cost-efficient, measurable, and trusted by users.

You will work closely with Product, Engineering, Design, Security, Compliance, Operations, and business stakeholders to identify high-impact AI opportunities, make pragmatic architecture decisions, and deliver production-ready AI experiences at scale.

Requirements
Key Responsibilities

Build AI Solutions for Business Impact

  • Design, build, and launch GenAI-powered applications including AI assistants, copilots, document intelligence, workflow automation, and decision-support solutions
  • Identify high-impact opportunities where AI can improve productivity, operational efficiency, service quality, customer experience, and business outcomes
  • Take AI applications from concept through production, collaborating with Product, Engineering, Design, Security, and business teams
  • Lead hands-on technical execution across application architecture, model selection, prompt engineering, retrieval, orchestration, APIs, data pipelines, and user-facing experiences
  • Translate business requirements into scalable and measurable machine learning and AI solutions
  • Establish success metrics and continuously optimize solutions based on real-world user feedback and business impact

Build Enterprise-Grade AI Systems

  • Architect reliable GenAI applications using modern approaches such as RAG, agentic workflows, tool use, structured outputs, retrieval, grounding, and fine-tuning where appropriate
  • Design systems that effectively combine frontier models, open-source models, smaller task-specific models, and deterministic components based on the specific use case
  • Develop strong grounding mechanisms using enterprise knowledge and relevant business data
  • Build production systems with appropriate observability, monitoring, versioning, fallback mechanisms, security, privacy, and operational ownership
  • Design for reliability, scalability, latency, cost efficiency, and maintainability
  • Stay current with advances in AI/ML and apply emerging techniques pragmatically where they deliver meaningful improvements

Evaluation, Quality & LLMOps

  • Define practical evaluation frameworks for GenAI applications covering accuracy, relevance, groundedness, safety, latency, cost, user trust, adoption, and business impact
  • Establish automated and human-in-the-loop evaluation processes for AI applications
  • Use LLM evaluation and observability platforms such as LangFuse, Arize, or similar tools
  • Monitor production performance and identify opportunities to improve model quality, reliability, and efficiency
  • Establish appropriate safeguards, fallback paths, and quality controls for production AI systems

Technical Leadership

  • Provide technical leadership across the AI/ML application development lifecycle
  • Make pragmatic architecture and technology decisions while balancing quality, speed, security, and cost
  • Mentor engineers and contribute to engineering standards, best practices, and technical direction
  • Partner with cross-functional teams to ensure AI solutions are usable, secure, reliable, and aligned with business objectives
  • Take ownership of production outcomes, including launch quality, reliability, user feedback, adoption, and measurable impact

Required Experience & Qualifications

  • 8+ years of experience building applied AI/ML-based intelligent software systems
  • 2+ years of practical Generative AI application experience
  • At least one production GenAI application that has been deployed to real users at meaningful scale
  • Proven experience taking GenAI solutions beyond PoC/prototype into production
  • Strong ownership of production quality, reliability, cost optimization, user feedback, adoption, and measurable business impact
  • Strong understanding of designing LLM applications using an appropriate combination of:
    • RAG
    • Agentic workflows
    • Tool use
    • Structured outputs
    • Retrieval and grounding
    • LLM orchestration
    • Frontier and open-source models
    • Fine-tuning
    • Task-specific models
    • Deterministic systems

  • Experience with modern AI application frameworks and LLMOps tools such as LangGraph, LangChain, LlamaIndex, and leading LLM APIs
  • Strong programming and software engineering capabilities with the ability to build and deploy production-quality AI applications
  • Experience using AI-native development tools such as Cursor, Claude Code, or similar tools is preferred, with strong judgment around code quality, security, and production reliability

Good-to-Have Experience

  • GraphRAG
  • Long-context architectures
  • Model routing
  • Semantic and intelligent caching
  • Model cascades
  • PEFT / LoRA / QLoRA
  • Knowledge retrieval and grounding
  • Model distillation
  • Open-source model deployment
  • Advanced LLM evaluation and observability
  • Enterprise AI security and governance

Must-Have Skills

  • Machine Learning
  • Generative AI (GenAI)
  • Production AI/ML Systems
  • LLM Applications
  • Python / Software Engineering
  • AI Application Architecture

Good-to-Have Skills

  • End-to-End Production AI
  • Fine-Tuning
  • RAG
  • Agentic AI
  • LLMOps
  • LangGraph / LangChain / LlamaIndex
  • Model Evaluation & Observability
  • GraphRAG
  • PEFT / LoRA / QLoRA

More Info

Job Type:
Industry:
Employment Type:

Key Skills

LangChain

Generative AI

AI Application Architecture

LLMOps

Python Software Engineering

Model Evaluation

LangGraph

Agentic AI

LLM Applications

RAG

LlamaIndex

Observability

Production AI ML Systems

About Company

Similar Jobs

9-14 yrs
INR 333,333 - 666,667 per month
Bengaluru
Skills:
Machine Learning
12-14 yrs
Bengaluru, India
Skills:
JavaMachine LearningIntegration TestingDockerPythonworkflow management toolsgRPC-based web servicesBatch Processing
6-9 yrs
Bengaluru, India
Skills:
Distributed State API DesignLLMOps Agent EvaluationTechnical AI LeadershipProduction Agentic EngineeringCognitive Telemetry
8-10 yrs
Bengaluru, India
Skills:
data engineering snowflake KubernetesMachine LearningTensorflowNosqlApache AirflowPytorchDockerPythonAWSJavaScalaSqlGoogle CloudAzure MLJenkinsMLopsApache KafkaDatabricksAzureSnowparkAWS SagemakerdbtCircleCIScikit-learnGitLab CIMLFLow