Search Jobs

Search by job, company or skills

MLOps Engineer
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

About The Role

We're looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You'll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What You'll Do

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.

Our current stack

You'll Work With Technologies Including

  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What We're Looking For

  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.

Nice to have

  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.

What Success Looks Like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

GPU infrastructure

LangFuse

MLflow

vLLM

agent workflows

LLM applications

Knative

vector retrieval

KServe

model-serving workloads

RAG architectures

About Company