MLOps Engineer

Evlo AI · Phoenix, AZ

Spotted 41m agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

About The Role

The role owns the infrastructure and tooling that turns experimental ML work into reliable, production-scale systems — CI/CD for models, scalable training pipelines, monitoring, and cost-efficient GPU utilization across the platform.

You will sit at the intersection of ML engineering and platform engineering, enabling data scientists and ML engineers to ship models faster while keeping training and serving infrastructure stable, observable, and reproducible.

Key Responsibilities

  • Design and maintain CI/CD pipelines for ML workflows using GitLab CI, GitHub Actions, or Jenkins, including automated testing, model versioning, and staged rollouts
  • Build and operate containerized training and inference workloads on Kubernetes, leveraging GPU scheduling, autoscaling, and resource quotas to control cloud spend
  • Deploy and manage model serving infrastructure using tools like KServe, Triton Inference Server, TorchServe, or SageMaker endpoints
  • Implement ML observability: model drift detection, latency/throughput dashboards, and automated alerting with Prometheus, Grafana, and custom metrics pipelines
  • Orchestrate data and training pipelines with Airflow, Kubeflow, or Metaflow, ensuring reproducibility and lineage across environments
  • Manage model registries and artifact stores (MLflow, Weights & Biases) and enforce experiment tracking standards across teams
  • Partner with data scientists and ML engineers to debug production incidents, optimize batch/online inference performance, and reduce infrastructure costs

What We Are Looking For

  • 3–7 years of experience in MLOps, ML engineering, DevOps, or platform engineering supporting ML workloads in production
  • Strong Python and Bash skills; proficiency with Docker, Kubernetes, and at least one major cloud (AWS, GCP, or Azure)
  • Hands-on experience building CI/CD pipelines for ML systems, including model versioning, canary deployments, and rollback strategies
  • Experience with workflow orchestration tools (Airflow, Kubeflow, Metaflow) and ML experiment tracking (MLflow, W&B)
  • Working knowledge of distributed training and serving concepts: data parallelism, batching, quantization, and GPU resource management
  • BS/MS in Computer Science, Engineering, or related field, or equivalent practical experience
  • Bonus: Experience with Terraform/Infrastructure-as-Code, LLM inference optimization (vLLM, Triton), feature stores (Feast, Tecton), or SOC 2/production security hardening for ML platforms
Interested in this role?Continue on Evlo AI's careers page.
Apply on Evlo AI