Back to jobs

MLOps Engineer

Evlo AI · Washington, DC

Spotted 1d agoFull-time

Job details

Employment
Full-time
Level
Entry level
Experience
2+ years
Education
Bachelor's degree
Posted
Oct 9, 2026
Last confirmed open
Oct 9, 2026
Job description

About this role

About The Role

The role owns the infrastructure and pipelines that turn ML experiments into reliable production systems — CI/CD for models, scalable training infrastructure, feature stores, and low-latency serving environments.

Working at the intersection of data science and platform engineering, this position sits on a team where deployment reliability, model reproducibility, and infrastructure cost efficiency are all first-class concerns, not afterthoughts.

Key Responsibilities

  • Design and maintain CI/CD pipelines for ML workflows using tools like GitHub Actions, GitLab CI, or Jenkins, with automated testing and validation gates for models and data
  • Build and operate model serving infrastructure (KServe, Seldon, SageMaker endpoints, or Triton) supporting low-latency inference at production scale
  • Implement experiment tracking and model registry workflows with MLflow, Weights & Biases, or SageMaker Experiments to guarantee reproducibility across environments
  • Build monitoring and observability for production ML: data drift detection, latency/throughput dashboards, and automated retraining triggers using tools like Evidently, Prometheus, and Grafana
  • Automate data and feature pipelines in partnership with data engineering, leveraging Airflow, Kubeflow, or Dagster orchestration
  • Manage Kubernetes-based ML workloads, optimizing GPU utilization, autoscaling policies, and cloud spend across training and inference clusters
  • Establish and document MLOps standards — model versioning, rollback procedures, and incident response — and champion them across data science teams

What We Are Looking For

  • 3–6 years of experience in MLOps, DevOps, or ML engineering, with at least 2 years deploying and operating ML systems in production
  • Strong Python and Bash skills; proficiency with Docker and Kubernetes for containerizing and orchestrating ML workloads
  • Hands-on experience with at least one cloud platform (AWS, GCP, or Azure) and its ML-specific services
  • Practical experience with ML workflow tooling: MLflow, Kubeflow, Airflow, or equivalent pipelines and registries
  • Solid grasp of production ML concerns: model versioning, canary deployments, drift monitoring, and retraining strategies
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
  • Bonus: Experience with GPU cost optimization, distributed training (Ray, Dask), infrastructure-as-code (Terraform), or serving LLMs at scale
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn