MLOps Engineer
Spotted 2h agofulltime
Job description
About this role
Employer-provided description, formatted for easier reading.
About The Role
The role owns the ML infrastructure layer that keeps models running reliably in production - CI/CD for ML, scalable training pipelines, model serving infrastructure, and observability across the entire ML lifecycle.
You will work directly with ML engineers and data scientists to turn experimental models into production systems, and own the platform decisions that determine how fast the team ships.
Key Responsibilities
- Build and maintain CI/CD pipelines for ML workflows using tools like GitHub Actions, Kubeflow, or MLflow, enabling rapid and safe model releases
- Design and operate model serving infrastructure on Kubernetes (KServe, Seldon, or Triton), handling autoscaling, canary deployments, and rollbacks
- Orchestrate large-scale training and batch inference pipelines with Airflow, Kubeflow Pipelines, or Ray, on AWS or GCP
- Implement model monitoring and observability: data drift detection, latency/throughput dashboards, and automated alerting with tools like Evidently, Grafana, or Datadog
- Establish experiment tracking and model registry standards (MLflow, Weights & Biases) to ensure reproducibility across teams
- Optimize infrastructure cost and performance - GPU utilization, spot instance strategies, and inference quantization/batching
- Partner with data scientists to debug production issues, improve deployment velocity, and codify MLOps best practices across the team
What We Are Looking For
- 3–6 years of experience in MLOps, ML platform engineering, or backend/DevOps engineering with significant ML infrastructure exposure
- Strong Python and Go (or similar) skills, with solid software engineering fundamentals in production systems
- Hands-on experience with Kubernetes in production, including deploying and scaling stateful and GPU-accelerated workloads
- Deep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and their ML-specific services
- Experience with workflow orchestration (Airflow, Kubeflow, Prefect) and model serving frameworks (TorchServe, Triton, KServe, Seldon)
- BS in Computer Science, Engineering, or equivalent practical experience; MS a plus
- Bonus: Experience with feature stores (Feast, Tecton), LLM inference optimization (vLLM, TensorRT-LLM), or contributing to open-source MLOps tooling
Interested in this role?Continue on Evlo AI's careers page.
Apply on Evlo AI