MLOps Engineer
Spotted 1d agoFull-time
Job details
- Employment
- Full-time
- Level
- Entry level
- Experience
- 2+ years
- Education
- Bachelor's degree
- Posted
- Oct 9, 2026
- Last confirmed open
- Oct 9, 2026
Job description
About this role
About The Role
The role owns the infrastructure and pipelines that turn ML experiments into reliable production systems — CI/CD for models, scalable training infrastructure, feature stores, and low-latency serving environments.
Working at the intersection of data science and platform engineering, this position sits on a team where deployment reliability, model reproducibility, and infrastructure cost efficiency are all first-class concerns, not afterthoughts.
Key Responsibilities
- Design and maintain CI/CD pipelines for ML workflows using tools like GitHub Actions, GitLab CI, or Jenkins, with automated testing and validation gates for models and data
- Build and operate model serving infrastructure (KServe, Seldon, SageMaker endpoints, or Triton) supporting low-latency inference at production scale
- Implement experiment tracking and model registry workflows with MLflow, Weights & Biases, or SageMaker Experiments to guarantee reproducibility across environments
- Build monitoring and observability for production ML: data drift detection, latency/throughput dashboards, and automated retraining triggers using tools like Evidently, Prometheus, and Grafana
- Automate data and feature pipelines in partnership with data engineering, leveraging Airflow, Kubeflow, or Dagster orchestration
- Manage Kubernetes-based ML workloads, optimizing GPU utilization, autoscaling policies, and cloud spend across training and inference clusters
- Establish and document MLOps standards — model versioning, rollback procedures, and incident response — and champion them across data science teams
What We Are Looking For
- 3–6 years of experience in MLOps, DevOps, or ML engineering, with at least 2 years deploying and operating ML systems in production
- Strong Python and Bash skills; proficiency with Docker and Kubernetes for containerizing and orchestrating ML workloads
- Hands-on experience with at least one cloud platform (AWS, GCP, or Azure) and its ML-specific services
- Practical experience with ML workflow tooling: MLflow, Kubeflow, Airflow, or equivalent pipelines and registries
- Solid grasp of production ML concerns: model versioning, canary deployments, drift monitoring, and retraining strategies
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
- Bonus: Experience with GPU cost optimization, distributed training (Ray, Dask), infrastructure-as-code (Terraform), or serving LLMs at scale
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn