MLOps Engineer
About this role
Employer-provided description, formatted for easier reading.
Cortex Real-Time Inference Engineering
Lead
Role Overview
We are seeking an
Inference Engineering
Lead
to design, build, and scale our next-generation, low-latency model serving platform. You will bridge data science and production engineering, establishing resilient, autoscaling architectures that deliver high availability and sub-second latency for real-time predictive workflows.
Key Responsibilities
- Inference Platform Engineering:
Architect low-latency, highly available model-serving architectures for real-time inference.
- Deployment & CI/CD:
Establish automated CI/CD pipelines (GitOps) and zero-downtime deployment patterns (Canary, Blue/Green, Shadow).
- Autoscaling & Capacity Control:
Design advanced, metrics-driven autoscaling to handle traffic spikes efficiently and optimize cloud costs.
- Performance Optimization:
Benchmark, profile, and optimize inference models and APIs to strictly adhere to low-latency SLAs.
- Observability & SLOs:
Implement robust monitoring and alerting (Prometheus/Grafana) for system health, data quality, and model drift.
- Resilience & Fault Tolerance:
Build self-healing, fault-tolerant infrastructure with fallback mechanisms across hybrid environments (Cloud & On-Prem).
Key Qualifications
- MLOps & Serving:
Deep expertise in online inference architectures, feature stores, and frameworks like Triton, TorchServe, vLLM, Seldon Core, or KServe.
- Kubernetes & Cloud:
Advanced proficiency in Kubernetes (HPA, custom metrics), microservices, and hybrid cloud/on-prem deployments (AWS/GCP/Azure).
- API & Languages:
Strong programming skills in Python, Go, C++, or Java; high proficiency in gRPC and REST APIs.
- Performance & Reliability:
Hands-on experience with load testing (Locust, K6), latency reduction (quantization, GPU acceleration), and observability tooling.
- Experience:
5+ years in Software/DevOps Engineering, with 3+ years dedicated to production MLOps and ML pipelines. BS/MS in CS or related field preferred.