Loading open roles
Loading open roles
Loading role
Recroots · posted 6 months ago
Key Responsibilities
· ML Platform Architecture: Design and scale distributed ML platform for data, feature store, training, and inferencing; build multi-region, containerized infrastructure using Kubernetes and Terraform.
· Training & Inferencing at Scale: Optimize GPU/TPU/CPU utilization for large-scale training and real-time inferencing; implement distributed training, model parallelism, and caching with Ray.
· Performance & Governance: Drive FinOps-aligned architecture, auto-scaling, and efficiency; enable observability, SLA/SLO tracking, and incident management.
· Collaboration & Leadership: Partner with cross-functional teams to deliver high-impact ML solutions; mentor engineers and set best practices for ML systems design.
Technology Stack
ML Frameworks: TensorFlow, PyTorch
Serving: Triton, TensorFlow Serving, TorchServe
Distributed Compute: Ray, Kubernetes, Spark, GPU/TPU optimization
Lifecycle: MLflow and AirFlow
Infra: AWS/GCP/Azure, Terraform, CI/CD
What We’re Looking For
· 9–14 years in ML platform/distributed systems.
· Strong in Python, Kubernetes, GPU/TPU optimization.
· Proven ability to design fault-tolerant, high-throughput ML systems.