Oct 2024 – Present
Aug 2023 – Oct 2024
0-1 built a centralized multi-region, multi-cluster control-plane for running large scale batch workloads natively on Kubernetes. The service features cross-cloud and cross-region support, along with heterogeneous resource discovery and management. Designed a truly serverless architecture for batch workloads in a modern, cloud-native way.
0-1 built Apple’s Batch Inference Service, supporting large-scale LLM batfch inferences for 150+ teams.
Designed and deployed on AWS EKS with a multi-cluster GPU infrastructure. Led API server design, implemented priority scheduling
and preemption using Apache YuniKorn for GPU resource management.
Apr 2022 to Aug 2023