Fundation Model Infrastructure, AIML, Apple

Oct 2024 – Present

  • Apple Intelligence: build large ML compute infrastructure for Fundation Model Training and Inference. Manage large TPU fleet, responsible for workload scheduler, work on training/inference resilience and efficiency.

Aug 2023 – Oct 2024

  • 0-1 built a centralized multi-region, multi-cluster control-plane for running large scale batch workloads natively on Kubernetes. The service features cross-cloud and cross-region support, along with heterogeneous resource discovery and management. Designed a truly serverless architecture for batch workloads in a modern, cloud-native way.

  • 0-1 built Apple’s Batch Inference Service, supporting large-scale LLM batfch inferences for 150+ teams. Designed and deployed on AWS EKS with a multi-cluster GPU infrastructure. Led API server design, implemented priority scheduling and preemption using Apache YuniKorn for GPU resource management.

Apr 2022 to Aug 2023

  • Scaled Apache Spark on Kubernetes at Apple, enabling large-scale batch processing across EKS and GKE. Worked on Batch Processing Gateway, leveraging Apache YuniKorn for resource management and scheduling. Leveraging Karpenter for autoscaling, instance lifecycle management, and Kubernetes version upgrades.