Meta
Feb 2024 – PresentSoftware Engineer, AI Infrastructure · Knowledge Distillation & AI Production Systems
- Leading fleet-scale GPU inference migration for knowledge distillation models in an 84% over-subscribed accelerator fleet: $4M/yr realized, $12M/yr in flight, $50M/yr projected at full scale.
- Fit models that used to need 4+ RDMA-equipped hosts onto cheaper non-RDMA hardware by combining CPU and SSD embedding offloading with multi-host partitioning, across NVIDIA H100/A100 and AMD MI350X/MI300X.
- Designed and shipped a gang-level canary and automated-rollback system for KD inference releases, preventing 21 bad-package rollouts across 14 high-revenue models.
- Own release reliability for 4,000+ production AI models. Raised release config-test coverage from 0% to 90%+ with simulator-based validation.