Physical Intelligence08/25/2024
Инженер инфраструктуры машинного обучения
Salary not specified
MARKET
15,900 ₽median for this role
Data Scientist · 115 jobs with disclosed pay
5,000half of the offers: 12,796–20,50043,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьОфис
Responsibilities
- 01Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
- 02Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
- 03Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization
- 04Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments
- 05Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost
- 06Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale
- 07Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics
Requirements
- 01Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms
- 02Hands-on large-scale training experience in JAX (preferred), PyTorch
- 03Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines
- 04Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)
- 05Ability to debug and optimize performance bottlenecks across the training stack
- 06Strong cross-functional communication and ownership mindset