Physical Intelligence08/25/2024

Инженер инфраструктуры машинного обучения

Salary not specified
MARKET
15,900median for this role
Data Scientist · 115 jobs with disclosed pay
5,000half of the offers: 12,796–20,50043,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьОфис

Responsibilities

  • 01Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
  • 02Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
  • 03Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization
  • 04Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments
  • 05Manage compute resources: Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost
  • 06Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale
  • 07Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics

Requirements

  • 01Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms
  • 02Hands-on large-scale training experience in JAX (preferred), PyTorch
  • 03Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines
  • 04Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)
  • 05Ability to debug and optimize performance bottlenecks across the training stack
  • 06Strong cross-functional communication and ownership mindset