OpenAI9 дней назад
Simulation Infrastructure Engineer
Зарплата не указана
Полная занятостьSan Francisco
Обязанности
- 01Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible
- 02Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable
- 03Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks
- 04Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), solve engine-level scaling (parallelization, batching multiple runs per engine), and optimize cloud/GPU runtime reliability
- 05Produce metrics and tooling for measuring simulation health, throughput, fidelity regressions, and cost; create presubmit / canary tests that catch sim regressions early
- 06Implement artifact versioning, environment immutability (images / asset versions), experiment provenance, and policies for resource quotas and cost control across the sim farm
- 07Work closely with Sim Environments, Sim Realism, research, and ops to close the loop—ensuring simulation improvements directly translate into better model evaluation and training results
Требования
- 01Have deep software engineering & infra experience: you’ve built CI/CD at scale, authored reliable pipelines, and shipped production services that coordinate many moving parts
- 02Are comfortable with distributed compute and cloud GPU workloads: you know how to get many sims running concurrently (scheduling, batching, GPU orchestration) and optimize throughput/cost
- 03Have built or maintained HIL/SIL workflows or other sim↔hardware integrations and understand the operational challenges of bridging software and hardware testbeds
- 04Are strong with automation, observability and metrics: you enjoy instrumenting systems, defining meaningful KPIs, and surfacing regressions early
- 05Can design APIs and developer ergonomics so research and SWE teams can easily submit jobs, reproduce experiments, and interpret results
- 06Have experience with Python/C++/Rust, container orchestration (Kubernetes), distributed task queues, and CI systems; bonus if you’ve worked with RL tooling, task generators, or large-scale data pipelines
- 07Enjoy collaborating across teams to turn experimental simulation work into dependable production tooling
Условия
- 01This role is based in San Francisco, CA
- 02Requires in-person 4 days a week