OpenAI16.05.2026

Tech Lead, Deployment & Operations — Custom Infrastructure

Зарплата не указана
Полная занятостьУдалёнка

Обязанности

  • 01Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments
  • 02Own the path from hardware bring-up and validation through production deployment, operational readiness, and sustained fleet support
  • 03Partner closely with silicon, systems, software, infrastructure, networking, data center, supply chain, and external partner teams to ensure successful deployment at scale
  • 04Define deployment processes, operational playbooks, technical readiness criteria, escalation paths, and reliability practices for new hardware platforms
  • 05Drive cross-functional execution across lab bring-up, rack/system integration, data center deployment, fleet monitoring, debugging, and issue resolution
  • 06Stay hands-on technically through architecture reviews, deployment planning, failure analysis, operational debugging, and critical system-level decision-making
  • 07Identify gaps in tooling, observability, automation, validation coverage, and operational processes, and build plans to close them
  • 08Establish clear metrics for deployment readiness, reliability, performance, maintainability, and operational health
  • 09Build a strong engineering culture grounded in ownership, technical rigor, operational excellence, and high-velocity execution
  • 10Ensure OpenAI’s custom hardware platforms can be deployed and operated reliably, repeatably, and safely at scale
  • 11Be a contributor and technical driver for the architecture and design of future ML systems

Требования

  • 018+ years of engineering experience in hardware systems, infrastructure, data center deployment, production operations, systems engineering, silicon bring-up, or related technical domains
  • 02Strong technical depth in one or more of: hardware deployment, data center operations, rack-scale systems, silicon bring-up, systems validation, fleet operations, reliability engineering, infrastructure automation, or hardware/software integration
  • 03Experience bringing complex hardware systems from development or validation into production environments
  • 04Experience working closely with silicon, systems, software, infrastructure, networking, or data center teams
  • 05Experience with deployment planning, operational readiness, incident response, debugging, and root-cause analysis for production systems
  • 06Experience building tooling, automation, observability, or operational processes that improve deployment quality and fleet reliability
  • 07Demonstrated ability to hire, develop, and lead senior technical talent
  • 08Ability to move fluidly between people leadership, technical strategy, and hands-on operational problem solving
  • 09Strong written and verbal communication skills, especially in high-urgency, cross-functional technical environments
  • 10Experience working in fast-moving environments

Условия

  • 01Compensation Range: $342K - $445K USD
Tech Lead, Deployment & Operations — Custom Infrastructure · Rekru