OpenAI24 дня назад

Network Operations Engineer, AI Networking

Зарплата не указана
Полная занятостьSan Francisco

Обязанности

  • 01Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers
  • 02Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR)
  • 03Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks
  • 04Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact
  • 05Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance
  • 06Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams
  • 07Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure
  • 08Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues
  • 09Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection
  • 10Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures
  • 11Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency
  • 12Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency

Требования

  • 01Bachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience
  • 025+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure
  • 03Experience supporting production network environments with high-availability requirements
  • 04Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS
  • 05Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs
  • 06Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links
  • 07Experience performing software upgrades, hardware maintenance, and production change management
  • 08Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams

Условия

  • 01Participate in a 24x7 on-call rotation supporting mission-critical AI infrastructure
  • 02Support time-sensitive production incidents, maintenance windows, capacity expansions, and network changes with a focus on service availability and minimal customer impact
  • 03This role requires up to 30% travel to data center locations for new turnups and acceptance activities, as needed
Network Operations Engineer, AI Networking · Rekru