OpenAI24 дня назад
Network Operations Engineer, AI Networking
Зарплата не указана
Полная занятостьSan Francisco
Обязанности
- 01Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers
- 02Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR)
- 03Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks
- 04Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact
- 05Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance
- 06Support new AI cluster deployments, data center expansions, and infrastructure migrations in partnership with deployment and engineering teams
- 07Partner with cloud service providers (CSPs), colocation providers, Smart Hands teams, and hardware vendors to maintain production infrastructure
- 08Perform root-cause analysis (RCA) for production incidents and drive permanent corrective actions that eliminate recurring issues
- 09Build and maintain monitoring, telemetry, dashboards, and alerting to improve network observability and proactive issue detection
- 10Develop and improve operational runbooks, playbooks, troubleshooting documentation, and standard operating procedures
- 11Automate repetitive operational tasks using Python and infrastructure automation frameworks to reduce toil and improve efficiency
- 12Continuously identify opportunities to improve service reliability, scalability, operational maturity, and engineering efficiency
Требования
- 01Bachelor’s degree in Computer Science, Network Engineering, or a related discipline, or equivalent practical experience
- 025+ years of experience operating large-scale data center, cloud, AI, or HPC network infrastructure
- 03Experience supporting production network environments with high-availability requirements
- 04Hands-on experience with one or more of the following platforms: Cisco NX-OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS
- 05Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs
- 06Experience troubleshooting physical infrastructure, including fiber optics, transceivers, DAC/AOC cables, and high-speed Ethernet links
- 07Experience performing software upgrades, hardware maintenance, and production change management
- 08Excellent analytical and troubleshooting skills, with the ability to communicate technical risk clearly across teams
Условия
- 01Participate in a 24x7 on-call rotation supporting mission-critical AI infrastructure
- 02Support time-sensitive production incidents, maintenance windows, capacity expansions, and network changes with a focus on service availability and minimal customer impact
- 03This role requires up to 30% travel to data center locations for new turnups and acceptance activities, as needed