SpaceXAI4 дня назад
Network Engineer (Supercomputer Infrastructure) - Memphis
Зарплата не указана
Southaven
Обязанности
- 01Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks
- 02Design and maintain supercomputer data center and campus networks in accordance with company network standards
- 03Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams
- 04Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond
- 05Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC)
- 06Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules)
- 07Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives
- 08Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations
- 09Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference
- 10Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures
- 11Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks
- 12Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects
- 13Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks
Требования
- 01Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience; OR 5+ years of professional network engineering experience in lieu of a degree
- 02Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments
- 03Functional experience with multiple network vendors in production or lab environments
- 04Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor
- 05Ability to pass applicable background checks for site access
- 06Ability to work in tight quarters; physical dexterity is necessary to perform job functions
- 07Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs
Условия
- 01Flexible schedule as required by compute schedules
- 02Work in a small, highly motivated team with a flat organizational structure
- 03Hands-on contribution directly to the company’s mission
- 04Leadership opportunities based on initiative and consistent delivery of excellence