SpaceXAI4 дня назад

Network Engineer (Supercomputer Infrastructure) - Memphis

Зарплата не указана
Southaven

Обязанности

  • 01Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks
  • 02Design and maintain supercomputer data center and campus networks in accordance with company network standards
  • 03Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams
  • 04Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond
  • 05Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC)
  • 06Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules)
  • 07Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives
  • 08Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations
  • 09Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference
  • 10Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures
  • 11Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks
  • 12Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects
  • 13Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks

Требования

  • 01Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience; OR 5+ years of professional network engineering experience in lieu of a degree
  • 02Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments
  • 03Functional experience with multiple network vendors in production or lab environments
  • 04Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor
  • 05Ability to pass applicable background checks for site access
  • 06Ability to work in tight quarters; physical dexterity is necessary to perform job functions
  • 07Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs

Условия

  • 01Flexible schedule as required by compute schedules
  • 02Work in a small, highly motivated team with a flat organizational structure
  • 03Hands-on contribution directly to the company’s mission
  • 04Leadership opportunities based on initiative and consistent delivery of excellence