SpaceXAI10 days ago

Site Reliability Engineer - Data Center

Salary not specified
Memphis

Responsibilities

  • 01Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's datacenter environment
  • 02Run security scanning and CVE / vulnerability analysis on firmware and related components
  • 03Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor
  • 04Investigate and diagnose hardware failures, including grey failures (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis
  • 05Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions
  • 06Collaborate with Datacenter Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time
  • 07Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime
  • 08Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base
  • 09Participate in on-call rotations and incident response for hardware-related issues in the Memphis datacenter

Requirements

  • 01Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience)
  • 022+ years of experience in hardware reliability engineering, preferably in high-performance computing or datacenter environments
  • 03Proven expertise in firmware analysis, hardware specifications review, and release validation
  • 04Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols
  • 05Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software
  • 06Familiarity with datacenter hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies
  • 07Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar)
  • 08Excellent problem-solving skills with a data-driven approach to reliability engineering
  • 09Ability to work collaboratively with cross-functional teams, including operations technicians

What we offer

  • 01On-call rotations and incident response for hardware-related issues
  • 02Work in Memphis datacenter
  • 03Flat organizational structure
  • 04Hands-on contribution to company's mission