SpaceXAI10 days ago
Site Reliability Engineer - Data Center
Salary not specified
Memphis
Responsibilities
- 01Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's datacenter environment
- 02Run security scanning and CVE / vulnerability analysis on firmware and related components
- 03Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor
- 04Investigate and diagnose hardware failures, including grey failures (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis
- 05Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions
- 06Collaborate with Datacenter Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time
- 07Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime
- 08Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base
- 09Participate in on-call rotations and incident response for hardware-related issues in the Memphis datacenter
Requirements
- 01Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience)
- 022+ years of experience in hardware reliability engineering, preferably in high-performance computing or datacenter environments
- 03Proven expertise in firmware analysis, hardware specifications review, and release validation
- 04Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols
- 05Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software
- 06Familiarity with datacenter hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies
- 07Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar)
- 08Excellent problem-solving skills with a data-driven approach to reliability engineering
- 09Ability to work collaboratively with cross-functional teams, including operations technicians
What we offer
- 01On-call rotations and incident response for hardware-related issues
- 02Work in Memphis datacenter
- 03Flat organizational structure
- 04Hands-on contribution to company's mission