SpaceXAI3 дня назад

Инженер по анализу отказов оборудования - Мемфис

Зарплата не указана
РЫНОК
15 044медиана по профессии
Backend Developer · 82 вакансий с указанной зарплатой
5 092половина предложений: 11 885–20 1031,2 млн
Работодатель не указал зарплату — сравните с рынком сами.
Southaven

Обязанности

  • 01Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's data center environment
  • 02Run security scanning and CVE / vulnerability analysis on firmware and related components
  • 03Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor
  • 04Investigate and diagnose hardware failures, including 'grey failures' (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis
  • 05Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions
  • 06Collaborate with Data Center Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time
  • 07Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime
  • 08Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base
  • 09Participate in on-call rotations and incident response for hardware-related issues in the Memphis data center

Требования

  • 01Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience)
  • 022+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments
  • 03Proven expertise in firmware analysis, hardware specifications review, and release validation
  • 04Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols
  • 05Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software
  • 06Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies
  • 07Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar)
  • 08Excellent problem-solving skills with a data-driven approach to reliability engineering
  • 09Ability to work collaboratively with cross-functional teams, including operations technicians

Условия

  • 01Flat organizational structure
  • 02Hands-on work environment contributing directly to the company's mission
  • 03Leadership opportunities for those showing initiative
  • 04On-call rotations and incident response for hardware-related issues in the Memphis data center