Crusoe20 дней назад

Старший инженер-программист (облачная инфраструктура)

Зарплата не указана
РЫНОК
15 167медиана по профессии
Cloud Engineer · 5 вакансий с указанной зарплатой
6 283половина предложений: 13 358–15 250150 000
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьОфис

Обязанности

  • 01Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems
  • 02Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X
  • 03Execute component-level diagnosis and remediation for failed or degraded hardware
  • 04Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware
  • 05Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance
  • 06Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan
  • 07Perform firmware and BIOS upgrades across the GPU fleet
  • 08Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems
  • 09Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows
  • 10Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions
  • 11Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team

Требования

  • 01Ability to code in Golang
  • 02Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments
  • 03Deep understanding of GPU architectures and hands-on experience with GPU-based systems
  • 04Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms
  • 05Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE)
  • 06Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing
  • 07Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities
  • 08Experience working with enterprise server hardware, power delivery, and cooling systems
  • 09Strong analytical and problem-solving skills
  • 10Excellent communication and collaboration skills
  • 11Ability to work independently in a fast-paced data center or operations environment

Условия

  • 01Hybrid work schedule
  • 02Industry competitive pay
  • 03Restricted Stock Units in a fast growing, well-funded technology company
  • 04Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • 05Employer contributions to HSA accounts
  • 06Paid Parental Leave
  • 07Paid life insurance, short-term and long-term disability
  • 08Teladoc
  • 09401(k) with a 100% match up to 4% of salary
  • 10Generous paid time off and holiday schedule
  • 11Cell phone reimbursement
  • 12Tuition reimbursement
  • 13Subscription to the Calm app
  • 14MetLife Legal
  • 15Company paid commuter benefit; $300 per pay period
  • 16Compensation will be paid in the range of $215,000 - $260,000