Crusoe20 дней назад
Старший инженер-программист (облачная инфраструктура)
Зарплата не указана
РЫНОК
15 167 ₽медиана по профессии
Cloud Engineer · 5 вакансий с указанной зарплатой
6 283половина предложений: 13 358–15 250150 000
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьОфис
Обязанности
- 01Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems
- 02Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X
- 03Execute component-level diagnosis and remediation for failed or degraded hardware
- 04Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware
- 05Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance
- 06Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan
- 07Perform firmware and BIOS upgrades across the GPU fleet
- 08Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems
- 09Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows
- 10Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions
- 11Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team
Требования
- 01Ability to code in Golang
- 02Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments
- 03Deep understanding of GPU architectures and hands-on experience with GPU-based systems
- 04Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms
- 05Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE)
- 06Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing
- 07Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities
- 08Experience working with enterprise server hardware, power delivery, and cooling systems
- 09Strong analytical and problem-solving skills
- 10Excellent communication and collaboration skills
- 11Ability to work independently in a fast-paced data center or operations environment
Условия
- 01Hybrid work schedule
- 02Industry competitive pay
- 03Restricted Stock Units in a fast growing, well-funded technology company
- 04Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
- 05Employer contributions to HSA accounts
- 06Paid Parental Leave
- 07Paid life insurance, short-term and long-term disability
- 08Teladoc
- 09401(k) with a 100% match up to 4% of salary
- 10Generous paid time off and holiday schedule
- 11Cell phone reimbursement
- 12Tuition reimbursement
- 13Subscription to the Calm app
- 14MetLife Legal
- 15Company paid commuter benefit; $300 per pay period
- 16Compensation will be paid in the range of $215,000 - $260,000