Crusoe19 дней назад

Senior Hardware Systems Engineer, Performance

Зарплата не указана
Полная занятостьОфис

Обязанности

  • 01Drive the end-to-end lifecycle of next-generation compute platforms, including evaluation, bring-up, validation, deployment, and production readiness
  • 02Define and execute performance characterization and validation strategies for CPU, GPU, and accelerated computing platforms
  • 03Conduct in-depth workload characterization studies across training and inference - dense, MoE, long-context, and multimodal models to understand compute, memory, communication, and I/O behavior on target platforms
  • 04Translate workload and platform insights into cluster-level tuning and configuration recommendations: topology, parallelism strategy, scheduling, power, and software stack settings to maximize delivered performance and efficiency
  • 05Build and maintain workload performance profiles and reference configurations that guide how clusters are deployed, tuned, and scaled for specific model families and workload classes
  • 06Analyze system and workload performance, identify bottlenecks, and work across hardware and software layers to drive improvements
  • 07Lead complex system-level debugging across compute, memory, storage, networking, accelerators, and platform firmware
  • 08Partner with vendors and internal engineering teams on prototyping, qualification, NPI, and production readiness of new technologies
  • 09Collaborate across hardware, firmware, networking, software, infrastructure, reliability, and operations teams to resolve complex platform issues
  • 10Use data and system-level insights to influence platform architecture, technology selection, hardware roadmaps, and long-term infrastructure strategy

Требования

  • 015-6+ years of experience in hardware systems engineering, platform engineering, performance engineering, ML systems engineering, infrastructure engineering, or related areas
  • 02Hands-on experience with large-scale GPU or accelerated computing infrastructure for AI/ML or HPC workloads
  • 03Hands-on experience with distributed training and/or inference workloads at scale, including parallelism strategies and performance tuning across the hardware/software stack
  • 04Experience with workload benchmarking, performance profiling, and system performance optimization across hardware and software layers
  • 05Strong understanding of modern server and accelerator architectures, including CPU, GPU, memory, storage, networking, and high-speed interconnects such as PCIe, InfiniBand, or NVLink
  • 06Hands-on experience with system bring-up, validation, performance characterization, and root-cause analysis of complex hardware/software issues
  • 07Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages
  • 08Ability to analyze system behavior using telemetry, benchmarks, profiling tools, and other quantitative data
  • 09Experience working across multiple engineering disciplines, including hardware, firmware, software, networking, and infrastructure teams
  • 10Strong analytical and problem-solving skills with the ability to operate effectively in ambiguous and rapidly evolving environments
  • 11Excellent technical communication skills and experience collaborating with internal engineering teams, customers, and external technology partners
  • 12Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience