Crusoe22.07.2026

Principal Engineer, Conductor Platform (CAPE)

Зарплата не указана
РЫНОК
15 044медиана по профессии
Backend Developer · 82 вакансий с указанной зарплатой
5 092половина предложений: 11 885–20 1031,2 млн
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьОфис

Обязанности

  • 01Build a unified observability plane correlating GPU, networking, storage, orchestration, and workload signals
  • 02Develop the fleet as one logical computer with a single health model, scheduler, and source of truth
  • 03Implement closed-loop autonomy for diagnosing, deciding, and remediating without human intervention
  • 04Maximize goodput as an objective function through scheduling, placement, and maintenance decisions
  • 05Predict failures hours ahead by forecasting GPU, NVLink, optics, and thermal degradation
  • 06Detect straggler and silent failures by isolating exact rank, GPU, and node from collective-operation signals
  • 07Schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom
  • 08Create a digital twin of the fleet to simulate failures, scheduling policies, and remediation logic
  • 09Develop agentic operations that propose and execute fixes with guardrails and audit trails
  • 10Implement zero-trust, fully auditable multi-tenancy with identity-scoped and policy-checked actions
  • 11Establish self-qualifying hardware through automated burn-in for new and repaired nodes

Требования

  • 0110+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation)
  • 02Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and autonomous decision systems
  • 03Hands-on fluency with GPU/HPC infrastructure including GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior
  • 04Track record of designing and shipping large-scale observability or telemetry platforms correlating signals across compute, network, and storage layers
  • 05Comfort operating in ambiguity and defining architecture and standards for new systems
  • 06Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar)
  • 07Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus
  • 08Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus

Условия

  • 01Competitive compensation and equity packages
  • 02Restricted Stock Units
  • 03Paid time off, paid holidays & leave of absence programs
  • 04Comprehensive health, dental & vision insurance
  • 05Employer contributions to HSA account
  • 06Paid parental leave
  • 07Paid life insurance, short-term and long-term