Crusoe10 days ago

Senior Staff Deployment Automation Engineer

RUB 20,833–25,000per month · before tax
MARKET
23,750median for this role
DevOps Engineer · 6 jobs with disclosed pay
13,000half of the offers: 17,708–26,97989,259
At market · above 50% of offers
Полная занятостьОфис

Responsibilities

  • 01Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe's AI Cloud Stack
  • 02CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications
  • 03Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads
  • 04Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc
  • 05Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary
  • 06Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments
  • 07Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts

Requirements

  • 0112+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field
  • 02Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes
  • 03Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres
  • 04Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters
  • 05Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack
  • 06Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios
  • 07Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU)
  • 08Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context
  • 09Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems

What we offer

  • 01Competitive compensation and equity packages
  • 02Restricted Stock Units
  • 03Paid time off, paid holidays & leave of absence programs
  • 04Comprehensive health, dental & vision insurance
  • 05Employer contributions to HSA account
  • 06Paid parental leave
  • 07Paid life insurance, short-term and long-term disability
  • 08Professional development & tuition reimbursement
  • 09Mental health & wellness support
  • 10Commuter benefits (parking & transit)
  • 11Cell phone stipend
  • 12401(k) Retirement plan with company match up to 4% of salary
  • 13Volunteer time off
  • 14Global travel insurance & emergency assistance
  • 15Daily meals allowance
  • 16Additional perks & programs specific to location
  • 17Onsite work in San Francisco, Sunnyvale, Bellevue