Crusoe10 days ago
Senior Staff Deployment Automation Engineer
RUB 20,833–25,000per month · before tax
MARKET
23,750 ₽median for this role
DevOps Engineer · 6 jobs with disclosed pay
13,000half of the offers: 17,708–26,97989,259
At market · above 50% of offers
Полная занятостьОфис
Responsibilities
- 01Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe's AI Cloud Stack
- 02CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications
- 03Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads
- 04Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc
- 05Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary
- 06Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments
- 07Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts
Requirements
- 0112+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field
- 02Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes
- 03Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres
- 04Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters
- 05Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack
- 06Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios
- 07Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU)
- 08Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context
- 09Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems
What we offer
- 01Competitive compensation and equity packages
- 02Restricted Stock Units
- 03Paid time off, paid holidays & leave of absence programs
- 04Comprehensive health, dental & vision insurance
- 05Employer contributions to HSA account
- 06Paid parental leave
- 07Paid life insurance, short-term and long-term disability
- 08Professional development & tuition reimbursement
- 09Mental health & wellness support
- 10Commuter benefits (parking & transit)
- 11Cell phone stipend
- 12401(k) Retirement plan with company match up to 4% of salary
- 13Volunteer time off
- 14Global travel insurance & emergency assistance
- 15Daily meals allowance
- 16Additional perks & programs specific to location
- 17Onsite work in San Francisco, Sunnyvale, Bellevue