GitLab20 дней назад

Инженер по надежности систем (Site Reliability Engineer) – Выделенные хосты для раннеров

Зарплата не указана
Bangalore

Обязанности

  • 01Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments
  • 02Develop and maintain infrastructure as code using Terraform, contributing to common modules and deployment tooling
  • 03Write Go code for runner tooling and autoscaling components
  • 04Build and improve GitLab CI/CD pipelines for deployments, upgrades, and testing
  • 05Define and monitor service level objectives for CI job execution and build Grafana dashboards and alerts
  • 06Participate in an on-call rotation, handle incidents, and automate recurring toil
  • 07Run performance and scale testing and tune autoscaling parameters
  • 08Write documentation and runbooks for team operations

Требования

  • 01Professional experience operating production infrastructure on AWS at scale
  • 02Strong infrastructure-as-code experience with Terraform
  • 03Proficiency in Go or strong experience in another systems language with willingness to work in Go
  • 04Practical knowledge of CI/CD systems and job execution
  • 05Experience with observability practices and tools like Prometheus, Grafana, OpenSearch
  • 06Experience with on-call rotations and incident management for customer-facing systems
  • 07Strong problem-solving skills, excellent written communication, comfort working asynchronously across time zones

Условия

  • 01Flexible Paid Time Off
  • 02Equity Compensation & Employee Stock Purchase Plan
  • 03Growth and Development Fund
  • 04Parental Leave
  • 05Benefits to support health, finances, and well-being
  • 06Remote work across Americas, EMEA, APAC time zones