MongoDB10 дней назад

Staff Site Reliability Engineer

Зарплата не указана
Bengaluru

Обязанности

  • 01Provide technical leadership for the operational foundations that enable deployment at scale of AI applications
  • 02Own the reliability architecture of the platform as it expands across regions and cloud providers
  • 03Set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline
  • 04Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices
  • 05Set operational standards for the team: on-call quality, incident response, and SLO discipline
  • 06Mentor and technically develop the SRE team
  • 07Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure

Требования

  • 0110+ years of experience working on software and operating distributed systems with deep Kubernetes expertise, including designing or evolving multi-cluster platforms
  • 02Proficiency in Python, Go, or a similar programming language
  • 03Understand workload isolation at the systems level: containers, virtual machines, and the trade-offs between them for running untrusted code
  • 04Possess a customer-focused mindset
  • 05Value efficiency in processes and operations and display a strong preference for automation over manual processes
  • 06Be intimately familiar with the infrastructure primitives of at least one of AWS, GCP, or Azure, and comfortable reasoning about differences between them
  • 07Have a track record of driving infrastructure architecture across teams and mentoring engineers

Условия

  • 01Hybrid working model based in Bengaluru
  • 02Participation in a 24/7 on-call rotation
  • 03Employee affinity groups
  • 04Fertility assistance
  • 05Generous parental leave policy
  • 06Supportive and enriching culture focused on employee wellbeing
Staff Site Reliability Engineer · Rekru