MongoDB10 дней назад
Staff Site Reliability Engineer
Зарплата не указана
Bengaluru
Обязанности
- 01Provide technical leadership for the operational foundations that enable deployment at scale of AI applications
- 02Own the reliability architecture of the platform as it expands across regions and cloud providers
- 03Set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline
- 04Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices
- 05Set operational standards for the team: on-call quality, incident response, and SLO discipline
- 06Mentor and technically develop the SRE team
- 07Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
Требования
- 0110+ years of experience working on software and operating distributed systems with deep Kubernetes expertise, including designing or evolving multi-cluster platforms
- 02Proficiency in Python, Go, or a similar programming language
- 03Understand workload isolation at the systems level: containers, virtual machines, and the trade-offs between them for running untrusted code
- 04Possess a customer-focused mindset
- 05Value efficiency in processes and operations and display a strong preference for automation over manual processes
- 06Be intimately familiar with the infrastructure primitives of at least one of AWS, GCP, or Azure, and comfortable reasoning about differences between them
- 07Have a track record of driving infrastructure architecture across teams and mentoring engineers
Условия
- 01Hybrid working model based in Bengaluru
- 02Participation in a 24/7 on-call rotation
- 03Employee affinity groups
- 04Fertility assistance
- 05Generous parental leave policy
- 06Supportive and enriching culture focused on employee wellbeing