MongoDB10 дней назад
Senior Site Reliability Engineer
Зарплата не указана
Gurugram
Навыки
PythonGoKubernetesAWSGoogle Cloud Platform (GCP)AzureLinuxTCP/IPDNSTLSIstioCilium
Обязанности
- 01Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads
- 02Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
- 03Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
- 04Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
- 05Mentor early-career SREs and contribute to the team’s operational practices as it grows
Требования
- 01Strong background in software development and operating distributed systems
- 026+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language
- 03Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues
- 04Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
- 05Strong understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing
- 06Customer-focused mindset and strong verbal and written technical communication skills, with a desire to collaborate with colleagues
- 07Strong bias for efficient processes, operational simplicity, and automation over manual work
- 08Eager to learn, with a strong technical background
Условия
- 01Based in Gurugram for hybrid working model
- 02Participate in a 24/7 on-call rotation