MongoDB10 дней назад

Senior Site Reliability Engineer

Зарплата не указана
Gurugram

Навыки

PythonGoKubernetesAWSGoogle Cloud Platform (GCP)AzureLinuxTCP/IPDNSTLSIstioCilium

Обязанности

  • 01Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads
  • 02Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
  • 03Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
  • 04Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
  • 05Mentor early-career SREs and contribute to the team’s operational practices as it grows

Требования

  • 01Strong background in software development and operating distributed systems
  • 026+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language
  • 03Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues
  • 04Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
  • 05Strong understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing
  • 06Customer-focused mindset and strong verbal and written technical communication skills, with a desire to collaborate with colleagues
  • 07Strong bias for efficient processes, operational simplicity, and automation over manual work
  • 08Eager to learn, with a strong technical background

Условия

  • 01Based in Gurugram for hybrid working model
  • 02Participate in a 24/7 on-call rotation