Okta7 дней назад

Staff Site Reliability Engineer

Зарплата не указана
Bengaluru

Обязанности

  • 01Design, build, and operate large-scale cloud infrastructure and production services
  • 02Participate in an on-call rotation supporting highly available customer-facing systems
  • 03Lead incident response efforts and drive post-incident reviews focused on systemic improvements
  • 04Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
  • 05Partner with engineering teams to improve service availability, scalability, performance, and resilience
  • 06Continuously improve observability through metrics, logging, tracing, dashboards, and alerting
  • 07Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies
  • 08Eliminate operational toil through automation, tooling, and platform engineering
  • 09Improve deployment safety and operational workflows through CI/CD and GitOps practices
  • 10Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities
  • 11Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security
  • 12Lead complex reliability initiatives spanning multiple engineering teams
  • 13Guide engineers in adopting operational best practices and reliability engineering principles
  • 14Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing
  • 15Influence architecture and operational decisions through data-driven recommendations and engineering expertise
  • 16Drive projects from conception through production rollout and long-term operational ownership
  • 17Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation
  • 18Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity

Требования

  • 01Strong experience operating large-scale production services in AWS and/or GCP
  • 02Deep expertise with Kubernetes in production environments
  • 03Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues
  • 04Extensive experience with Infrastructure as Code technologies such as Terraform and Helm
  • 05Strong software engineering skills in Golang and/or Python
  • 06Experience building automation and internal engineering platforms
  • 07Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies
  • 08Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management
  • 09Experience with observability platforms, monitoring strategies, and production telemetry
  • 10Experience with or strong interest in AI-assisted engineering and operational automation
  • 11Strong expertise operating customer-facing production systems
  • 12Experience leading incident response and driving operational improvements
  • 13Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning
  • 14Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices
  • 15Proven ability to balance reliability, scalability, security, and engineering velocity
  • 16Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design
  • 17Experience implementing operational controls and best practices in regulated environments