Okta3 дня назад

Staff Site Reliability Engineer - Ecosystem

Зарплата не указана
Bengaluru

Обязанности

  • 01Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP
  • 02Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments
  • 03Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines)
  • 04Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability
  • 05Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers
  • 06Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP
  • 07Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response
  • 08Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams
  • 09Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning
  • 10Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity)
  • 11Participate in the on-call rotation, using incidents as learning opportunities to enhance systems and processes

Требования

  • 018+ years in SRE, DevOps, or Infrastructure Engineering roles
  • 023–5 years of experience with Kubernetes (EKS/GKE) and related ecosystem tools (Helm, Karpenter, etc.) in production
  • 033–5 years of experience with AWS and GCP
  • 043–5 years using Terraform to manage multi-cloud infrastructure
  • 053+ years of coding experience in Python, Go, or similar languages
  • 06Proven track record leading high-impact projects, specifically migration projects (ECS → EKS/GKE) and enabling microservice architectures
  • 07Experience implementing SLOs/SLIs, performing root cause analyses, and improving operational resilience
  • 08Prior work in SaaS or high-scale, cloud-native environments is a strong plus
  • 09Strong Linux and security fundamentals
  • 10Bachelor’s degree in Computer Science or equivalent hands-on experience

Условия

  • 01#LI_Hybrid
  • 02The Okta Experience Supporting Your Well-Being
  • 03Driving Social Impact
  • 04Developing Talent and Fostering Connection + Community
  • 05We are intentional about connection
  • 06Our global community, spanning over 20 offices worldwide, is united by a drive to innovate
  • 07Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one