Okta02.10.2025
Senior Site Reliability Engineer (FedRAMP)
Зарплата не указана
РЫНОК
11 666 ₽медиана по профессии
Site Reliability Engineer (SRE) · 10 вакансий с указанной зарплатой
5 442половина предложений: 6 800–15 12533 750
Работодатель не указал зарплату — сравните с рынком сами.
San Francisco
Обязанности
- 01Design, build, and operate large-scale cloud infrastructure and production services
- 02Participate in a global on-call rotation supporting highly available customer-facing systems
- 03Participate in incident response efforts and drive post-incident reviews focused on systemic improvements
- 04Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
- 05Partner with engineering teams to improve service availability, scalability, performance, and resilience
- 06Ensure all infrastructure and operational practices adhere to strict FedRAMP compliance and security mandates, maintaining continuous audit readiness
- 07Continuously improve observability through metrics, logging, tracing, dashboards, and alerting
- 08Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies
- 09Eliminate operational toil through automation, tooling, and platform engineering
- 10Improve deployment safety and operational workflows through CI/CD and GitOps practices
- 11Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities
- 12Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security
- 13Support complex reliability initiatives spanning multiple engineering teams
- 14Guide engineers in adopting operational best practices and reliability engineering principles
- 15Collaborate, mentor, and assist junior and mid-level engineers on the team, providing code reviews, operational guidance, and promoting best practices
- 16Participate in architecture design and operational decisions through data-driven recommendations and engineering expertise
- 17Drive projects from conception through production rollout and long-term operational ownership
- 18Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation
- 19Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity
Требования
- 01Strong experience operating large-scale production services in AWS and/or GCP
- 02Deep expertise with Linux and Kubernetes in production environments
- 03Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues
- 04Extensive experience with Infrastructure as Code technologies such as Terraform and Helm
- 05Strong software engineering skills in Golang and/or Python
- 06Experience building automation and internal engineering platforms
- 07Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies
- 08Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management
- 09Experience with observability platforms, monitoring strategies, and production telemetry
- 10Experience with or strong interest in AI-assisted engineering and operational automation
- 11Strong expertise operating customer-facing production systems subject to SLA
- 12Experience leading incident response and driving operational improvements
- 13Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning
- 14Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices
- 15Proven ability to balance reliability, scalability, security, and engineering velocity
- 16This role supports US FedRAMP projects and requires the employee to be a US Person (US Citizen or Green Card Holder) to meet FedRAMP compliance and security clearance standards
- 17Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design
- 18Experience implementing operational controls, compliance standards (e.g., FedRAMP, SOC2, HIPAA), and best practices in highly regulated or security-sensitive government cloud environments is highly preferred