OktaНовая13 часов назад
Staff Site Reliability Engineer, Federal (TS/SCI)
Зарплата не указана
Washington
Навыки
AWSGCPKubernetesTerraformHelmGolangPythonPostgreSQLRedisOpenSearchMySQLCassandraDNSTLSCI/CDGitOps
Обязанности
- 01Design, build, and operate large-scale cloud infrastructure and production services
- 02Participate in an on-call rotation supporting highly available customer-facing systems
- 03Lead incident response efforts and drive post-incident reviews focused on systemic improvements
- 04Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
- 05Partner with engineering teams to improve service availability, scalability, performance, and resilience
- 06Continuously improve observability through metrics, logging, tracing, dashboards, and alerting
- 07Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies
- 08Eliminate operational toil through automation, tooling, and platform engineering
- 09Improve deployment safety and operational workflows through CI/CD and GitOps practices
- 10Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities
- 11Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security
- 12Lead complex reliability initiatives spanning multiple engineering teams
- 13Guide engineers in adopting operational best practices and reliability engineering principles
- 14Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing
- 15Influence architecture and operational decisions through data-driven recommendations and engineering expertise
- 16Drive projects from conception through production rollout and long-term operational ownership
- 17Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation
- 18Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity
Требования
- 01Active U.S. TS/SCI clearance with Full Scope Poly
- 02Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6)
- 03Strong experience operating large-scale production services in AWS and/or GCP
- 04Deep expertise with Kubernetes in production environments
- 05Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues
- 06Extensive experience with Infrastructure as Code technologies such as Terraform and Helm
- 07Strong software engineering skills in Golang and/or Python
- 08Experience building automation and internal engineering platforms
- 09Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies
- 10Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management
- 11Experience with observability platforms, monitoring strategies, and production telemetry
- 12Experience with or strong interest in AI-assisted engineering and operational automation
- 13Strong expertise operating customer-facing production systems
- 14Experience leading incident response and driving operational improvements
- 15Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning
- 16Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices
- 17Proven ability to balance reliability, scalability, security, and engineering velocity
- 18Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design
- 19Experience implementing operational controls and best practices in regulated or security-sensitive environments is a plus
- 20Demonstrated success leading complex engineering initiatives across multiple teams
- 21Strong collaboration and communication skills
- 22Experience working effectively within globally distributed engineering organizations spanning multiple timezones and cultures
- 23Experience mentoring engineers and elevating technical capabilities within an organization
- 24Ability to influence technical direction through expertise, partnership, and execution
- 25Experience operating SaaS platforms serving large-scale customer workload