Pinterest1 день назад
Sr. Site Reliability Engineer, tvScientific
Зарплата не указана
San Francisco
Обязанности
- 01Ensuring the reliability, availability, and performance of production infrastructure and platform services
- 02Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads
- 03Managing GitOps-based deployment workflows using ArgoCD and Helm
- 04Driving infrastructure provisioning and change management through Terraform/Terragrunt
- 05Building and supporting CI/CD automation and deployment workflows using GitHub Actions
- 06Leading incident response efforts, root cause analysis, and post-incident improvement initiatives
- 07Reducing operational toil through scripting, tooling, and process automation
- 08Advancing observability practices across logs, metrics, traces, dashboards, and alerting
- 09Supporting secure secrets integration, IAM-aware operations, and platform guardrails
- 10Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes
Требования
- 014+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure
- 02Strong hands-on experience operating AWS in production environments
- 03Deep expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration
- 04Proven experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns
- 05Experience implementing and operating ArgoCD within a GitOps delivery model
- 06Strong hands-on experience with Helm
- 07Strong experience with Terraform/Terragrunt for infrastructure provisioning and environment management
- 08Solid scripting and automation skills using Bash and/or Python
- 09Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions
- 10Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems
- 11Experience with monitoring, alerting, and observability in production environments
- 12Demonstrated ownership mindset with experience handling incidents, resolving production issues, and driving follow-through after outages
- 13Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams
- 14Bachelor’s degree in computer science, engineering, a related field or equivalent experience
- 15Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
- 16Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
- 17High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables
Условия
- 01US based applicants only
- 02Salary range: $139,764 — $287,749 USD
- 03This position is not eligible for relocation assistance
- 04Remote work option available