Flexport6 дней назад

Старший инженер-программист: Platform SRE

Зарплата не указана
РЫНОК
13 625медиана по профессии
Site Reliability Engineer (SRE) · 10 вакансий с указанной зарплатой
5 442половина предложений: 6 800–15 91726 667
Работодатель не указал зарплату — сравните с рынком сами.
Amsterdam

Обязанности

  • 01Designing, building and operating reliable, scalable infrastructure in AWS and other cloud platforms
  • 02Evolving our container orchestration (Kubernetes), Helm charts, deployment systems (Argo) and the guardrails around them to ensure engineers can safely deploy, run, understand and debug their services
  • 03Evolving our observability stack (Openmetrics, Datadog) to make insights into our systems as good and easy to understand as possible
  • 04Setting standards for our Infrastructure-as-Code codebases (Terraform)
  • 05Creating and supporting developers with CI/CD pipelines (Buildkit, Github Actions), build tooling (Gradle, Bazel, npm/pnpm/bun, Go and Cargo) and artifact repositories (Artifactory, ECR, github)
  • 06Improving our incident response tooling to ensure we can efficiently handle incidents when they occur
  • 07Crossing organizational boundaries to align with product teams and ship changes that impact the whole engineering org

Требования

  • 0110+ years of software engineering experience, site reliability engineering or infrastructure engineering experience, with significant time spent on production system infrastructure
  • 02Deep knowledge of cloud infrastructure, with hands-on AWS experience
  • 03Strong Infrastructure-as-Code skills, preferably with Terraform and a track record of creating safe and reusable infrastructure patterns
  • 04Clickops should be your tool of last resort
  • 05Solid software engineering skills and the ability to build automation and production tooling in a general-purpose programming language
  • 06Experience leading incident response, conducting blameless post-mortems or COE processes and executing on the follow ups
  • 07An independent mindset: you thrive in ambiguity, find the highest-leverage problems yourself, and don't wait to be told what to do
  • 08Sound judgment around infrastructure: you will make the right decisions around security, IAM, networking and change management
  • 09Excellent communication: you can write a clear design or incident document, explain operational tradeoffs, and align stakeholders across teams on platform-wide changes
  • 10A generally helpful attitude: everyone should know you as the person who gets stuff unblocked, points them in the right direction and makes their work easier
  • 11A drive for impact: you're not satisfied until our systems are reliable and every engineer is confident operating what they own

Условия

  • 01Office presence 3 days a week for collaboration, whiteboarding and shipping
  • 02Access to latest hardware and software, including frontier AI models
  • 03Agile, flexible work practices decided by teams
  • 04Opportunity to contribute to one of the fastest-growing companies with global impact