Writer1 день назад

Infrastructure engineer (UK)

Зарплата не указана
Полная занятостьУдалёнка

Обязанности

  • 01Bring deep focus to one problem at a time while moving between SRE, DevOps, Infrastructure, and Platform work over a quarter or two, owning substantial initiatives such as on-call posture, release pipeline, multi-region Terraform layout, or internal platform surface
  • 02Challenge the status quo and remove toil before adding features by automating operational tasks and infrastructure management with Python or Go, rejecting unsuitable tools, and treating manual on-call work as a defect to be designed out
  • 03Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently with Kubernetes, Helm, Terraform, and supporting cloud and AI tooling that backs WRITER's high-traffic platform
  • 04Run AI agents (Claude Code, Droid, Codex, internal skills) in your daily loop to investigate incidents, draft Terraform/Helm changes, write runbooks, scaffold tooling, and review PRs, building an agentic setup where humans and digital teammates share skills, context, and on-call workflows
  • 05Lead incident response, post-mortems, and root-cause analyses, tracing failures to underlying problems, applying learnings to architecture, and preventing recurrence of the same incident
  • 06Own the reliability, performance, and efficiency of WRITER's core services end-to-end by defining and upholding SLOs and error budgets, carrying the on-call pager, and standing behind outcome metrics
  • 07Balance short-term critical work with 6–12-month platform direction, shipping immediate fixes while shaping multi-year observability, cost, and reliability investments
  • 08Collaborate cross-functionally with product, security, and engineering peers to provide expert guidance on system design for reliability, performance, and scalability, connecting infrastructure agenda to product and revenue context and disagreeing with evidence

Требования

  • 015+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
  • 02Experience running containerisation in production (a real cluster, not a lab) with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred), plus good proficiency in Python or Go for automation and tooling
  • 03AI is part of how you ship: agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, you have built or adopted AI-assisted workflows others now use, and you have strong opinions on where it's unreliable
  • 04Demonstrated ability to challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems by reasoning from constraints and failure modes and naming tradeoffs in business terms

Условия

  • 01Hybrid position based out of New York City or London hubs
  • 02Reports to the Director of Engineering
Infrastructure engineer (UK) · Rekru