Harvey1 день назад

Staff Software Engineer, Production Engineering

Зарплата не указана
Полная занятостьУдалёнка

Навыки

AWSAzureGoogle Cloud PlatformKubernetesTerraformPulumiInfrastructure-as-CodeIAMmonitoringloggingalerting

Обязанности

  • 01Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads
  • 02Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations
  • 03Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost
  • 04Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions
  • 05Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely
  • 06Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship
  • 07Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance
  • 08Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads
  • 09Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth
  • 10Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation
  • 11Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring
  • 12Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls
  • 13Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi
  • 14Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform
  • 15Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements

Требования

  • 0110+ years of software, infrastructure, site reliability, or production engineering experience
  • 02Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform
  • 03Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability
  • 04Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics
  • 05Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi
  • 06Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations
  • 07Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response
  • 08Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices
  • 09A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority
  • 10Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders
  • 11A systems-thinking mindset and a passion for building simple, reliable, and scalable infrastructure