Harvey1 день назад
Senior Software Engineer, Production Engineering
Зарплата не указана
Полная занятостьУдалёнка
Обязанности
- 01Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads
- 02Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations
- 03Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost
- 04Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions
- 05Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely
- 06Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship
- 07Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance
- 08Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads
- 09Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth
- 10Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation
- 11Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring
- 12Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls
- 13Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi
- 14Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform
- 15Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements
Требования
- 015+ years of software, infrastructure, site reliability, or production engineering experience
- 02Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform
- 03Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability
- 04Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics
- 05Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi
- 06Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations
- 07Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response
- 08Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices
- 09A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority
- 10Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders
- 11A systems-thinking mindset and a passion for building simple, reli