Cohere10 дней назад

Engineering Manager, GPU Infrastructure

Зарплата не указана
Полная занятостьОфис

Обязанности

  • 01Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement
  • 02Manage performance, career development, and hiring for team members
  • 03Conduct regular 1:1s and team meetings to ensure alignment and address challenges
  • 04Provide technical guidance and support to team members on complex infrastructure problems
  • 05Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling
  • 06Oversee the implementation of workload scheduling and queuing, hardware fault detection, and performance optimization systems
  • 07Collaborate with cloud providers and MLEs to adapt our training and inference stack to bleeding-edge GPU architectures
  • 08Ensure infrastructure reliability, scalability, and security across all GPU environments
  • 09Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions
  • 10Work with the research teams on training software stack adaptation for new GPU architectures
  • 11Coordinate with Capacity EPM and Finance to manage capacity of a rapidly growing compute footprint
  • 12Interface with Legal and Security teams on compliance requirements
  • 13Collaborate with other infrastructure teams on shared goals and dependencies
  • 14Establish observability and monitoring frameworks for GPU utilization, performance, and reliability
  • 15Drive practices and policies to automate cluster provisioning and management
  • 16Drive cost optimization initiatives while maintaining performance standards
  • 17Manage vendor relationships and contract negotiations for hardware and cloud services

Требования

  • 01Experience managing engineering or SRE teams with a focus on technical mentorship and growth
  • 02Strong communication skills to translate complex technical concepts for diverse audiences
  • 03Ability to make data-informed decisions under pressure
  • 04Experience working in remote, distributed teams
  • 05Commitment to fostering an inclusive and collaborative team culture
  • 06Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing environments
  • 07Proven experience with Kubernetes at scale: deployment, management, and troubleshooting cloud-native clusters for AI workloads in multi-cloud environments
  • 08Knowledge of infrastructure monitoring tools (Prometheus, Grafana)
  • 09Familiarity with Terraform, ArgoCD, or other IaC tools
  • 10Experience with cost optimization and capacity planning for GPU infrastructure
  • 11Track record of collaborating with AI researchers or ML engineers to solve infrastructure challenges
  • 12Strong problem-solving abilities with a data-driven approach
  • 13Passion for enabling AI research through robust infrastructure
  • 14Collaborative mindset with a focus on cross-team success
  • 15Willingness to learn and adapt in a fast-paced, evolving environment

Условия

  • 01A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch
  • 02Full health and dental benefits, including a separate budget for mental health
  • 03RRSP matching, 401K, Pension Scheme
  • 04100% Parental Leave top-up for up to 6 months, for either parent
  • 05Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit
  • 06Education & learning stipend for conferences, courses, and coaching
  • 076 weeks of paid vacation (30 working days!)
  • 08Budget for traveling to other offices if you are remote, plus an annual company offsite
  • 09Cohere is remote-friendly
  • 10For those in the office: a daily lunch program, plenty of snacks, and regular community and social events
  • 11For those not near an office: a co-working benefit
Engineering Manager, GPU Infrastructure · Rekru