Cohere10 дней назад
Engineering Manager, GPU Infrastructure
Зарплата не указана
Полная занятостьОфис
Обязанности
- 01Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement
- 02Manage performance, career development, and hiring for team members
- 03Conduct regular 1:1s and team meetings to ensure alignment and address challenges
- 04Provide technical guidance and support to team members on complex infrastructure problems
- 05Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling
- 06Oversee the implementation of workload scheduling and queuing, hardware fault detection, and performance optimization systems
- 07Collaborate with cloud providers and MLEs to adapt our training and inference stack to bleeding-edge GPU architectures
- 08Ensure infrastructure reliability, scalability, and security across all GPU environments
- 09Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions
- 10Work with the research teams on training software stack adaptation for new GPU architectures
- 11Coordinate with Capacity EPM and Finance to manage capacity of a rapidly growing compute footprint
- 12Interface with Legal and Security teams on compliance requirements
- 13Collaborate with other infrastructure teams on shared goals and dependencies
- 14Establish observability and monitoring frameworks for GPU utilization, performance, and reliability
- 15Drive practices and policies to automate cluster provisioning and management
- 16Drive cost optimization initiatives while maintaining performance standards
- 17Manage vendor relationships and contract negotiations for hardware and cloud services
Требования
- 01Experience managing engineering or SRE teams with a focus on technical mentorship and growth
- 02Strong communication skills to translate complex technical concepts for diverse audiences
- 03Ability to make data-informed decisions under pressure
- 04Experience working in remote, distributed teams
- 05Commitment to fostering an inclusive and collaborative team culture
- 06Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing environments
- 07Proven experience with Kubernetes at scale: deployment, management, and troubleshooting cloud-native clusters for AI workloads in multi-cloud environments
- 08Knowledge of infrastructure monitoring tools (Prometheus, Grafana)
- 09Familiarity with Terraform, ArgoCD, or other IaC tools
- 10Experience with cost optimization and capacity planning for GPU infrastructure
- 11Track record of collaborating with AI researchers or ML engineers to solve infrastructure challenges
- 12Strong problem-solving abilities with a data-driven approach
- 13Passion for enabling AI research through robust infrastructure
- 14Collaborative mindset with a focus on cross-team success
- 15Willingness to learn and adapt in a fast-paced, evolving environment
Условия
- 01A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch
- 02Full health and dental benefits, including a separate budget for mental health
- 03RRSP matching, 401K, Pension Scheme
- 04100% Parental Leave top-up for up to 6 months, for either parent
- 05Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit
- 06Education & learning stipend for conferences, courses, and coaching
- 076 weeks of paid vacation (30 working days!)
- 08Budget for traveling to other offices if you are remote, plus an annual company offsite
- 09Cohere is remote-friendly
- 10For those in the office: a daily lunch program, plenty of snacks, and regular community and social events
- 11For those not near an office: a co-working benefit