Together AI8 дней назад
Technical Support Engineer (GPU Clusters) - US Weekends
Зарплата не указана
РЫНОК
50 000 ₽медиана по профессии
IT Support / Helpdesk · 7 вакансий с указанной зарплатой
6 250половина предложений: 27 093–50 00065 000
Работодатель не указал зарплату — сравните с рынком сами.
Remote
Обязанности
- 01Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge Kubernetes GPU clusters; ensure swift and effective solutions every time
- 02Act as a customer facing SRE to ensure our customer’s Kubernetes clusters remain healthy and stable
- 03Become a product expert in our GPU Cluster service, serving as the last line of technical defense before issues are escalated to Engineering and Product teams
- 04Monitor GPU cluster health and proactively communicate hardware issues to customers (thermal throttling, BMC failures, missing GPUs, and NVLink/InfiniBand degradation) with clear remediation steps
- 05Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management
- 06Investigate and resolve storage and networking issues such as Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies on bare-metal and VM environments
- 07Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction
- 08Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together’s roadmap (e.g., future models to support)
- 09Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs to facilitate knowledge sharing with team and customers
- 10Be flexible in providing support coverage during holidays, nights and weekends as required by business needs to ensure consistent and reliable service for our customers
Требования
- 013+ years of experience in a customer-facing technical role with at least 1 year in a support function for an AI service or supporting a mission-critical API in SaaS
- 02Experience as an SRE or DevOps engineer working with Kubernetes
- 03Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments
- 04Advanced knowledge with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible) high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting and programming languages
- 05Experience with HPC/Slurm cluster environments — node draining, job scheduling, maintenance workflows
- 06Familiarity with high-speed networking concepts — InfiniBand, RDMA, network interface diagnostics
- 07Experience with distributed storage systems (e.g., Weka, NFS) and troubleshooting I/O and bandwidth issues
- 08Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters
- 09Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution
- 10Ability to work cross-functionally with teams such as Sales, Engineering, Support, Product and Research to drive customer success
- 11Strong sense of ownership and willingness to learn new skills to ensure both team and customer success
- 12Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders
- 13Ability to operate in dynamic environments, adept at managing multiple projects, and comfortable with frequent context switching and prioritization
Условия
- 01Fulltime position working US daytime hours
- 024-day shift, 10 hours per day, with 2 additional hours of on-call coverage on Saturdays and Sundays
- 03Role would start as a Monday to Friday role for the first few months to allow for ramping up and learning from teammates
- 04After being considered fully ramped, the role would transition to the 4-day weekend shift
- 05Competitive compensation, startup equity, health insurance, and other benefits
- 06Flexibility in terms of remote work
- 07US base salary range for this full-time position is: $160K - $230K + equity + benefits