Together AI8 дней назад
Технический специалист поддержки (Inference) - US Weekends
Зарплата не указана
РЫНОК
13 796 ₽медиана по профессии
IT Support / Helpdesk · 24 вакансий с указанной зарплатой
6 250половина предложений: 9 792–42 50064 000
Работодатель не указал зарплату — сравните с рынком сами.
Remote
Обязанности
- 01Engage directly with customers to tackle and resolve complex technical challenges involving cutting-edge GPU clusters and inference and fine-tuning services, ensuring swift and effective solutions
- 02Act as a customer-facing SRE to ensure customer’s Inference endpoints (running on Kubernetes) remain healthy, stable, and performant
- 03Become a product expert in all Gen AI solutions, serving as the last line of technical defense before issues are escalated to Engineering and Product teams
- 04Assist with hardware and platform migrations by validating system health and traffic routing
- 05Monitor dashboards to detect anomalies and escalate with data-backed analysis
- 06Manage customer-facing communications during incidents and degradations; translate deep technical findings (latency regressions, provider issues, network reachability drops) into clear, evidence-backed updates without exposing platform internals
- 07Contribute infrastructure changes for model deployment, capacity rebalancing, and cluster configuration; execute infrastructure changes via pull requests (infra-as-code) for tasks such as endpoint configuration, model bringup/bringdown, and capacity scaling
- 08Flag engine-level bugs with logs and reproduction steps for engineering
- 09Collaborate seamlessly across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders both internally and externally to ensure the highest levels of customer satisfaction
- 10Transform customer insights into action by identifying patterns in support cases and working with Engineering and Go-To-Market teams to drive Together’s roadmap (e.g., future models to support)
- 11Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs to facilitate knowledge sharing with team and customers
- 12Be flexible in providing support coverage during holidays, nights and weekends as required by business needs to ensure consistent and reliable service for our customers
Требования
- 016+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, with at least 1 year in a support role for an AI service
- 02Experience as an SRE or DevOps engineer working with Kubernetes
- 03Strong technical background, with knowledge of AI, ML, GPU technologies and their integration into high-performance computing (HPC) environments
- 04Advanced, production-level experience with infrastructure services (e.g., Kubernetes, SLURM), infrastructure as code solutions (e.g., Ansible), high-performance network fabrics, NFS-based storage management, and container infrastructure
- 05Familiarity with operating storage systems in HPC environments such as Vast and Weka
- 06Proven ability to diagnose complex network-layer issues and read traces
- 07Strong knowledge of Python, TypeScript, and/or JavaScript with testing/debugging experience using curl and Postman-like tools
- 08Demonstrated expertise with observability tooling (e.g., Prometheus, Grafana) at scale
- 09Deep familiarity with REST API debugging and HTTP semantics
- 10Experience with LLM inference frameworks and LoRA fine-tuning and common training failure modes
- 11Experience with Infrastructure as Code and Git-based workflows
- 12Background in GPU cluster management
- 13Cloud platform experience (AWS, GCP, and/or Azure)
- 14Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters
- 15Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution
- 16Ability to work cross-functionally with teams such as Sales, Engineering, Support, Product and Research to drive customer success
- 17Strong sense of ownership and willingness to learn new skills to ensure both team and customer success
- 18Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders
- 19Ability to operate in dynamic environments, adept at managing multiple projects, and comfortable with frequent context switching and prioritization
Условия
- 01Full-time position, US daytime hours
- 024-day shift, 10 hours per day, with 2 additional hours on-call coverage on Saturdays and Sundays
- 03Works both weekend days (Saturday and Sunday) plus two additional weekdays
- 04Starts as Monday‑to‑Friday for initial months, then transitions to 4‑day weekend shift after ramp‑up
- 05Competitive compensation, startup equity, health insurance