Together AI8 дней назад

Инженер по развертыванию (Inference & Post-Training) - владение китайским

Зарплата не указана
РЫНОК
65 900медиана по профессии
AI Researcher · 19 вакансий с указанной зарплатой
13 833половина предложений: 30 208–85 64097 160
Работодатель не указал зарплату — сравните с рынком сами.
Singapore

Обязанности

  • 01Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
  • 02Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and latency targets
  • 03Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production
  • 04Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic accounts — monitoring and optimizing endpoint configurations, helping customers get the most out of the platform, and collaborating to ensure we hit critical milestones
  • 05Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure the right inference and post-training configurations are in place from day one to improve time-to-value
  • 06Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from the field. Contribute back to the product where needed to support customer requirements or drive a better experience. Drive early feature and research adoption with strategic logos

Требования

  • 015+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows
  • 02Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang); ability to diagnose and resolve performance issues at the engine level
  • 03Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
  • 04Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO; ability to advise on system design
  • 05Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets
  • 06Strong Python skills; comfortable working in production environments
  • 07Must be a permanent resident or citizen of Singapore

Условия

  • 01Competitive compensation
  • 02Startup equity
  • 03Health insurance
  • 04Flexibility in terms of remote work
  • 05Salary ranges determined by location, level and role