Together AI19 days ago
Инженер по развертыванию (Inference & Post-Training) - владение китайским
Salary not specified
MARKET
65,900 ₽median for this role
AI Researcher · 37 jobs with disclosed pay
5,600half of the offers: 16,833–78,900143,300
The employer didn't disclose pay — compare with the market yourself.
Singapore
Responsibilities
- 01Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
- 02Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and latency targets
- 03Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production
- 04Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic accounts — monitoring and optimizing endpoint configurations, helping customers get the most out of the platform, and collaborating to ensure we hit critical milestones
- 05Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure the right inference and post-training configurations are in place from day one to improve time-to-value
- 06Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from the field. Contribute back to the product where needed to support customer requirements or drive a better experience. Drive early feature and research adoption with strategic logos
Requirements
- 015+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows
- 02Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang); ability to diagnose and resolve performance issues at the engine level
- 03Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
- 04Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO; ability to advise on system design
- 05Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets
- 06Strong Python skills; comfortable working in production environments
- 07Must be a permanent resident or citizen of Singapore
What we offer
- 01Competitive compensation
- 02Startup equity
- 03Health insurance
- 04Flexibility in terms of remote work
- 05Salary ranges determined by location, level and role