Modal06.07.2026
Специалист по технической разработке (Inference Research)
Зарплата не указана
РЫНОК
26 250 ₽медиана по профессии
AI Researcher · 36 вакансий с указанной зарплатой
5 600половина предложений: 16 785–53 975118 000
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьNew York
Обязанности
- 01Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic
- 02Train custom speculators against real production traffic and feed what you learn back into target models
- 03Work directly with customers alongside our Forward Deployed Engineers to deploy and tune models, and bring what you learn back into the research
- 04Carry and expand collaborations with outside research labs
- 05Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine
- 06Help shape the research agenda
Требования
- 01A research-leaning or systems background in LLM inference, with work you can point to
- 02Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
- 03A record of shipping research or systems that other people build on, whether in a lab or in industry
- 04The drive to independently take a research bet from idea to result, working in the open with the rest of the team
- 05Ability to work in-person, in our NYC or San Francisco office