Modal07/06/2026
Специалист по технической разработке (Inference Research)
Salary not specified
MARKET
65,900 ₽median for this role
AI Researcher · 37 jobs with disclosed pay
5,600half of the offers: 16,833–78,900143,300
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьNew York
Responsibilities
- 01Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic
- 02Train custom speculators against real production traffic and feed what you learn back into target models
- 03Work directly with customers alongside our Forward Deployed Engineers to deploy and tune models, and bring what you learn back into the research
- 04Carry and expand collaborations with outside research labs
- 05Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine
- 06Help shape the research agenda
Requirements
- 01A research-leaning or systems background in LLM inference, with work you can point to
- 02Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
- 03A record of shipping research or systems that other people build on, whether in a lab or in industry
- 04The drive to independently take a research bet from idea to result, working in the open with the rest of the team
- 05Ability to work in-person, in our NYC or San Francisco office