ElevenLabs9 дней назад

Research Engineer - Inference

Зарплата не указана
РЫНОК
26 250медиана по профессии
AI Researcher · 36 вакансий с указанной зарплатой
5 600половина предложений: 16 785–53 975118 000
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьУдалёнка

Обязанности

  • 01Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure
  • 02Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels
  • 03Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters
  • 04Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics

Требования

  • 01Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications
  • 02Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang)
  • 03The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it

Условия

  • 01Remote work with option to work from offices in London, New York, San Francisco, and Warsaw
  • 02Annual discretionary stipend for professional development
  • 03Annual discretionary stipend for social travel to meet colleagues
  • 04Annual company offsite in new locations
  • 05Monthly co-working stipend if not near main hubs