Sesame15.03.2025

ML Model Serving Engineer

Зарплата не указана
Полная занятостьОфис

Обязанности

  • 01Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models.
  • 02Partner with ML infrastructure and training engineers to build a fast, cost-effective, accurate, and reliable serving layer to power a new consumer product category.
  • 03Modify and extend LLM serving frameworks like VLLM and SGLang to take advantage of the latest techniques in high-performance model serving.
  • 04Work with the training team to identify opportunities to produce faster models without sacrificing quality.
  • 05Use techniques like in-flight batching, caching, and custom kernels to speed up inference.
  • 06Find ways to reduce model initialization times without sacrificing quality.

Требования

  • 01Expert in some differentiable array computing framework, preferably PyTorch.
  • 02Expert in optimizing machine learning models for serving reliably at high throughput, with low latency.
  • 03Significant systems programming experience; ex. Experience working on high-performance server systems—you’d be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase.
  • 04Significant performance engineering experience; ex. Bottleneck analysis in high-scale server systems or profiling low-level systems code.
  • 05Always up to date on the latest techniques for model serving optimization.
  • 06Familiarity with high-performance LLM serving; ex. experience with VLLM, SGlang deployment, and internals.
  • 07Experience with a public cloud platform such as GCP, AWS, or Azure.
  • 08Experience deploying and scaling inference workloads in the cloud using Kubernetes, Ray, etc.
  • 09You like to ship and have a track record of leading complex multi-month projects without assistance.
  • 10You’re excited to learn new things and work in a multitude of roles.

Условия

  • 01401(k) max employer match: 3.5% of compensation
  • 02100% employer-paid health, vision, and dental benefits for you and your dependents
  • 03Unlimited PTO and sick time
  • 04Flexible spending account with employer matching up to $1,650/year (medical FSA)
  • 05Guardian Employee Assistance Program (EAP)
  • 06Opportunity to share in the company's success with competitive stock options
  • 07Full-time position
ML Model Serving Engineer · Rekru