Scale AI7 дней назад

AI Infrastructure Engineer, Serving Platform

Зарплата не указана
London

Навыки

PythonGoRustC++LLM servingrate limitingtoken streamingload balancingbudgetsreasoningtool callingprompt templatesDockerKubernetesAWSGCPTerraformvLLMSGLangTensorRT-LLMtext-generation-inference

Обязанности

  • 01Design and build platforms for scalable, reliable, and efficient serving of LLMs
  • 02Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale
  • 03Build an internal platform to empower LLM capability discovery
  • 04Collaborate with researchers and engineers to integrate and optimize models for production and research use cases
  • 05Conduct architecture and design reviews to uphold best practices in system design and scalability
  • 06Develop monitoring and observability solutions to ensure system health and performance
  • 07Lead projects end-to-end, from requirements gathering to implementation, in a cross-functional environment

Требования

  • 014+ years of experience building large-scale, high-performance backend systems
  • 02Strong programming skills in one or more languages (e.g., Python, Go, Rust, C++)
  • 03Experience with LLM serving and routing fundamentals (e.g., rate limiting, token streaming, load balancing, budgets)
  • 04Experience with LLM capabilities and concepts such as reasoning, tool calling, prompt templates
  • 05Experience with containers and orchestration tools (e.g., Docker, Kubernetes)
  • 06Familiarity with cloud infrastructure (AWS, GCP) and infrastructure as code (e.g., Terraform)
  • 07Proven ability to solve complex problems and work independently in fast-moving environments
  • 08Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference (nice to have)