Scale AI7 дней назад
AI Infrastructure Engineer, Serving Platform
Зарплата не указана
London
Навыки
PythonGoRustC++LLM servingrate limitingtoken streamingload balancingbudgetsreasoningtool callingprompt templatesDockerKubernetesAWSGCPTerraformvLLMSGLangTensorRT-LLMtext-generation-inference
Обязанности
- 01Design and build platforms for scalable, reliable, and efficient serving of LLMs
- 02Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale
- 03Build an internal platform to empower LLM capability discovery
- 04Collaborate with researchers and engineers to integrate and optimize models for production and research use cases
- 05Conduct architecture and design reviews to uphold best practices in system design and scalability
- 06Develop monitoring and observability solutions to ensure system health and performance
- 07Lead projects end-to-end, from requirements gathering to implementation, in a cross-functional environment
Требования
- 014+ years of experience building large-scale, high-performance backend systems
- 02Strong programming skills in one or more languages (e.g., Python, Go, Rust, C++)
- 03Experience with LLM serving and routing fundamentals (e.g., rate limiting, token streaming, load balancing, budgets)
- 04Experience with LLM capabilities and concepts such as reasoning, tool calling, prompt templates
- 05Experience with containers and orchestration tools (e.g., Docker, Kubernetes)
- 06Familiarity with cloud infrastructure (AWS, GCP) and infrastructure as code (e.g., Terraform)
- 07Proven ability to solve complex problems and work independently in fast-moving environments
- 08Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference (nice to have)