Crusoe10 дней назад
Applied AI Inference Engineer
Зарплата не указана
Полная занятостьОфис
Обязанности
- 01Bring current inference techniques into production and refine them
- 02Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches
- 03Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems
- 04Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models
- 05Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic
- 06Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service
- 07Build and support the software and product features around the inference stack in a production setting, using one or more general-purpose languages, with Python preferred given how central it is to ML work
- 08Experiment quickly: take fuzzy goals, shape them into clear specs and focused proofs of concept, run fast experiments to find what works, and ship well-tested results without delay
- 09Own delivery end to end, from the first experiment through to the optimization running in production, keeping the underlying performance goals, clear specs, and follow-through front of mind, and drafting features and product requirement documents together with other engineering and product teams
- 10Work through ambiguity and make sound calls on tradeoffs and tooling, steering away from complexity that is not needed
Требования
- 01A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
- 02Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python
- 03Familiarity with methods for optimizing LLMs for high throughput / low latency inference
- 04Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level
- 05A firm grasp of how GPUs are built and how they behave
- 06Clear interest and hands-on experience with large language models
- 07A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models
- 08Strong communication skills, particularly when explaining hard technical topics to customers and teammates
Условия
- 01Competitive compensation and equity packages
- 02Restricted Stock Units
- 03Paid time off, paid holidays & leave of absence programs
- 04Comprehensive health, dental & vision insurance
- 05Employer contributions to HSA account
- 06Paid parental leave
- 07Paid life insurance, short-term and long-term disability
- 08Professional development & tuition reimbursement
- 09Mental health & wellness support
- 10Commuter benefits (parking & transit)
- 11Cell phone benefits