OpenAI11 days ago

Machine Learning Engineer, API Multicloud

Salary not specified
MARKET
15,900median for this role
Data Scientist · 115 jobs with disclosed pay
5,000half of the offers: 12,796–20,50043,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьSan Francisco

Responsibilities

  • 01Partner with strategic customers and internal teams to define target model behaviors, diagnose failure modes, and translate real-world needs into training, evaluation, and system requirements
  • 02Build and scale production ML systems for model customization, post-training, and fine-tuning-as-a-service workflows
  • 03Investigate whether training and customization workflows are producing the intended outcomes, and identify changes to data, evaluation, training, or infrastructure that improve performance
  • 04Partner with backend and infrastructure engineers to integrate ML capabilities into AWS-native API environments
  • 05Feed learnings from partner deployments back into the platform by proposing and implementing improvements to post-training systems, tooling, APIs, and developer workflows
  • 06Work closely with Research and Applied teams to bring model improvements, training workflows, and evaluation best practices into production
  • 07Help design systems that allow strategic partners and enterprise customers to safely customize OpenAI models for high-value use cases
  • 08Debug and improve complex systems spanning model behavior, training data, APIs, distributed infrastructure, and customer-facing product surfaces
  • 09Operate with high ownership in a 0→1 environment where requirements are ambiguous, systems are evolving quickly, and reliability matters

Requirements

  • 01Master's or PhD in Computer Science, Machine Learning, or a related field, or equivalent practical experience
  • 027+ years of professional engineering experience in relevant ML, infrastructure, or product-driven engineering roles
  • 03Strong ML engineering experience building, training, fine-tuning, evaluating, or deploying production AI systems, with hands-on experience in deep learning, transformer models, and frameworks like PyTorch or TensorFlow
  • 04Familiarity with training and fine-tuning large language models, including methods like supervised fine-tuning, distillation, preference optimization, reinforcement learning, or other post-training techniques
  • 05Strong software engineering fundamentals, including data structures, algorithms, systems design, and high-quality production code in Python, Rust, or similar languages
  • 06Experience with model customization, evaluation systems, data pipelines, distributed systems, cloud infrastructure, or production ML platform tradeoffs
  • 07Ability to operate across model behavior, APIs, and infrastructure, while collaborating closely with Research, Safety, product engineering, infrastructure, and external technical partners
  • 08Comfort moving quickly through ambiguity, owning problems end-to-end, and learning whatever is needed to get the job done

What we offer

  • 01Bonus: experience with AWS, Kubernetes, agents, tool use, runtime environments, AI developer platforms, or speech models