DeepL9 дней назад
Senior Research Scientist | Model Steering
Зарплата не указана
Полная занятостьУдалёнка
Обязанности
- 01Drive the development of translation models that are steerable conditioned on user preferences, rules and context
- 02Drive hands-on research and development on post-training for our core translation models: supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning tuned to translation quality
- 03Build reward models and evaluator models for translation, including rubric- and reference-based grading, and investigate and mitigate reward hacking and quality-estimation failure modes
- 04Drive an agenda toward models that ingest multimodal content and context to increase translation quality
- 05Own the full lifecycle of model delivery: prototyping, ablations, training, evaluation, optimization, and production deployment, working closely with engineering to ship into real-time systems at scale
- 06Establish strong practices for evaluation, reproducibility, monitoring, and continuous model improvement in production
- 07Mentor researchers and engineers, promote hands-on collaboration, and raise the bar for model quality
Требования
- 01Proven experience making large models steerable and instruction-following by identifying the most effective method to instill a given behavior, drawing from instruction tuning, latent space methods, steering vectors, and/or constrained encoding and decoding methods
- 02Deep, hands-on expertise in LLM post-training (SFT, DPO), knowledge distillation (teacher-student training), and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, and reward modeling)
- 03Strong data-centric instincts for building synthetic-data and preference-data pipelines, LLM-as-judge generation, data curation and filtering, and reasoning about data mixtures and ablations
- 04Experience designing evaluation and reward signals using automatic metrics, LLM-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation
- 05A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production while staying grounded in product impact and real-world quality
- 06Ownership of a substantial research direction with strong execution, and experience mentoring others on a fast-moving, applied research team
- 07Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow), and the ability to communicate clearly and align research with product and engineering priorities