Synthesia07/17/2026

Staff Research Engineer - Multimodal Generative Modelling

Salary not specified
MARKET
15,900median for this role
Data Scientist · 115 jobs with disclosed pay
5,000half of the offers: 12,796–20,50043,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьУдалёнка

Responsibilities

  • 01Shape the roadmap to create new model capabilities and unlock new functionality for the customer base, on both short and long time horizons
  • 02Propose novel multi-modal system architectures (especially text and voice)
  • 03Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis
  • 04Design solutions that reinforce emotional expressiveness and natural interaction
  • 05Implement and bring designs to life, from pretraining through post-training
  • 06Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness
  • 07Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements
  • 08Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs
  • 09Curate new datasets to complement existing data
  • 10Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality
  • 11Ship models to production with optimized runtime to serve customers, and address their feedback thereafter

Requirements

  • 01The ability to bring novel ideas and designs that advance the field of interactive multimodal systems
  • 02Strong understanding of generative modelling, ideally applied to sequential or multimodal data
  • 03Hands-on experience with large language models or similar transformer-based architectures
  • 04High proficiency in PyTorch, including distributed training and model optimization
  • 05A solid grasp of time-series modeling and tokenization, preferably in the context of audio, speech, or video
  • 06A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently
  • 07Proven experience training deep learning models end-to-end, from data preparation through evaluation
  • 08Strong general software engineering skills, enabling contributions to a large, shared research infrastructure
  • 09Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it
  • 10Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints
  • 11Working on LLMs with large scale trainings leading to models with decent reasoning capabilities
  • 12Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment
  • 13Collaborating across modalities or teams to ship a unified system