Synthesia07/17/2026
Staff Research Engineer - Multimodal Generative Modelling
Salary not specified
MARKET
15,900 ₽median for this role
Data Scientist · 115 jobs with disclosed pay
5,000half of the offers: 12,796–20,50043,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьУдалёнка
Responsibilities
- 01Shape the roadmap to create new model capabilities and unlock new functionality for the customer base, on both short and long time horizons
- 02Propose novel multi-modal system architectures (especially text and voice)
- 03Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis
- 04Design solutions that reinforce emotional expressiveness and natural interaction
- 05Implement and bring designs to life, from pretraining through post-training
- 06Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness
- 07Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements
- 08Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs
- 09Curate new datasets to complement existing data
- 10Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality
- 11Ship models to production with optimized runtime to serve customers, and address their feedback thereafter
Requirements
- 01The ability to bring novel ideas and designs that advance the field of interactive multimodal systems
- 02Strong understanding of generative modelling, ideally applied to sequential or multimodal data
- 03Hands-on experience with large language models or similar transformer-based architectures
- 04High proficiency in PyTorch, including distributed training and model optimization
- 05A solid grasp of time-series modeling and tokenization, preferably in the context of audio, speech, or video
- 06A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently
- 07Proven experience training deep learning models end-to-end, from data preparation through evaluation
- 08Strong general software engineering skills, enabling contributions to a large, shared research infrastructure
- 09Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it
- 10Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints
- 11Working on LLMs with large scale trainings leading to models with decent reasoning capabilities
- 12Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment
- 13Collaborating across modalities or teams to ship a unified system