OpenAI26.06.2026
Agent Post-Training, Personality
Зарплата не указана
Полная занятостьSan Francisco
Обязанности
- 01Develop a rigorous understanding of what makes an agent a great collaborator across professional, creative, technical, and everyday work
- 02Turn qualitative judgments about model behavior into concrete hypotheses, evals, graders, and training interventions
- 03Study explicit and implicit user signals to understand which behaviors create trust, satisfaction, continued use, and successful outcomes
- 04Work with human experts and trainers to produce high-quality, tasteful rollouts and preference data that capture excellent collaborative behavior
- 05Improve reward models and RL objectives for model behaviors
- 06Work with pretraining and early-training teams on data mixtures, objectives, synthetic data, and other upstream choices that shape downstream personality
- 07Build sustainable pipelines for updating older training data as our understanding of excellent model behavior evolves
- 08Partner closely with ChatGPT, Codex, and other product teams to turn consumer insight into model improvements and validate them in real workflows
- 09Own projects end to end, from observing a subtle behavioral failure through experimentation, training, evaluation, and launch
Требования
- 01Think instinctively from the user’s perspective and care deeply about how models feel to work with, not only how they perform on benchmarks
- 02Can translate subjective-seeming product questions into falsifiable hypotheses and rigorous evaluations without losing the nuance that made the question important
- 03Care about preserving individuality, adaptability, and behavioral diversity rather than optimizing every model toward one narrow style
- 04Want to shape how frontier agents communicate, collaborate, and build trust with millions of people
- 05Have strong technical foundations in machine learning, software engineering, statistics, behavioral science, HCI, or a related field, and can quickly learn across unfamiliar parts of the stack
- 06Have strong taste for model behavior: you can look at user feedback and can explain why one response feels thoughtful, natural, and useful while another does not
- 07Have experience with LLMs, post-training, RL/RLHF, reward modeling, evals, synthetic data, pretraining data, or production ML systems
- 08Are excited by ambiguous capability problems where the signal is noisy, the failures are qualitative, and the solution may involve data, training, evals, product changes, or all of the above
- 09Can work effectively with researchers, engineers, product teams, designers, domain experts, human-data teams and safety boundaries, and can communicate clearly with each group
- 10Like building load-bearing systems and processes when that is what the team needs, even if the work is not glamorous
- 11Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users