OpenAI26.06.2026

Agent Post-Training, Personality

Зарплата не указана
Полная занятостьSan Francisco

Обязанности

  • 01Develop a rigorous understanding of what makes an agent a great collaborator across professional, creative, technical, and everyday work
  • 02Turn qualitative judgments about model behavior into concrete hypotheses, evals, graders, and training interventions
  • 03Study explicit and implicit user signals to understand which behaviors create trust, satisfaction, continued use, and successful outcomes
  • 04Work with human experts and trainers to produce high-quality, tasteful rollouts and preference data that capture excellent collaborative behavior
  • 05Improve reward models and RL objectives for model behaviors
  • 06Work with pretraining and early-training teams on data mixtures, objectives, synthetic data, and other upstream choices that shape downstream personality
  • 07Build sustainable pipelines for updating older training data as our understanding of excellent model behavior evolves
  • 08Partner closely with ChatGPT, Codex, and other product teams to turn consumer insight into model improvements and validate them in real workflows
  • 09Own projects end to end, from observing a subtle behavioral failure through experimentation, training, evaluation, and launch

Требования

  • 01Think instinctively from the user’s perspective and care deeply about how models feel to work with, not only how they perform on benchmarks
  • 02Can translate subjective-seeming product questions into falsifiable hypotheses and rigorous evaluations without losing the nuance that made the question important
  • 03Care about preserving individuality, adaptability, and behavioral diversity rather than optimizing every model toward one narrow style
  • 04Want to shape how frontier agents communicate, collaborate, and build trust with millions of people
  • 05Have strong technical foundations in machine learning, software engineering, statistics, behavioral science, HCI, or a related field, and can quickly learn across unfamiliar parts of the stack
  • 06Have strong taste for model behavior: you can look at user feedback and can explain why one response feels thoughtful, natural, and useful while another does not
  • 07Have experience with LLMs, post-training, RL/RLHF, reward modeling, evals, synthetic data, pretraining data, or production ML systems
  • 08Are excited by ambiguous capability problems where the signal is noisy, the failures are qualitative, and the solution may involve data, training, evals, product changes, or all of the above
  • 09Can work effectively with researchers, engineers, product teams, designers, domain experts, human-data teams and safety boundaries, and can communicate clearly with each group
  • 10Like building load-bearing systems and processes when that is what the team needs, even if the work is not glamorous
  • 11Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users
Agent Post-Training, Personality · Rekru