OpenAI26.06.2026

Agent Post-Training, Computer Use Research

Зарплата не указана
Полная занятостьSan Francisco

Обязанности

  • 01Design and run experiments that improve agentic model behavior for complex computer use
  • 02Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis
  • 03Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions
  • 04Partner with Codex and ChatGPT product teams to understand what users need and translate product signal into model improvements
  • 05Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior
  • 06Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs
  • 07Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness
  • 08Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness
  • 09Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes

Требования

  • 01Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field
  • 02Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems
  • 03Ability to move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next
  • 04Comfortable working across research, product, infrastructure, data, evals, and safety boundaries
  • 05Ability to communicate clearly with different groups
  • 06Willingness to build load-bearing systems and processes when needed
Agent Post-Training, Computer Use Research · Rekru