Cursor13 дней назад

Software Engineer, Pretraining

Зарплата не указана
РЫНОК
26 250медиана по профессии
AI Researcher · 36 вакансий с указанной зарплатой
5 600половина предложений: 16 785–65 900143 300
Работодатель не указал зарплату — сравните с рынком сами.
Полная занятостьОфис

Обязанности

  • 01Build and own high-throughput, fully telemetered data pipelines that process frontier-scale data with end-to-end traceability
  • 02Train and ship models that classify, rank, filter, clean, and identify data at extreme throughput
  • 03Design and run scaling-ladder experiments on data-mixture, repeatability, and quality depth
  • 04Partner tightly with Data Acquisition to hunt down missing or low-quality sources, and with the training teams to close the loop on what actually moves loss and downstream evals
  • 05Build the platform that turns raw web, code, multimodal, and acquired data into training-ready datasets for frontier pretraining runs
  • 06Own the pipelines, orchestration, and tooling that make pretraining data iteration fast, reliable, observable, and reproducible at scale
  • 07Create clear signals for data quality, lineage, freshness, and pipeline health so researchers can trust what goes into each run
  • 08Build and scale the web crawling systems that discover, schedule, fetch, and parse high-quality documents across the open web for initial training
  • 09Improve URL seeding, scoring, and fair host scheduling so crawl capacity lands on the hosts and pages that matter most for model quality
  • 10Raise crawl success and parsing quality — defeating antibot failures, improving extractors, and capturing content we previously could not get cleanly
  • 11Debug and harden complex crawl infrastructure end-to-end for availability, recovery, and ingestion lag, and automate delivery of crawl datasets into the data pipeline

Требования

  • 01You have a strong infrastructure or data platform background, and ideally a spike of outlier depth somewhere (crawling/search infra is a plus, not a hard requirement)
  • 02You are a high-slope engineer who has moved unusually fast — for example, staff-level ownership within a few years — or you bring deep domain experience
  • 03You are able to architect and ship end-to-end with high ownership, debug complex systems independently, and work alongside AI agents
  • 04You have strong intuitions about large-scale distributed systems
  • 05You're excited to learn how pre-training data shapes model quality, and want the ownership and visibility that comes with building systems that feed frontier training runs