Together AI4 days ago

Старший инженер-программист, инфраструктура для вывода моделей / вычислительная инфраструктура

Salary not specified
MARKET
14,583median for this role
Backend Developer · 47 jobs with disclosed pay
5,092half of the offers: 11,042–19,804308,000
The employer didn't disclose pay — compare with the market yourself.
India

Responsibilities

  • 01Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions
  • 02Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop
  • 03Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically
  • 04Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service
  • 05Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform
  • 06Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks

Requirements

  • 01Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living
  • 02Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution
  • 03Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines)
  • 04Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts
  • 05A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship

What we offer

  • 01Remote in India