Together AI3 days ago
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Salary not specified
MARKET
14,583 ₽median for this role
Backend Developer · 47 jobs with disclosed pay
5,092half of the offers: 11,042–19,804308,000
The employer didn't disclose pay — compare with the market yourself.
London
Responsibilities
- 01Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions.
- 02Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop.
- 03Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
- 04Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
- 05Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform.
- 06Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks.
Requirements
- 01Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
- 02Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
- 03Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
- 04Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
- 05A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.
What we offer
- 01Equal Opportunity Employer offering equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.