Together AI4 days ago
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Salary not specified
MARKET
14,583 ₽median for this role
Backend Developer · 47 jobs with disclosed pay
5,092half of the offers: 11,042–19,804308,000
The employer didn't disclose pay — compare with the market yourself.
Amsterdam
Responsibilities
- 01Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions.
- 02Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop.
- 03Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
- 04Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
- 05Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform.
- 06Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks.
Requirements
- 01Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
- 02Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
- 03Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
- 04Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
- 05A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.
- 06Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
- 07Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
- 08Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
- 09Systems programming in Rust or Go.