Together AI13 дней назад
Senior Software Engineer — Infra Agent Systems (Remote, India)
Зарплата не указана
РЫНОК
14 500 ₽медиана по профессии
Backend Developer · 80 вакансий с указанной зарплатой
5 092половина предложений: 11 708–19 3981,2 млн
Работодатель не указал зарплату — сравните с рынком сами.
India
Обязанности
- 01Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
- 02Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
- 03Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
- 04Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
- 05Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
- 06Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
- 07Turn what agents learn in production into reliable, reviewed software and automation.
Требования
- 015+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
- 02Strong systems design skills and experience owning significant systems from design through production.
- 03Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding
- 04Knowledge graphs or graph data modeling
- 05Search, retrieval, ranking, RAG, or semantic search systems
- 06Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
- 07Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
- 08Comfortable working across languages such as Go, TypeScript, Python, or Rust.
- 09Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
- 10Graph databases
- 11Event-driven systems and messaging platforms such as NATS or Kafka
- 12Observability platforms such as Prometheus and Grafana
- 13Building evaluation frameworks or improving the quality and reliability of LLM-powered systems
Условия
- 01Remote based in India
- 02Equal Opportunity Employer