Together AI13 дней назад

Senior Software Engineer — Infra Agent Systems (Remote, India)

Зарплата не указана
РЫНОК
14 500медиана по профессии
Backend Developer · 80 вакансий с указанной зарплатой
5 092половина предложений: 11 708–19 3981,2 млн
Работодатель не указал зарплату — сравните с рынком сами.
India

Обязанности

  • 01Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • 02Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • 03Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
  • 04Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
  • 05Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
  • 06Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
  • 07Turn what agents learn in production into reliable, reviewed software and automation.

Требования

  • 015+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
  • 02Strong systems design skills and experience owning significant systems from design through production.
  • 03Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding
  • 04Knowledge graphs or graph data modeling
  • 05Search, retrieval, ranking, RAG, or semantic search systems
  • 06Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
  • 07Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
  • 08Comfortable working across languages such as Go, TypeScript, Python, or Rust.
  • 09Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
  • 10Graph databases
  • 11Event-driven systems and messaging platforms such as NATS or Kafka
  • 12Observability platforms such as Prometheus and Grafana
  • 13Building evaluation frameworks or improving the quality and reliability of LLM-powered systems

Условия

  • 01Remote based in India
  • 02Equal Opportunity Employer