SpaceXAI4 дня назад

Site Reliability Engineer - Memphis

Зарплата не указана
РЫНОК
12 500медиана по профессии
Site Reliability Engineer (SRE) · 11 вакансий с указанной зарплатой
5 442половина предложений: 7 266–15 66726 667
Работодатель не указал зарплату — сравните с рынком сами.
Southaven

Обязанности

  • 01Own monitoring architecture and signal quality: what we alert on, suppress, and trust.
  • 02Consume NOC noise-disposition feedback to drive suppression and redesign.
  • 03Treat alert noise as a design failure, not an operator failure.
  • 04Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • 05Run blameless postmortems and drive corrective actions to closed, not filed.
  • 06Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • 07Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current.
  • 08Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • 09Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • 10Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.

Требования

  • 01Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 025+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • 03Proven large-scale incident command experience and calm technical leadership on a bridge.
  • 04Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • 05Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • 06Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • 07Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • 08Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • 09Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
  • 10Experience in AI/ML infrastructure or supercomputing environments.
  • 11Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • 12Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • 13Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • 14Prior work in a fast-paced startup or tech company like SpaceXAI.

Условия

  • 01SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.
  • 02Our team is small, highly motivated, and focused on engineering excellence.
  • 03This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
  • 04We operate with a flat organizational structure.
  • 05All employees are expected to be hands-on and to contribute directly to the company’s mission.
  • 06Leadership is given to those who show initiative and consistently deliver excellence.
  • 07Work ethic and strong prioritization skills are important.
  • 08All employees are expected to have strong communication skills.
  • 09They should be able to concisely and accurately share knowledge with their teammates.