SpaceXAI4 дня назад
Site Reliability Engineer - Memphis
Зарплата не указана
РЫНОК
13 625 ₽медиана по профессии
Site Reliability Engineer (SRE) · 10 вакансий с указанной зарплатой
5 442половина предложений: 6 800–15 91726 667
Работодатель не указал зарплату — сравните с рынком сами.
Southaven
Обязанности
- 01Own monitoring architecture and signal quality: what we alert on, suppress, and trust.
- 02Consume NOC noise-disposition feedback to drive suppression and redesign.
- 03Treat alert noise as a design failure, not an operator failure.
- 04Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
- 05Run blameless postmortems and drive corrective actions to closed, not filed.
- 06Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
- 07Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current.
- 08Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
- 09Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
- 10Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.
Требования
- 01Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
- 025+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
- 03Proven large-scale incident command experience and calm technical leadership on a bridge.
- 04Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
- 05Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
- 06Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- 07Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
- 08Excellent problem-solving skills with a data-driven approach to reliability engineering.
- 09Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
- 10Experience in AI/ML infrastructure or supercomputing environments.
- 11Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
- 12Experience running game days, dependency mapping, and closed-loop corrective action programs.
- 13Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
- 14Prior work in a fast-paced startup or tech company like SpaceXAI.
Условия
- 01SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.
- 02Our team is small, highly motivated, and focused on engineering excellence.
- 03This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
- 04We operate with a flat organizational structure.
- 05All employees are expected to be hands-on and to contribute directly to the company’s mission.
- 06Leadership is given to those who show initiative and consistently deliver excellence.
- 07Work ethic and strong prioritization skills are important.
- 08All employees are expected to have strong communication skills.
- 09They should be able to concisely and accurately share knowledge with their teammates.