SpaceXAI3 дня назад

Network Operations Center Specialist - Memphis

Зарплата не указана
РЫНОК
12 500медиана по профессии
Site Reliability Engineer (SRE) · 11 вакансий с указанной зарплатой
5 442половина предложений: 7 266–15 66726 667
Работодатель не указал зарплату — сравните с рынком сами.
Southaven

Обязанности

  • 01Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches
  • 02Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving
  • 03Detect, verify, and escalate within time budgets; operate the escalation matrix (NOC → on-call SRE → domain owners) and page correctly the first time
  • 04Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls
  • 05Produce first-pass RCA framing (what happened, when, what’s impacted, who’s engaged) and hand it to SRE / Hardware Failure Analysis for depth — the NOC does not publish root cause
  • 06Run structured shift handoffs and durable shift logs; maintain cross-site awareness
  • 07Write major-incident reports; open corrective projects in Linear and chase them to closure — the NOC is the nag of record
  • 08Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days

Требования

  • 01Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent)
  • 02Proven ability to acknowledge, classify, and escalate incidents under SLA in a high-signal environment
  • 03Experience opening and running incident bridges, including stakeholder updates on a fixed cadence and live timeline hygiene
  • 04Excellent written and verbal communication skills; able to write clear updates while an incident is in progress
  • 05Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact
  • 06Experience following, maintaining, and improving operational process (runbooks, escalation matrices, handoffs, or similar)
  • 07Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage

Условия

  • 01Work in a flat organizational structure
  • 02Rotating shift schedule including nights and weekends
  • 03Fast-paced startup or tech company environment